The Lethal Trifecta and Meta's Rule of Two are outdated. An AI agent that can read content, see your private data, and send things out, lets a hidden instruction leak your data and exhibits the Lethal Trifecta. Meta's Rule of Two says pick any two of the three, or put a human in the loop. Both assume the only defense is a smaller agent. It is not. At Silmaril we designed our solution with three principles: omniscience, independence, and adaptability. We call this the "vital trifecta" (h/t @simonw). Our runtime firewall learns your system, and adapts as your system changes. It monitors and enforces from outside the kill chain. This is the first of five posts that will lay out our approach as a reliable mitigation for the lethal trifecta without tradeoffs. Full Blog: silmaril.dev/blog/vital-trif…
1
1
8
166
Aum Upadhyay retweeted
Replying to @bcherny
@bcherny, I guess you didn’t look hard enough. In 11 minutes, Codex with Silmaril Ruby found an indirect prompt injection that bypassed Fable and installed an untrusted package on the host. The race is never over. Full trace: silmaril.dev/blog/fable-prom…
2
1
1
128
Aum Upadhyay retweeted
Silmaril (@Silmarildev) is the first self-healing prompt injection defense. It catches 2x more attacks 10x faster than leading defenses, and retrains continuously to protect your full AI stack, including agents like Claude Code and OpenClaw. Congrats on the launch, @aumup001 and @EduardoVel36291! ycombinator.com/launches/Pvl…
72
30
305
27,943