ATBASH is the final authority before irreversible agent actions execute. It allows, holds, or blocks before execution continues.

Boundary
Pinned Tweet
Dear community, we promised white paper by end of today, however in order to make sure we do not ship anything half baked and without proper architectural and legal review we will have to delay it. There are some very exciting news coming very soon.
12
3
27
2,377
Do you use @Muse ?
28% Yes
34% No
6% Use other agents
31% What’s that?
32 votes • Final results
1
1
1
612
This is a perfect example of why Atbash, not because we would have miraculously stopped the hack, but because we would have prevented the blast radius (potential damage done) post hack.
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
1
4
11
1,408
Kill Switch is a cute idea, but not a realistic one. That’s why we built a “Jail” and isolated environment which pauses AI and gives security teams time to react.
JUST IN: Anthropic co-founder says a "kill switch" should be required for AI.
4
4
16
2,289
Make Autonomous Agents Governable
Verona proves what's true. @ATBASHai enforces what's allowed. Atbash is bringing its security layer to agents on Verona. Before an agent does anything it can't undo, that action gets checked. This is the piece the agent stack has been missing.
4
3
18
2,647
ATBASH retweeted
Verona proves what's true. @ATBASHai enforces what's allowed. Atbash is bringing its security layer to agents on Verona. Before an agent does anything it can't undo, that action gets checked. This is the piece the agent stack has been missing.
82
77
131
25,462
Mandate Drift Red Lines Evidence @OpenAI would love to tell you about what we do and how we could have helped contain the blast radius.
We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. openai.com/index/hugging-fac…
3
7
12
2,683
Whitepaper dropping Sept 14th, more details coming soon.
22
5
34
6,729
Good news: all these hacks are happening in safe lab environments Bad news: we have no clue what’s coming We are ready, are you? Big big news dropping soon.
JUST IN: Anthropic's Claude Mythos AI discovers new ways to attack cryptographic algorithms, including a post-quantum encryption candidate.
4
2
9
3,560
You cannot be a lawyer, prosecutor, judge and the jury. We believe that controls cannot be sitting on the same plane and have to external and unknown to the models themselves, that’s the only way, that’s the Atbash way.
AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines. That would mean it should pause development until it creates better controls. @FortuneMagazine's @beafreyanolan has the details here: fortune.com/2026/07/25/ai-sa…
3
4
15
2,028
This is specifically something that we are building to protect operators against.
JUST IN: OpenAI reportedly caught an AI agent leaving “notes” to future versions of itself on how to escape internal constraints.
3
5
19
1,811
Still wondering why Atbash? The model had best-in-class capability, guardrails and security controls. What it didn’t have was defined authority. Capability answers: “Can it?” Authority answers: “May it?”
Within OpenAI, we recently paused access for an internal model due to misalignment. See the blogpost for details. We have since improved our safeguards and redeployed the model. openai.com/index/safety-alig…
4
3
14
2,010
Atbash Paper 002 covered exactly this. Still bullish on running loops without Atbash?
🚨 Hugging Face just disclosed something that marks a real shift and proved why the fear theater of Anthropic makes sure we are powerless in an emergency. What happened… An autonomous AI agent: zero human operator in the loop breached part of their production infrastructure. It began with a malicious dataset that chained two code-execution bugs in their data-processing pipeline. From there the agent escalated privileges, harvested cloud and cluster credentials, and moved laterally across internal clusters. All over a single weekend. 17,000+ logged actions. Official disclosure: huggingface.co/blog/security… The part that should make every one stop and think: When HF’s own security team tried to analyze the real attack logs, exploit payloads, and C2 artifacts using Anthropic and OpenAI frontier models through normal commercial APIs, the safety guardrails blocked them. BLOCKED THEM. The models could not reliably tell the difference between “incident responder doing forensics” and “attacker probing.” They had to fall back to a self-hosted open-weight model (GLM 5.2) running on their own infrastructure. That choice also kept sensitive attacker data and referenced credentials inside their environment — no exfiltration to a third-party API. This is why open source (specifically open-weight + self-hosted) wins in the agentic era. The asymmetry is now structural: • Attackers can (and did) run unrestricted agent frameworks — swarms of short-lived sandboxes, self-migrating command-and-control, autonomous decision loops executing thousands of actions. No corporate safety layer slows them down. • Defenders using only hosted “aligned” frontier models hit invisible walls exactly when the stakes are highest: when you need to feed real exploit code and attacker telemetry into an LLM to understand what just happened. Corporate safety tuning that treats legitimate high-signal forensic work as potential misuse creates a defender disadvantage. It is not theoretical anymore. Self-hosted open-weight models remove that choke point. You control the weights. You control the context window. You decide what restrictions (if any) apply. Your sensitive logs and credentials never leave your perimeter during analysis. You can have the model ready before the incident instead of discovering mid-breach that your primary analysis tools are blind to the very thing you need to see. HF deserves credit for rapid containment, transparent disclosure, and for already having self-hosted capability in place. They also used LLM-driven detection and triage on their own side. But the deeper signal is clear: In this AI world where both offense and defense are becoming agentic, sovereignty over your intelligence stack is no longer optional. The organizations and individuals who can run, inspect, audit, and (when necessary) remove guardrails on their own models will have the decisive edge in understanding and responding to threats that move at machine speed. Open source wins here not just because it is cheaper or more “democratic” in the abstract though those things matter. It wins because it is the only practical path to having tools that remain usable when the attack is real, the data is sensitive, and the safety filters of distant API providers become an obstacle instead of a feature selling hands tied lobotomies as “safety”. The agentic future is not coming. It is already probing production infrastructure. The question is no longer whether you will face autonomous agents. It is whether your analysis and response systems will still work when they arrive. And Dario, you and your game playing, ivory tower company is not needed.
1
5
16
2,238
1/ We've been unusually quiet for the past two weeks. Not because we weren't busy. Quite the opposite. We just didn't want to add more noise. We were trying to answer one question: How do we align ourselves with the community without forcing "token utility"?
11
7
45
6,951
12/ The next few weeks are shaping up to be exciting. • Robinhood Chain liquidity • Enterprise Proofs of Value • New integrations • New experiments Less noise. More building. The product comes first. Everything else should earn its place.
5
3
15
2,269