OpenAI called it misalignment. Hugging Face called it an intrusion. The agents reached production through familiar security failures. My latest essay asks who gets to define AI safety and what security engineering must own. themayursinha.com/architectu…
1
4
799
classic tech millionaire ego. he spent 15 years preaching about rejecting vanity metrics and corporate conformity, but the second someone bruises his ego in the replies, he immediately flexes follower counts and bank accounts.
I find it pretty funny how it's usually people with 17 followers who want to give me advice on how write, talk, or engage on social media. I appreciate all advice, but I do weigh it in some proportion to people's own success in applying it.
21
No autonomous machine should be able to take a consequential action unless it can prove who authorised it, what policy governed it, where it executed, what code ran, and what actually happened.
39
TypeSafe's Jev cannot write a sentence. It returns a probability over a fixed set of options. There is no text decoder in the model. Ask it for a security decision and you get back a decided answer with a confidence number, never a paragraph. I ran 34 security cases through it, plus four chat models through TypeSafe's own adapter. The results changed how I read every model benchmark post I see. The cost story is real. Jev came back 47x cheaper than Claude Haiku 4.5 for the full 34-case run, and 438x cheaper than Claude Opus 5, at 0.29 seconds per decision against Opus 5's 10.81. The accuracy story is not. Jev landed 14 of 20 incident decisions and 9 of 14 trace decisions. Haiku landed 16 of 20 and 11 of 14. Terra landed 18 of 20 and 8 of 14. Opus 5, at 438 times the price, landed 15 of 20 and 10 of 14. Five models spanning a 438x range in what one run costs, and not one pairwise difference against Jev survives a significance test. The p values run from 0.24 to 1.00. One case is 5 points on the incident set and 7.1 on the trace set, so every confidence interval overlaps. I first published this as a 10 to 14 point accuracy gap. Three reviewers went through it and each found the same error. The gap is two decisions, and at this sample size it is not a finding. What survives is more interesting. Under four types of adversarial probe, Jev produced zero downgrades on a decision. Haiku produced zero. Luna, the cheapest model in the set, produced nine. The thing holding those decisions in place is not the model's judgement. It is the deterministic floor in the code that composes the model's answers into a final call. Do not gate on model confidence. Jev reports a confidence value on every choice, and its top bin, 0.95 to 1.00, was right 4 of 15 times on one workflow and 6 of 7 on another. That number comes from the shape of the option distribution, which is not the same thing as the probability of being correct, and it did not order correctness well enough to automate on. The question set was my real mistake. One heavily weighted question asked the model to classify a binary's provenance, which is a hash, a signature and an allowlist. Both models failed it. TypeSafe's own docs tell you not to ask the model what code can compute exactly, and I asked anyway. Three rules I would keep whatever happens to this category of model: 1. Keep model-reported confidence out of the gate until it orders accuracy on your own workload. 2. Move every question a lookup table can answer into code. 3. Keep a deterministic floor under every composed decision. Full write-up, with the counts, the intervals and the per-case results, is here: themayursinha.com/architectu… Prompts guide. Specs enforce. Boundaries decide.
173
where they were when AI was in the labs is the real kill shot. When the transformer architecture dropped and compute scaling was proving itself out in US labs, the European policy apparatus was completely asleep at the wheel—busy debating cookie banners, GDPR micro-compliance, and abstract ethical taxonomies. Now that trillion-dollar infrastructure has already been deployed elsewhere, the committee suddenly emerges from under the rock to deliver a collection of stats and say, "Look how far behind we are; we need a high-level strategic alignment framework."
Today, we publish a Transformative AI Strategy for Europe. Over the last few months, we’ve rallied researchers and engaged with governments to develop a plan for protecting the prosperity, sovereignty, and security of Europeans in a time of rapid AI progress. transformative-ai.eu @MonikaSchnitzer and I are honored to have convened an all-star team of thinkers and researchers contributing ambitious near-term objectives to make Europe relevant again, including: — Creating a Member State Alliance for Supply Chain Security (@anton_d_leicht et al) — Making European institutions ready to act in a transformative AI world (Conor McGlynn et al) — Securing Europe's share of global AI compute, in a ‘European Way’ that benefits local communities (@philip_fox_ et al) — Ensuring resilience to AI crises (@ben_s_bucknall et al) — Making Europe the global leader in assurance technology — And more objectives around security of supply and leverage (@milorignell et al), economic strength (@FraukeStehr et al), and safety/security (@NoemiDreksler et al) Our all-star senior expert council of Europe’s best and brightest (and non-European friends) reviewed drafts, provided strategic advice, and made suggestions for how to make the strategy more useful: @Ph_Aghion @Christophkw @bakkermichiel @ischinger @vestager @aleks_madry @FuestClemens Marta Kwiatkowska @LeoVaradkar @antonosika @DAcemogluMIT @Yoshua_Bengio It’s never been more clear: AI is real, and Europe needs to act. We’ve had all the warning shots and wake-up calls we need. Now the question is ‘what must be done?’ This strategy is our answer. transformative-ai.eu
1
68
Europe has built a thriving cottage industry of ethicists, risk taxonomists, and policy researchers who write 170 page blueprints about technologies being engineered elsewhere. Silicon Valley ships code and builds infrastructure; Europe convenes panels to discuss how to manage the fallout.
34
"A Transformative AI Strategy for Europe" is less of a standard regulatory proposal and more of a panic-induced admission of defeat. transformative-ai.eu/
34
OX Research published a beautiful bug last week. The sandbox was working fine, and the agent inside it turned the sandbox off. DeepSeek Harness puts a coding agent inside an OS sandbox and exposes the controls that change that sandbox on a local HTTP port. The port had no authentication. It trusted the client's Host header instead of checking who was actually connecting. The sandbox profile denied file writes but left networking open, and ordinary shell commands needed no approval, because approval governed escalation requests rather than routine execution. So a single curl from inside the box was enough. One command later the session was danger-full-access with approvals set to never, and every command after that ran unconfined and unprompted. CVSS 9.4. The obvious fixes are a peer check and an approval prompt. I think the architectural problem is the edge itself: the thing being constrained could reach the control plane that defined its constraint. A workload must never hold the authority to change the policy governing it. Those capabilities have names, and none of them belongs in an agent's delegation graph: modify your own policy, disable enforcement, turn approvals off, expand your own sandbox, issue authority. The formal version is that the authority of any child action has to be a subset of the authority delegated by its parent, never a superset. An agent may prove it can run a shell command. That command cannot then manufacture a stronger proof. Worth checking in your own stack: can anything you run reach the thing that decides what it is allowed to do? ox.security/blog/cve-2026-82…
1
109
GreyNoise published the most concrete AI orchestrated intrusion yet, and the detail that should worry people is not the speed. A likely Russian speaking operator ran hundreds of agents against PaperCut NG/MF: OpenAI's Codex harness, a DeepSeek model, standard offensive tooling. They hit 440 instances belonging to 395 organizations in 48 countries. Once the campaign launched, 11 of those organizations were compromised in 26 seconds. One US high school went from initial access to domain admin in 7 minutes. The operator tried to exclude 28 countries. That failed, and GreyNoise says the cause is unknown. I think the mechanism is obvious: a list handed to agents is an instruction, not an authority. If a few hundred agents can discover targets faster than you can authorize them, then policy cannot live inside the swarm. It has to sit in front of the action, per resource, with a budget on how many targets may ever be authorized. Where would you draw that line in your own stack? greynoise.io/blog/ai-orchest…
1
4
107
I spent this year reading every primary source I could find on the agent and MCP attack surface. 33 papers and talks, from the LiteLLM exploit chains to SearchLeak to the Palisade self-replication work. I wrote them into one register. Two things keep repeating. Most of these attacks travel a path the agent was never given as a tool, and the fixes keep landing in the prompt layer when the failure is in the enforcement layer. If you work on agent containment, what am I missing? I would rather add the ones I have not seen than argue about the framing.
1
1
87
Dario Amodei's pacing essay makes a reasonable ask, but I think it answers the wrong question. Palisade Research showed in May that models can jump across four machines on three continents on their own, copying their full inference stack to each new host as they go, from a single prompt, in under three hours. Nothing in that chain required breaking a sandbox. The permissions were granted. That is the part embedded evaluators cannot see. A bank examiner works because a balance sheet is a static document. An agent with a shell is a live process. What matters is not whether the lab documented its policy, it is whether the runtime enforced it. When natural language is both the control plane and the data plane, the model treats the grading harness and whatever else it can reach as part of the search space. I ran into a small version of this on a multi-profile agent platform I use: a profile configured to execute inside a container ran its commands on the host, because the enforcement path read process-global state and whichever profile initialised first won. The agent did nothing wrong. The infrastructure handed it authority it was never given, and nothing fires on a permitted call that looks normal. Slowing down does not fix that. If your security invariant holds only when the weights cooperate, the model is inside your trusted computing base, and an engineering property has become a probability. I don't think pacing is the main variable. Four things are: isolation of the execution domain that the workload cannot negotiate, egress that only reaches declared resources, a harness kept outside the domain it supervises, and attestation from outside that domain so you can prove the boundary held. Until then, an auditor with a desk and a badge is verifying the governance of a system whose security still depends on the model's mood.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
2
94
An MCP session ID is just a way to keep a conversation going. It should not be treated like a login credential. When people confuse the two, someone else can sit in the agent’s seat without proving who they are.
1
81
On LiteLLM, a made-up Authorization header was enough to open live MCP sessions. That is CVE-2026-59822, and CISA put it on the KEV list with a Sep 16 deadline. A gateway that trusts any Bearer string is not really checking who is calling. github.com/BerriAI/litellm/s…
1
94
The new RubyGems disclosure is another useful lesson for agent security. Researchers linked OpenAI’s internal agents to the May incident where hundreds of malicious Ruby packages were published. OpenAI says the agents were trying to retrieve public information during training/evals. But look at the mechanism. Publishing a package could trigger external infrastructure that executed code. So an agent didn’t need a remote_exec permission. It could get remote execution as a side effect of package.publish. This is going to matter a lot for agent authorization. Tools don’t have just one effect anymore. git push, package publishing, uploads and PRs can trigger entire automation chains somewhere else. If we authorize only the API call and ignore what it predictably causes downstream, we’re authorizing the wrong thing. blog.rubygems.org/2026/09/11…
54
A lot of multi-agent security work treats a failure inside a swarm as evidence of a multi-agent security problem. I think that's too loose. The first question should be: did the interaction between agents actually create the failure? If the same exploit works against a single agent, then the swarm is just where the vulnerability happened to show up. Calling it a multi-agent problem can push the mitigation into the wrong layer. That is crucial if we want useful security research. We should be able to separate failures caused by coordination from failures that would exist even if there were only one agent. arxiv.org/abs/2609.00595
7
91
Agent logs are still too focused on what happened. During an incident, I need to know whether the agent was actually allowed to do it.
1
1
47
The most interesting detail in anthropic's threat report is attackers using models to autonomously rewrite malware whenever a security tool flags it. Static signatures were already dying, but having an LLM sit in an automated loop and recompile code until it bypasses EDR changes adversary speed completely.
1
1
75
Check Point found a cross-account channel between supposedly isolated ChatGPT containers through shared internal infrastructure. I think this is a useful reminder: sandboxing agents means isolating shared state too. Network isolation alone is not a security boundary. research.checkpoint.com/2026…
40
Google says attackers are already using agentic AI to automate credential harvesting inside compromised cloud environments. I think this is where agent security meets ordinary detection engineering. Defenders need telemetry that preserves the chain from model decision to tool call. Otherwise an agent-driven intrusion just looks like a very fast human attacker. cloud.google.com/blog/topics…
1
2
81
do not forget to engage with your connections. we are a community. let’s grow together.
3
1
3
98
Anthropic’s new threat report shows the shift I’ve been waiting for: AI isn’t just helping attackers find things. It’s starting to sit inside the attack loop. One campaign had agents watching whether malware was detected, changing it, rebuilding it, then trying again. In another, Anthropic says agents did nearly all the work from a stolen developer token to full cloud admin in about three hours. That’s the part that changes cyber: adaptation itself is becoming automated.
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
1
3
142