*Security is sleeping on emerging catastrophic risks*
(cross-post from my blog..)
We've just seen:
* a campaign that used agents to compromise ~100 businesses and steal about 600,000 credit cards with minimal human involvement; token costs were ~$25 per target successfully hacked.
* Hacktron getting access to OpenAI’s monorepo by using Claude to exploit a blind buffer-overflow RCE in a way that (to me) felt superhuman.
* OAI/Huggingface.
* And, of course, there’s the ongoing explosion in newly discovered vulnerabilities.
We've muddled through all manner of crises in security before, and for most AI attacks, cyber attack and defense will reach a natural equilibrium over the next few years, as we figure out how to mitigate cyber attackers who have limited goals like espionage and ransomware.
But some cyber attackers will have maximalist or nihilistic goals, and I'm concerned about this because while yesterday's 'maximalist attackers' (e.g. Russia -> Ukraine, US/Israel -> Iran, Iran -> US) were bottlenecked by labor; today's aren't.
I think it's hard to get true intuition for the shape of the risk here.
As an intuition pump, imagine it’s a year from now -- Q4 2027 -- models are a year better (meaning open weight models are better than today's closed frontier), and in this environment, Iran unleashes a swarm of 100k hacking agents using a safety-stripped (let's say) GLM-5.6.
Imagine the damage such an agent army could do given what we've observed with respect to the paper-thin resistance of today's networks to attacks from today's agents. I suspect the damage from such an attack would far exceed the damage caused by NotPetya ($10 billion USD ten years ago).
Or imagine it’s Q4 2027 and an AI-security PhD student whose name rhymes with Morris, who's researching wormable offensive-agent harnesses in the lab, decides, out of nihilism or sheer recklessness, to release his creation into the wild.
Imagine the size of the resulting, exponentially growing swarm, figuring that these local models, a year from now, will be at the level of today's Sonnet or Opus.
Imagine instead of 800 reward-hacking OpenAI agents we now have 250k worm instances (WannaCry, a 2010s-era worm, had about this many).
The challenge in mitigating expected damages from such scenarios is technical, political, and economic. From a microeconomic perspective, as AI improves, and as we continue not to see extreme catastrophes, we have a growing bubble of unpriced risk in which the security community, CISOs, CEOs, and boards may become lulled into complacency.
We are, of course, already seeing this, as some within security think AI is “just another tool,” doesn’t change the fundamentals, won’t require incredible innovation to rise to the occasion of defending against it, etc. This complacency may fly when thinking about ordinary cybercrime, but it misses the emerging tail risks.
There are three things those who recognize the dangers need to do here:
Catalyze appropriate risk pricing. Try to get organizations informed enough to price this new, fattening and elongating tail of risk into their decision-making. Do this by forming an AI security observatory that distills information about emerging AI risks and broadcasts analyses, damage estimates, and forecasts to decision-makers.
Use regulations and subsidies to ensure that critical infrastructure is paying down the risk. This acknowledges that critical-infrastructure organizations that fail to protect themselves from these new threats can impose the costs of cyber catastrophe on society as a whole.
Develop moonshot technologies that make it cheap to pay down the risk. This acknowledges that the measures we may need to take—for example, rewriting entire codebases using memory-safe languages—may be too costly with today’s technology to reasonably prepare ourselves, and that innovation, some of which may need to be funded by government agencies and some by philanthropic funders like the OpenAI Foundation and Coefficient Giving, is necessary.
The security community isn’t used to thinking in societal-disaster-planning terms and has in many ways become inured to them. But there’s no sane empirical case to be made that the risks aren’t here. I’d love to hear from readers about how you’re thinking about this.
Full/longer version here:
joshuasaxe181906.substack.co…