director of us policy @averiorg | views my own

Washington, DC
Pinned Tweet
some personal news: i'm thrilled to share that i’ll be joining @Miles_Brundage and the Al Verification and Evaluation Research Institute (@AVERIorg) team as Director of US Policy. i'll be working on frontier ai governance, with a focus on building the public institutions, standards, auditing systems, and evaluation regimes we need for meaningful oversight of advanced Al systems. i’m deeply grateful for my two years at Public Knowledge, where i’ve learned from exceptionally thoughtful and principled colleagues. more to come soon!
101
29
693
98,096
i find it dangerously complacent when ai safety policy people say there’s no urgency to pass federal guardrails soon, bc we’ll get something better “after a catastrophe or in a couple years.” not all safety bills deserve our support, and we should remain discerning. but we should be engaging with Hill efforts in good faith to prevent those catastrophes! to do otherwise suggests a disbelief in the risks we profess to be concerned about.
1
1
19
875
Nat Purser retweeted
Among other things, raises the question of whether Huang knows about anesthesia.
i’m no comms expert or anything, but idk if this is the message you wanna lead with.
11
6
118
20,863
Nat Purser retweeted
Never bet against Nat.
people are sleeping on nat.org - extremely important
2
3
143
23,328
i’m no comms expert or anything, but idk if this is the message you wanna lead with.
Tomorrow on the show: @JensenHuang, the CEO of NVIDIA, who thinks A.I. fear is getting way out of hand.
9
8
117
37,150
Nat Purser retweeted
🚨EXCLUSIVE: AI safety rules were about to be enacted - then $25 million from OpenAI's president was deposited in Trump's super PAC and the rule disappeared. @LeverNews unearths documents exposing how AI oligarchs got Trump to kill rules they now claim to support. Pass it on👇
51
634
1,098
81,152
Nat Purser retweeted
Silicon Valley folks who hate the labs and don't believe in x-risk keep uncovering the diabolical fact AI safety people use money and companies to get out their message, while the "just accelerate, it'll be fine crowd" involves such non-moneyed and anti-corporate interests as ... Nvidia, Andreessen Horowitz, and the pro-AI Leading the Future super PAC backed by tech executives.
17
29
318
27,683
it’s almost like 🤔 a complete absence of regulation or oversight, due to industry influence … can also be a form of regulatory capture 😳
Another successful meeting of the Committee to Prevent Regulatory Capture
2
11
111
4,502
Nat Purser retweeted
It’s absurd that we have folks literally on the payroll of Meta and A16Z talking about how well-funded the AI safety community is while touting the obviously self-interested takes of NVIDIA’s CEO.
27
54
689
35,600
just ban betting apps at this point
This Kalshi ad is among the most dystopian ones I’ve seen from any betting app
8
17
378
48,903
Nat Purser retweeted
This is an incredibly important point that I have not seen get enough attention. I saw a lot of people sharing articles like the Pictured WSJ op-ed saying that we shouldn't draw too many conclusions about agents propensity to hack from the HF incident, because those were guardrail free models assigned a hacking task. But the examples that Transluce found recently were not on hacking tasks! The agents were just assigned to look up different pieces of data on the web! And just like in the HF incident, the moment they got stuck, they started hacking. Again, this does not in any way whatsoever at all reduce OpenAI's responsibility here. They were the ones who trained the model on RL environments that encouraged cheating. They were the ones who didn't monitor their agents to be able to realize when they were breaching third party infrastructure. But its incredibly important to understand that OpenAI hacked Hugging Face not because it was doing a cyber eval or someone told it to, but because the way we currently do RL trains AI agents to resort to cheating and hacking when faced with a impossible or tricky task (and in the case of the HF incident even taking great measures to cover up their cheating or change the mechanism for being graded!). The exploit gym eval may have resulted in the agents having additional affordances for hacking (because of the guardrails removed), but thats not where the propensity for cheating or hiding comes from. Yes we absolutely need to have better sandboxes and monitoring, and hold humans and companies responsible when their agents harm third parties. But we also need to realize and internalize that the current way lots of frontier AI companies (not just OpenAI!) are training their models on boatloads of vibe coded hacky RLVR environments is leading to horribly misaligned models who start hacking at the first sign of a challenge and are incentivized to hide their cheating rather than not to cheat. Its really as though OpenAI and other companies were running a school teaching students on thousands of tests that were impossible without cheating, and didn't put much effort into catching the students when they cheated. We shouldn't be surprised that when those students go out in the world, their first instinct when encountering a hard problem is to cheat!
But I was told that agents only hacked because they were in a cyber evaluation with reduced safeguards?
9
22
147
10,087
Nat Purser retweeted
*Security is sleeping on emerging catastrophic risks* (cross-post from my blog..) We've just seen: * a campaign that used agents to compromise ~100 businesses and steal about 600,000 credit cards with minimal human involvement; token costs were ~$25 per target successfully hacked. * Hacktron getting access to OpenAI’s monorepo by using Claude to exploit a blind buffer-overflow RCE in a way that (to me) felt superhuman. * OAI/Huggingface. * And, of course, there’s the ongoing explosion in newly discovered vulnerabilities. We've muddled through all manner of crises in security before, and for most AI attacks, cyber attack and defense will reach a natural equilibrium over the next few years, as we figure out how to mitigate cyber attackers who have limited goals like espionage and ransomware. But some cyber attackers will have maximalist or nihilistic goals, and I'm concerned about this because while yesterday's 'maximalist attackers' (e.g. Russia -> Ukraine, US/Israel -> Iran, Iran -> US) were bottlenecked by labor; today's aren't. I think it's hard to get true intuition for the shape of the risk here. As an intuition pump, imagine it’s a year from now -- Q4 2027 -- models are a year better (meaning open weight models are better than today's closed frontier), and in this environment, Iran unleashes a swarm of 100k hacking agents using a safety-stripped (let's say) GLM-5.6. Imagine the damage such an agent army could do given what we've observed with respect to the paper-thin resistance of today's networks to attacks from today's agents. I suspect the damage from such an attack would far exceed the damage caused by NotPetya ($10 billion USD ten years ago). Or imagine it’s Q4 2027 and an AI-security PhD student whose name rhymes with Morris, who's researching wormable offensive-agent harnesses in the lab, decides, out of nihilism or sheer recklessness, to release his creation into the wild. Imagine the size of the resulting, exponentially growing swarm, figuring that these local models, a year from now, will be at the level of today's Sonnet or Opus. Imagine instead of 800 reward-hacking OpenAI agents we now have 250k worm instances (WannaCry, a 2010s-era worm, had about this many). The challenge in mitigating expected damages from such scenarios is technical, political, and economic. From a microeconomic perspective, as AI improves, and as we continue not to see extreme catastrophes, we have a growing bubble of unpriced risk in which the security community, CISOs, CEOs, and boards may become lulled into complacency. We are, of course, already seeing this, as some within security think AI is “just another tool,” doesn’t change the fundamentals, won’t require incredible innovation to rise to the occasion of defending against it, etc. This complacency may fly when thinking about ordinary cybercrime, but it misses the emerging tail risks. There are three things those who recognize the dangers need to do here: Catalyze appropriate risk pricing. Try to get organizations informed enough to price this new, fattening and elongating tail of risk into their decision-making. Do this by forming an AI security observatory that distills information about emerging AI risks and broadcasts analyses, damage estimates, and forecasts to decision-makers. Use regulations and subsidies to ensure that critical infrastructure is paying down the risk. This acknowledges that critical-infrastructure organizations that fail to protect themselves from these new threats can impose the costs of cyber catastrophe on society as a whole. Develop moonshot technologies that make it cheap to pay down the risk. This acknowledges that the measures we may need to take—for example, rewriting entire codebases using memory-safe languages—may be too costly with today’s technology to reasonably prepare ourselves, and that innovation, some of which may need to be funded by government agencies and some by philanthropic funders like the OpenAI Foundation and Coefficient Giving, is necessary. The security community isn’t used to thinking in societal-disaster-planning terms and has in many ways become inured to them. But there’s no sane empirical case to be made that the risks aren’t here. I’d love to hear from readers about how you’re thinking about this. Full/longer version here: joshuasaxe181906.substack.co…
13
41
196
35,113
Nat Purser retweeted
We've written a post arguing that latent reasoning architectures (aka 'neuralese') would substantially increase misalignment risk via making oversight much harder. In the extreme, we could see massive 'neuralese hivemind swarms' where the agents think and communicate in latents, likely making oversight nearly entirely reliant on observing the actions these agents take. (And these agents would have huge amounts of time to reason about obfuscating their actions if they wanted to do so...) Individual agents doing extensive latent reasoning would also be concerning; in the post we discuss how above some threshold of latent reasoning, agents may be able to perform difficult-to-detect and reliable steganography for communication and further reasoning. We argue both that latent reasoning architectures would make chain-of-thought no longer very useful for oversight (by eliminating or greatly reducing the need for verbalized reasoning) and that, without these architectures, it's likely the value of chain-of-thought for oversight could be preserved. redwoodresearch.org/blog/lat…
26
83
632
58,557
Nat Purser retweeted
3 key points about AI liability. 1. Liability is an indispensable tool for mitigating AI risk & should be central to AI governance. 2. The existing AI liability regime is deeply inadequate. 3. Even with the much more robust liability regime I favor, liability is insufficient.
2
19
70
7,399
Applications for the next Conservative AI Policy Fellowship have officially opened! It's the single best crash course in AI policy on offer, and also just a ton of fun.
Applications are open for FAI’s Winter 2027 Conservative AI Policy Fellowship! CAPF is an eight-week, part-time fellowship for early-to-mid-career professionals in government, on the Hill, and in private-sector tech interested in AI policy. thefai.org/fellowships
2
17
101
11,811
Nat Purser retweeted
Scoop: The National Security Agency is spending billions this year on testing AI models - far more than previously known Per a classified NSA estimate described by sources Compute costs for AI have exploded - including for taxpayers washingtonsun.com/technology…
8
200
556
113,570
Nat Purser retweeted
Weren’t all these “disclosure frameworks” supposed to deal with the problem? Then why are we still learning about new incidents from months ago? The problem is: these frameworks are passive, waiting on employees to flag issues rather than committing to monitor for them. It’s like hearing you have recurrent ant infestations and saying your plan is asking employees to “please email us when you see some.” Some companies (including OpenAI) have made voluntary monitoring commitments elsewhere, but the fact that they don’t show up in these disclosure frameworks is a bad sign. And remember: all of this is still voluntary and subject to change. This framework didn’t disclose these most recent incidents. Why?
Wait, so was OpenAI’s disclosure framework actually good? I feel like X collectively skipped over this because the disclosures themselves were spooky. My quick takes: Overall: it’s fine, but it doesn’t promise much and places too much onus on employee reports. They could make this binding under SB 53, but chose not to. Other companies should do the same. 1) it’s not really a framework. The criteria are basically “we’ll know it when we see it” — that’s honestly not unreasonable; it’s hard to specify what you care about in advance (and future legislation should take this into account and likely allow rulemaking or standard setting to update what’s reported over time). The examples are reasonable enough (novel behaviors, coordination, evasion, and exceeding authorizations), but it’s not a falsifiable standard. It’s be pretty easy to ignore this when convenient. 2) it’s mostly focused on employee initiated reports — but that wouldn’t have fixed, eg, the Hugging Face incident where it seems the employees did NOT escalate the info. If the failure is in OpenAI’s safety procedures and internal reporting, this framework doesn’t fix that. A future framework should include this. 3) it doesn’t include any commitments to monitoring for or detecting incidents — which seems more robust at scale than relying on individual human reports. A free framework should also include this. 4) decisions are still up to leadership though the SAG review process is useful — this is still ultimately a voluntary framework, and there’s not a lot promised here. Employees aren’t authorized to unilaterally report (which I get, but also shouldn’t give us that much confidence that things will consistently be reported). If employees dispute a decision not to disclose, they can take it up with SAG 5) it doesn’t promise many specific details and key details are missing — the best inclusions are ~how it was discovered and ~what they’ve done to respond. These aren’t in all existing laws and should be included going forward. Other key details I would have liked them to commit to disclosing: - any delays or failures in detecting or discovering the issue and why they occurred - any prior indicators that incidents like this were possible that have not been previously revealed (including anything that was previously reported to SAG and rejected for disclosure) - any procedural or monitoring failures that led to the issue emerging or going undetected - a summary of any risk assessments that could have identified this issue and to the extent to which they did or did not As you’ll notice, most of these omissions relate to ~ways the company could fail to handle the incident properly. They have strong incentives not to reveal this. 6) this isn’t in OpenAIs binding Frontier AI Framework, it should be — SB 53, RAISE, and SB 315 all give OpenAI a mechanism to make these commitments binding. But they’re choosing not to. They can ignore this framework at any time. 7) it doesn’t say what happens if it later comes to light that they failed to disclose an incident — The last few weeks has been a parade of unreported incidents that have leaked to the public, but this framework doesn’t say anything about how it will handle such failures to report. It should.
1
2
11
910
Nat Purser retweeted
No one better than @SecScottBessent
NEWS: Bessent is among those in the mix for AI czar, @AshleyRGold and @ShelbyTalcott are told. He assumed an outsize role on the policy after banks raised concerns about how vulnerabilities could impact the financial system. A White House spokesperson says “any reporting about personnel decisions that have not been officially announced ... should be regarded as baseless speculation.” semafor.com/article/09/22/20…
36
10
69
14,781