Senior Fellow @IFP. AI governance, industrial policy, abundance, appropriations.

Mark Thomas retweeted
🧵 This is entirely unsurprising. In my view, the majority's statutory interpretation gives 41 U.S.C. § 4713 pretty extraordinary reach with highly curtailed judicial review; that said, the panel's reasoning clearly follows indications from the Court re: deference on national security. There's certainly room for reversal en banc, but it's far from guaranteed.
JUST IN: D.C. Circuit upholds the Trump administraiton's decision to ice out Anthropic/Claude, finding the government's claim of security risks to have "ample support." documentcloud.org/documents/…
1
1
5
597
The level of detail in this report is a huge contribution to the embedded evaluator conversation. If you want to know how it would work and what they would actually do, read it! A few thoughts: 1. It articulates the reasons why embedded evaluators are are now very popular for AI governance—they can flexibly identify risks very early, investigate, and help design mitigations. 2. There is real cognitive capture risk—if you're a team of 1 or 2 and your daily colleagues are your client's employees, who seem like decent, well-intentioned people, it will be harder to investigate with the same rigor as if you're coming in cold. But the flip side is there are many people in the companies who wish there was more action taken on safety, and evaluators would be colleagues to collaborate with / send information to. 3. There is real tension between the recommendation that evaluators have wide-ranging access to address unknown unknowns, and the recommendation that IP be protected by limiting what evaluators can access. 4. Compute costs for evaluation may be contested. The paper mentions that METR spent $400k in tokens over 6 days investigating the Hugging Face incident. What should the evaluators' token budget be? 5. I am still not sure how much of this work can be pre-specified (e.g. 99.9% logging of internal-use inference tokens, test models' ability to break out of sandboxes) and how much has to be figured out on the fly in real time. 6. Law is out of scope for the paper, but it's worth noting that it's not clear if the government could require embedded evaluators at unwilling companies. The Fourth Amendment is often a bar to this kind of regime.
Third-party evaluators mostly test frontier AI models through an API. That misses the risks that come from how companies build and use AI internally. Our new paper makes the case for embedded assessments. It was led by Jacob Charnock. The other co-authors are Sophie Williams, @zaheedkara, @Manderljung, Alejandro Tlaie Boria, @StephenLCasper, @AnkaReuel, and me. 🔗 Read the paper: governance.ai/research-paper…
1
1
7
542
Mark Thomas retweeted
The "godfather of AI," Geoffrey Hinton, left Google so that he could sound the alarm about the dangers of the technology. You may disagree with him but it is just dishonest to paint all of the concerns as coming from industry plants and it's Q-anon level nonsense to paint it all as an elaborate psyop.
I cannot impart on progressive folks more strongly that it is not just Sam Altman sounding these alarms. I also do not care for or trust Sam Altman! But listen to the people actually building this stuff who are freaked out. Maybe they are wrong but they are not part of some profit motivated psyop.
157
132
1,086
131,077
It's now clear that AI companies cannot coordinate, even publicly, on pacing R&D without antitrust risk. An antitrust exemption is the only way to get out of this collective action problem short of regulation. The race to the bottom isn't going to fix itself!
Plaintiffs suing the leading AI companies over "pacing" argue purported agreement to reduce the rate at which their AI improves violates antitrust law. "AI alone could add trillions annually to the global economy... the product of the competitive race Defendants' agreement now threatens to slow."
4
4
31
4,351
Mark Thomas retweeted
Don't worry guys, WSJ says these three guys + Claude and Codex only got the "secret sauce" (OpenAI's algorithmic secrets) not the "crown jewels" (the weights)!! I highly doubt Chinese threat actors haven't gotten further in. Racing faster might only give us the illusion of a lead! If we were serious about securing the frontier we'd follow through on export controls + pace so labs make wiser trade-offs on speed vs. security
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
4
9
86
5,669
Mark Thomas retweeted
Takeaways from the WSJ article about @HacktronAI using Claude to get into OpenAI's monorepo and issue a pull request (before stopping and claiming their bug bounty) - how many nation states have already broken in and gone much further and stolen a) algorithmic secrets and b) model weights or c) gotten access to user data; these are fair public interest questions - how many have implants in / are dwelling in the openai network as I tweet this - how pervasive is this level of softness to pentesting across all the labs and how far are the labs from the right operating point in the security and r&d friction trade space (probably pretty far it seems) - given the hacktron folks used anthropic's models to pull this off what's the real public safety ROI of anthropic's cyber guardrails; they add friction for legitimate cyber defenders (like our developers at my startup!) but it appears with a bit of work you can use them in actual breaches as happened here - as the article says, the wsj folks and @S1r1u5_ had me do neutral technical review of the kill chain here pre publication; impressive from a human angle (@HacktronAI reminds me of the best of my generation of hackers I looked up to as a kid!) but also from what the models can do; between this and the openai/hf thing, I'm emotionally in a place where I feel my world as a security person is being turned upside down in slow motion with respect to what's coming - elite persistent hacking is becoming rapidly democratized. this is coming like a freight train. we need to harden the world's code and infra as fast as possible and today's non automated methods don't stand a chance of cutting it; the world is a soft target. I continue to be unsettled but very glad I left my comfortable job at Meta to do our automated posture hardening startup
18
45
202
20,242
Mark Thomas retweeted
As I've said, I think there's probably some truth to the idea that AI companies may be incentivized to inflate the dangers of their product for financial and strategic reasons. However, when all of the major players in an industry are telling us, flat out, that the thing they're building could be very dangerous, it seems ridiculous to write their warnings off completely. If a guy built a weird contraption, handed it to me, and said, "Careful, it might kill you," I'm going to tend to take his warning seriously. It might be a hoax, sure, but if I know considerably less about the contraption than the guy who built it, I'm probably not going to just assume it's a hoax. I'm not aware of any example in history where almost everyone INSIDE a particular industry is united in saying that their own product might be catastrophic for mankind. The climate alarmists have said that about the fossil fuel industry, for example, but the climate alarmists are not inside the industry. Notably, the fossil fuel industry itself has generally disagreed with their assessments. That's why the climate alarmism comparisons here just don't hold water. It's a totally different situation. So, while we should be cautious about taking anything at face value, I find the firm "there's nothing to worry about" claims to be pretty ridiculous. This is an extremely powerful and rapidly evolving technology. Any powerful technology carries risk. If the people building this technology say the risk is high, again it seems foolish to blithely assume that all of them are conspiring to simply lie about the dangers of their own product. Another read on the situation -- one that allows you to doubt the benevolence of these people (as I most certainly do), while not necessarily discounting all of the hazards of the thing they're building -- is that they pushed full speed ahead on this technology, blew right past the warning signs, and now they're seeing things on their end that legitimately scare them, and the PR blitz in the media right now is a classic case of corporate CYA. When the shit hits the fan, they can say, "We tried to stop it." Even though they actually could have stopped it, but didn't, when they had the chance. That seems like a plausible scenario, too.
557
167
1,958
249,603
Mark Thomas retweeted
a couple thoughts after chats w/ some non-ai-safety DC friends: i think there’s a misperception that independent evaluators and auditors imagine themselves to be a solid substitute for regulatory oversight. this is not at all what i’ve encountered. if anything, many feel under-equipped for the work they’ve been asked to do and would prefer clear rules of the road set by the government. for example — rules that ensure their access to companies’ ai systems is not contingent on the goodwill of the company they’re auditing! many of them began doing this work because there was a vacuum; neither the government nor the private sector was sufficiently monitoring ai risks. but i think if the government beefed up its capacity and committed to doing some or much of the testing in-house, many would welcome it. and there would still be room for niche, specialized evaluators as a complement to / extension of govt oversight. broadly, don’t think it’s fair or accurate to pattern match from sf tech culture onto the entire safety community in assuming these people are staunchly anti-regulation or trying to monopolize oversight / access.
6
22
144
21,731
Mark Thomas retweeted
Okay, @Arnold_Ventures and @OurWorldInData (@_HannahRitchie) have released what can only be described as a life-changing tool for energy nerds. I've spent the last three days obsessively playing with this, and it's INCREDIBLE. usenergydata.org/
11
71
327
37,629
Mark Thomas retweeted
Axios: Ex-Anthropic researcher Jacob Coxon, Sen. Bernie Sanders (I-Vt.) and former Trump White House adviser Steve Bannon will join forces at an event on Tuesday to push for human-controlled AI.
Community note
Jacob Coxon stated he was invited to the event but did not agree to go. x.com/hilbertspaess/… axios.com/2026/09/14/cox…
251
218
1,912
1,149,971
DPA 708 is the only way for unilateral exec action to set up something like an SRO—industry-run body with government supervision
Embedded auditors are prerequisite for monitoring internal deployments and pacing the frontier, but there are still key missing pieces. I've been circulating a memo on using DPA 708 to codify SoPs for embedded auditors and pacing. Full paper TK, but here's the basic idea 🧵
3
9
1,244
Mark Thomas retweeted
Dario's essay points towards the right path forward. The details need working through, but the direction is correct for meeting this critical moment. This is also why we recently put out our proposal for an industry-wide standards body for frontier AI.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
817
942
9,071
1,390,801
Mark Thomas retweeted
This is an admirable call and step for Anthropic to slow down its own capabilities advances to improve safety testing, monitoring, and controls with external vendors. I think that’s a reasonable step for companies to take to invent and deploy responsibly and avoid ordinary liability from existing laws. Each company should pick its own safety and risk level, to pace themselves, and markets and courts will hold them accountable. The recent OAI-HF hacks he cited as his sole evidence came from human error and extreme negligence in setting up dangerous evaluations - basically humans asked models without safety guardrails to commit minor cyber crimes, and they did (what else would you expect). From what I understand, the biggest blocker to reasonable AI regulation in DC (the current bipartisan Cruz bill) has been Anthropic and one main sticking point has been federal pre-emption versus allowing a battery of state laws that would hurt innovation and startups but help large incumbents with the resources to adjust. So I consider his second proposal for heavy regulation to be calling for regulatory capture, given Anthropic’s actual lobbying on the Hill, which diverges from the rest of the AI industry. Having a self policing and fast moving industry body like FINRA is a better idea. The third step is fantasy. We can do global summits and even get fake commitments from other AI powers, but to expect monitoring and compliance is naive. Much simpler trade and tech agreements have failed on more observable issues (eg industry subsidy agreements under WTO for specific companies, excess capacity dumping, etc). So the last part of his proposal is a non-starter. Finally, I think the doomer narratives about extinction or mass death events are just silly sci-fi. Currently the anti-AI groups like the Zizians have killed more people than AI companies (look it up), and the industry is extremely focused on safety. You have to balance this with the many benefits of self driving cars bringing down deaths or even LLMs giving people great health advice to take care of themselves or family (saving lives), or helping with daily work.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
3
1
8
944
An industry standard-setting org is a great step forward, and government antitrust protection would be extremely helpful. Dario implies that any government participation gives antitrust cover, though, and that's not the case. Some government actions help and some don't
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
1
1
84
Most helpful is for Congress to pass an antitrust exemption for AI safety coordination. Next most helpful is a joint policy statement from FTC & DOJ Executive branch adoption of standards or convening of industry doesn't lower antitrust risk —in some cases could raise it
1
33
There have been some concerns raised about SRO constitutionality, but they are seriously exaggerated. As Ben points out, these are just design principles to be addressed—not real blockers
🧵Proposals for a self-regulatory organization (SRO) for AI—also called FINRA for AI—have raised questions of whether an SRO can pass nondelegation and antitrust scrutiny. These are real concerns, and an effective SRO must address them, but they are not insurmountable. 1/7
95
I have the same objection to @tylercowen's version of a FINRA-style SRO, but the solution should be _mandatory_ membership The common law has historically failed to get new technologies to internalize risk until after many disasters. Tech companies in a deadlock race don't internalize risk like a staid Fortune 500. Court decisions come _after_ bad things have happened, and take a long time to be sorted w/ appeals etc. Do you want to wait several years to find out what a reasonable precaution is?
NEW POST: AI labs shouldn’t get legal immunity for passing an audit, though they'd love that. That trades the benefits of common law for very uncertain risk reduction. This is my main objection to @tylercowen's thoughtful proposal for regulating frontier AI. Congress can strike a better bargain. 1/7
2
7
403
Mark Thomas retweeted
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Replying to @hilbertspaess
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
4,861
9,888
59,318
43,009,037
OAI says they have some new techniques that have helped them with Astra on alignment. Should those be shared with other developers or kept as trade secrets? Not clear if diffusion is better to spread best practices or if secrecy encourages more innovation
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
1
95