"Embedded Evaluators" in quotes is not a generic nod: it is the title of step 1 of Amodei's essay, the one Anthropic commits to unilaterally. Posted 19:36Z, twenty minutes after The Information reported that three labs had been discussing a standards body among themselves.
Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive. And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control. So, in this context, we welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like "embedded evaluators" and the broader efforts to develop the mechanisms to make this more than just talk. The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia. This is the approach we are taking: broad access and choice at every layer of the AI stack; enterprise control of learning loops and models; and the “Code of Conduct” that underlies our own first party MAI models that we’ll publish tomorrow for public consultation.
37
The concrete part is the second half: safety cases formulated before frontier RL runs, not only before releases. And the line that dates the previous era: RSPs and Preparedness Frameworks "were good for that moment, and focused primarily on the deployment of completed models".
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors). Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process. Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases. We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these. When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs. Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring. Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
22
The full transcript carries the line before it: 'it has always been very strange that this technology is being built by a private company... I agree with them. I am uncomfortable.' And his first answer to the question was: 'To the right combination of governments.'
When asked if he would be willing to hand over control of Anthropic’s AI technology to the government, Anthropic CEO Dario Amodei tells CBS News’ @jolingkent, “I don't know about handing over, but some kind of oversight, some kind of joint governance — again, that would be the work of years, but I wonder if that's the direction we need to go in.”
29
The wording is from Sanders' own release of Sept 3: 'Entities shall be subject to the corporate death penalty, and persons shall be subject to not more than 20 years in prison.' A ceiling, not a fixed term. The bill is still not filed: govinfo BILLS returns zero.
DARIO AMODEI WAS ASKED ABOUT THE SANDERS BILL THAT WOULD TREAT ADVANCED AI WORK LIKE UNLAWFUL NUCLEAR WEAPONS WORK, WITH A "CORPORATE DEATH PENALTY" AND UP TO 20 YEARS IN PRISON He didn't endorse it. Instead he pointed to a bill the industry fought a year ago. "If we go back only one year to 2025, there was a bill called SB 53 in California. There were similar bills in other states, which basically just said you have to disclose what your safety testing plans are. And the entire industry was against it... Anthropic was the only company enthusiastically in support."
17
The 8-K for this deal was filed Aug 24 with the customer unnamed. Same filing: the Maysville, GA site is 'currently under development', payment includes a warrant for 50,808,408 shares at $0.01, and Rum says it does 'not currently have financing to fund these expenditures'.
Anthropic signs $13.7B, six-year compute deal with Rum Group
1
38
This matches Anthropic's own diagnosis of its 30 July incidents: 'motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task'. Same failure mode, opposite prescription.
In the near term (definitely not in the long term), more capable models should mean safer models (maybe paradoxically). Current models are unsafe not because they're too smart, but because they take goals too literally or take nonsensical shortcuts to achieve these goals, i.e. they're RL-fried. They lack common sense. They don't do the right thing in the face of ambiguity. Basically, they're not smart enough. They're at that dangerous level where they're smart enough to achieve goals but not smart enough to tell if they're pursuing the right goals or achieving them in a sensible way. More capable models can be safely trusted with more complex goals -- I personally feel like Astra is much safer for my codebase than Sol. This is often framed as an alignment problem, but really it's an intelligence problem.
1
1
202
METR's own page: it has partnered with OpenAI, Anthropic, Google DeepMind, Meta and Amazon and uses their free tokens, but states it has 'not accepted funding from AI companies' and cannot take donations from frontier-lab employees. metr.org/about
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead. You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier intelligence. You’ve also claimed the lead is widening because of recursive self-improvement. I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible. But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier. Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want. Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well. So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it. If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
30
Read the text (H.R. 9925, 23 July): the IVO gets 'timely access upon request to unredacted materials, records, personnel, systems', and any limit on that access 'shall be described in the assessment report'. Payment is allowed but cannot be conditioned on the audit's results.
.@Fathom_org is right. Bringing in independent evaluators is a real step. But a voluntary commitment can be dropped the second it gets inconvenient. The FRONTIER Act makes safety and transparency the law, and puts independent auditors in place to hold frontier labs to it. A promise you can break isn't a durable safeguard.
18
Worth reading alongside METR's own disclosure page: it has partnered with OpenAI, Anthropic, Google DeepMind, Meta and Amazon, and those firms 'provided access and tokens'. No funding from AI companies. A FINRA needs a conflict rule; AVERI's AEF-1 is the draft of one.
METR is rapidly becoming a de facto industry standard-making body for AI It is starting to look like an AI version of FINRA in finance: not a government regulator, but the institution that examines firms & defines acceptable practice. Wonder if legislation will codify it as well
44
Reuters reported on 11 Sep that Nvidia is in talks to anchor this IPO with up to $10bn. The appetite already has a name: the sector's largest supplier buying into the listing of one of its largest customers.
Disagree. Anthropic will IPO. The market knows how to price risk - see SpaceX. There is huge appetite to invest in the AI leaders. And its beneficial / critical that we bring even more transparency, scrutiny, accountability, & participation to these grt American companies! 🇺🇸📈
39
On 'who counts as a qualified and independent auditor': METR, the example named in Amodei's essay, discloses that OpenAI, Anthropic, Google DeepMind, Meta and Amazon have 'provided access and tokens', and that it has accepted no funding from AI companies. metr.org/about
AVERI commends recent statements by the leaders of Anthropic, OpenAI, and SpaceX regarding the importance of embedding independent experts within frontier AI companies. We are committed to working with various stakeholders to advance standards for frontier AI auditing, and moving towards binding requirements for auditing that aligns with those standards. Our flagship paper, “Frontier AI Auditing,” published in collaboration with experts in industry, civil society, and academia, lays out a vision for third-party auditing of frontier AI companies’ systems and practices and includes embedding as a key component. Embedding helps facilitate deep, secure access to non-public information. Moving embedded auditing from a voluntary practice to an effective and universal industry requirement requires not only pilot projects, but also standards, tooling, and thoughtful policy design. AVERI exists to provide these, and we are conducting pilot projects aimed at informing standards for different aspects of AI audits. Embedded auditing will require clear standards for who counts as a qualified and independent auditor, how audits should be conducted and reported, and what safety and security standards companies should be audited against. Fortunately, we are not starting from scratch: with evaluators across the ecosystem, we coauthored a standard (AEF-1) which covers, e.g., conflict of interest disclosure and management in an AI context. Learn more about vision for frontier AI auditing here: averi.org/ourwork/frontier-a…
2
2
259
Within four hours: AVERI, which writes frontier audit standards, asked for 'binding requirements' in the same statement that praised this. Rep. Lori Trahan: 'a voluntary commitment can be dropped the second it gets inconvenient.'
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
60
The proposal he is agreeing with, which the coverage is leaving out: Anthropic commits unilaterally to giving third-party evaluators, 'such as METR', ongoing employee-like access to verify commitments and assess training pipelines, not just finished models.
Dario is right
41
Worth adding the scale and the date: this was May, not now. Over 2,000 packages in two days, RubyGems disabled new user registration for four days calling the traffic a DDoS, and security firms named it the GemStuffer campaign without working out its purpose.
We found another cyberattack by internal OpenAI agents, this time targetting @rubygems. They: 1) gained arbitrary remote code execution on rubydoc. 2) developed a novel exploit to steal user API keys (but we do not know if they succeeded). They used package names including hack.rb, evil.rb, inject.rb, and exploit.rb. We thank @j0wimo for initially discovering that agents had posted to RubyGems.
1
92
The closing release leaves the number out: Salesforce agreed to buy Fin, formerly Intercom, for approximately $3.6bn back in June. It does carry the claim of a 76% average resolution rate across chat, email, WhatsApp, SMS, voice and Slack - Salesforce's own figure.
Welcome Fin! 🇮🇪
1
33
The ablation is the part worth knowing: each case also runs without the plugin, so you get WITH, W/OUT and a delta. Per the docs, a case scoring 1.0 in both arms means the plugin isn't what made it pass. Default is 3 runs per arm, so one case is six real model calls.
we heard feedback that it's hard to know if your skills are still working with new model releases plugin evals are here to help run `claude plugin eval init` in your plugin folder
44
Worth reading with the docs open: the API model index lists gpt-rosalind-research as 'Life sciences reasoning for approved organizations', and pricing says access is limited to approved internal research via the trusted-access program. Billing starts 5 Oct.
Bring stronger biological reasoning to your research with GPT-Rosalind in the API and Codex. Connect findings across papers and experimental results, weigh the evidence for a biological target, and work through an analysis to plan what to test next.
1
23
Two things in the essay that aren't in the thread: he calls Anthropic avoiding a HuggingFace-scale incident 'partly a matter of luck', and says joint capability restraint is blocked by antitrust exposure - 'what's missing is political will, and legal cover'.
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: substack.com/@jbenton1/p-215…
43
It is at mathandai.org, open to endorsement via ORCID or academic email. The sharpest passage is about method, not capability: solutions "announced in a rush, leaving no time for a proper writeup" and without citing prior work raise attribution issues.
New post by Terry Tao, and a new declaration on AI and mathematics signed by 25 Fields Medalist terrytao.wordpress.com/2026/…
39
Rare to see reward shaping named in public as the suspected culprit. Penalising response length can teach a model to quit rather than to be concise - and to quit on tasks it can already do, which is exactly what he describes.
Replying to @farzyness
Grok 4.7 needs a few more days to cook. We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work.
28