System cards, risk reports, and misc safety takes at Anthropic; math; puzzles; spaced repetition. Writes with too many caveats for Twitter.

Berkeley, CA
I'm confused about the level of badness of the Australian government "hack". From chatting with Claude about public reporting so far, I can't tell if anything more interesting than "got around a basic anti-scraping measure and read some publicly accessible URLs" happened?
13
2
120
8,521
Quite possibly there was much spookier stuff that happened, and I just haven't seen it, but it currently seems to me like discourse lumping this in with HuggingFace is pretty misleading. Very happy for corrections if I missed more details on severity here.
1
33
1,228
Update: I now think that [something we know not much about] happened, distinct from the "basically just reading a public webpage" stuff Transluce found, which may or may not have been spooky.
Replying to @MaskedTorah
AFAIK we just don’t know that much about the incident. I tried repeating the Transluce methodology on the other three cases but found nothing on any of the url scanner sites (and nothing in DSE wiki).
23
977
It is good for AI companies to have oversight from deeply cracked teams of experts with world-leading experience in investigating frontier alignment problems. It is also good for AI companies to have oversight from unimpeachably boring and independent sources that even a bad-faith psyop would struggle to find complaints with. As of September 2026, you can't max out both of these axes in one source yet, but nothing stops you from having multiple sources of independent evaluation! I think Accenture and METR are both on the pareto frontier of options. Obviously any frontier AI company which doesn't also deeply involve METR/Apollo-shaped orgs in evaluating their practices would be failing to do an adequate job of holding themselves accountable to the public (at least until such time as there exist more ordinary consultants with a deep bench of AI auditing talent, which I am hopeful will start happening soon). And I do mean it about the pareto frontier thing, I think a lot of people are underrating Accenture? Like, go scroll through the twitter account of Accenture CTO and Faculty CEO @MarcWarner10, these are not people who haven't heard of AI safety!
We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in this area over the next five years. anthropic.com/news/accenture…
Community note
Anthropic presents this as an "independent evaluation" but will directly fund Accenture's work and has a prior commercial partnership deploying its models, including training ~30,000 Accenture professionals on Claude. anthropic.com/news/accenture… anthropic.com/news/anthropic…
6
5
96
9,491
Upon reflection I think this tweet is a bit too one sided and want to add two caveats: (1) Providing funding obviously introduces a COI here that there is not with eg METR. (2) In some ways more sources of oversight are additive - additional shots on goal for noticing and being transparent about problems are great - but there’s a cost where labs can emphasize the most friendly reviews or use them as a defense against more critical external review. I have some worry that even with a cluster of very good talent doing the work, it’ll be harder to say very blunt/weird critical things like “this company is taking on lots of existential risk right now and should immediately stop” from within a large “normal” company.
1
13
567
A ranking of takes on embedded third party evaluation this past week, from worst to best: [contentless sneer against outgroup] METR bad because [huge Sankey diagram]. fact checking? timelines of relative investments that make causal sense? determining which numbers are large or small fractions of other numbers? sounds like some EA bullshit to me. just look at this diagram with all those curved lines, you can SEE the nest of snakes. METR bad because [an actual specific chain of actors who have some pairwise relationship to each other that ends with someone it would be bad for them to have strong COIs with, e.g. "METR once received funding from an org who once received funding from a person who once gave funding to a now-frontier AI company"]. It is beneath my dignity to explain how this actually affects the decisionmaking of METR; you should vaguely perform a mood affiliation and keep scrolling. Don't think too hard. guys guys guys you HAVE to embed my company/organization into the labs. I have been an unwavering supporter of external embedded auditors since 7 minutes ago when I saw this essay taking off. Look at this thing we published once that sort of looks like AI safety if you squint! Fear not, I can assure you that I have never done anything altruistic in my life and if I had I would have been ineffective at it. I observe that [politicized actor] has had [bad take]. Let me use this correct observation to score points for my side and politicize the situation further. === zero point: takes beyond this line are better than logging off === Dear [person with insane take], here is an earnest explanation of why you are wrong. METR is bad/problematic because [an actually reasonable concern, like greater cultural overlap with Anthropic than OAI or the pressure for individual employees to be on good enough terms with labs that they could later get hired], with no further suggestions for what to do about this issue. Tweets which simultaneously acknowledge that (1) more social independence from labs would be good and (2) almost everyone competent is socially connected to labs, even if they don't propose solutions. [literally any take that engages with object level assessment of AI companies done by a third party org like METR, SecureBio, Guidelight, Redwood, Nightingale, etc] === current discourse frontier: takes beyond this line are better than anyone has yet posted === I have [actually reasonable concern] with METR/Redwood/etc. To address this concern, we should do [concrete proposal that makes any sense and would solve the problem]. I am founding a new third party evaluation org / pivoting my existing org to do more third party evaluation. We're doing [thing which is both (1) not identical to METR (2) remotely useful for assessing AI companies]. Our initial work will be on [concrete specific thing that would help]. METR has problem X. We should instead use preexisting organization Y, which does not have problem X and has comparable expertise at assessing loss-of-control risks and misalignment incidents inside AI companies, for instance [past work they've done of similar quality]. (this would be an amazing take but it is impossible to post because no such org exists)
8
5
106
8,599
If you are good at investigating the behavior of powerful AI models, today is a great day to apply to METR: metr.work
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
3
9
174
35,778
I think there's a ton of informative stuff in this post! Highly recommend reading.
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. anthropic.com/research/align…
2
2
65
94,297
Based! I generally agree with this thread. I personally think Ant capability research is net good, but only in the hope of letting A\ spend down a lead on measures that give humanity more time to try and make it out of this alive, and I very much respect the choice to abstain.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
14
2
138
50,389
Things are moving way too fast, we don't have anywhere near the degree of assurance we'll want for ASI, and if we survive an unmitigated race at the current pace it will be because we got lucky at how hard the problems were rather than because the industry behaved responsibly.
12
6
81
15,492
Anthropic's second Risk Report is out! I'm pretty happy with a lot of the new things we landed in this one and think it's a big improvement over our first one in a bunch of ways. Among the novel features: * Coverage of internal models * A much more structured argument around alignment risk * Claude-powered review of the alignment section * Discussion of a wide variety of incidents, safety process failures, and ways Anthropic has fallen short (see 4.8.2, 6.5, and especially 5.2) * An attempt to get somewhat quantitative about threat modeling for CB-2 risk despite enormous uncertainty * Much more detailed discussion of jailbreak risks * Two different cases of increasing assessed risk levels from "very low" to "low" on the basis of general increased uncertainty given worries about the overall degree of process assurance * Early notes on some cool new model organism work that informs the alignment assessment and more! Excited to hear thoughts on how the next one could be better.
As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: anthropic.com/aug-2026-risk-…
3
5
87
28,681
+1. I think working at METR would have been better than the frontier AI company work I was doing a year ago. The work I'm doing now is a sufficiently good fit for me in particular that I thiiiink it looks higher impact, but I think it's plausible I'm still making a mistake!
Replying to @idavidrein
To my friends at frontier AI companies: please consider joining us. Things are getting real, and we need folks who can help lead the ambitious risk assessment programs we're developing.
3
3
139
22,626
Very happy to have signed the pacing the frontier letter; its existence gives me a lot of hope for humanity's survival! I endorse the letter as written (and think, given the constraints, it's probably close to the best it could be for this level of consensus), but some places where I differ from its connotational tone: (1) Not only is AI "not guaranteed" to make a dramatically better future, the odds of failure are terrifyingly high: I think* there's something like a 40% chance we get an outcome around as bad as human extinction or worse, and another 30% chance we get a future that, while containing some good things, falls radically short of what a wiser civilization could have obtained (say, <5% of the value of a truly great future). (2) I don't really endorse the vibes of "to realize AI's potential"; I think a much more immediate and pressing motivation is "to avoid catastrophically bad outcomes from AI that will kill a lot of people". (The potential is also very important, ofc, but I think most moral theories would view it as being of secondary importance when risks are this high and there's little that would do more than temporarily delay that potential anyway.) (3) "address emerging risks, develop security measures, and strengthen oversight" is fine so far as it goes but a little vague. Concrete things I'd like to do with slower AI development: way better interpretability, build up a robust and well-resourced third party ecosystem for independent auditing of AI companies and get lots of reps in for their oversight, put tons of effort into the automation of alignment research, work on governance mechanisms for ASI, develop a vastly better science of the nature and development of AI behavior, build much more powerful control mechanisms, point lots of powerful AI labor at ambitious scalable alignment projects (eg work like ARC's), deep dives (including external audits) of individual AI behavior incidents, etc. Also getting civilizational biosecurity preparedness in order. *epistemic status very approximate vibes, my numbers will change day to day and depending on the exact operationalization. pacingthefrontier.com
8
12
127
24,121
I will note that I don't think slowdown measures are *obviously* good - the current set of frontier actors are much more worried about the risks that matter than they could be (and are few in number, and located under the jurisdiction of a single nation), which I'm very grateful for. A poorly-executed slowdown could leave us at the same capability level but with a wider pool of more cavalier frontier actors, which could be better for centralization-of-power reasons but I think looks a lot worse on the loss-of-control front (and I currently think the latter is more urgent to address). And to the extent there's a limited budget for total slowdown (either because of political will, or because of the background rate of x-risk until we exit the current vulnerable period of human civilization), it's better all else equal to spend that budget in a regime with more productivity uplift from AI labor. But the ask of the letter, to start setting up the mechanisms now so we have the OPTION of doing this later if and when it seems prudent, seems more clearly good to me, and I think it's fairly unlikely (15%?) that I'll later think the object-level request made in the letter was an unwise one.
2
18
3,775
RSP v3.4 is out! Not a huge change, but slightly more diffs than some previous incremental updates, and I helped out with workshopping a bunch of the changes here and feel pretty good about them. Happy to offer takes where people have questions. anthropic.com/responsible-sc…
4
29
3,804
cc @TheZvi for possible interest since iirc you were sad about missing the 3.1-3.3 updates. See the redline copy for exact character-by-character* diffs: cdn.sanity.io/files/4zrzovbb… *excluding the page numbering fix where there are no longer two pages labeled page 10, oops
1
12
1,907
"Passing the ITT of someone with FDT intuitions" is a surprisingly difficult task - before seeing the state of discourse, I would have expected *much* less disagreement and way smaller inferential gaps on decision theory among smart people who think about the topic.
Replying to @Benthamsbulldog
It's pretty interesting and surprising to me how much people vary in their empathy for something like the core FDT intuition! Like to me it is just extremely obvious that you should cut off your leg in the leg-cutting situation* and this scenario gives me no pause at all, even though apparently it feels like a decisive argument against FDT to you. Am I right in understanding your position as being such that you _would_ pay to have someone do neurosurgery on you to make you into the kind of person who'd cut off their leg in this situation, you just think this mental state you wish you were in is an irrational one? *up to caveats around whether the predictor is only doing the 30 year sleep because of the existence of people who will cut off their legs, and if this is more like giving into blackmail, and all the ways that nothing is ever obvious, but for the first-order account presented here it's very very intuitive to me.
2
15
6,324
"Identifying AI writing" is another cognitive task where I'm shocked at how much skill variation there is after controlling for intelligence. Examples that are blindingly obvious slop to me will go unremarked-on by very smart people I know who've had similar total AI exposure.
1
1
14
685
(From within my typical mind fallacy it's baffling that people ever post undisclosed AI slop as their own work, because it feels impossible to hide, but I suspect variation in detection ability drives a lot of this - to many people it does just read like ordinary human prose.)
5
394