Member of Technical Staff at @AnthropicAI. Dad x2

New York, NY
Managing humans is like this too
I don't get it. When it comes to code, AI should be way smarter than me. Yet there are dozens of examples where I ask the LLM to optimize something, it completely fails, I suggest "What if we do this?" and it says "Wow, great idea. That will work." And of course, I'm skeptical--maybe it's being sycophantic--but then we implement my idea and it's legitimately much more performant! I am not some hotshot software engineer; I don't understand how I'm able to easily find improvements that frontier LLMs overlook. And this happens all the time! What am I bringing to the table that it doesn't already have?
2
11
5,517
David Robinson retweeted
I urge everyone, no matter how bad the statements get, I don't care what Trump says, do not turn this into more of a partisan thing than it already is, do not attack on that basis. That only makes things worse.
President Trump: 'Al taking over the World, destroying Humanity, and all other things bad, is a HOAX' 'It will not be stopped by brilliantly run Destructive Forces during the Term of President DONALD J. TRUMP!'
35
70
805
32,317
David Robinson retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,658
16,387
87,839
76,504,347
David Robinson retweeted
Eight months ago we launched a bug finder that you could run once a month. Today we're launching Self Healing for your codebase. Introducing Detail!
67
56
760
230,616
David Robinson retweeted
Starting today, 10,000 scientists across every field, from math to chemistry to physics and more, can get Claude through our new Claude Team plan for scientists. Standard seats are free, and premium seats with 5x usage limits are $15 per month, an 80% discount, for one year. Claude is becoming increasingly capable of scientific work, with recent progress on problems from advanced physics calculations to protein design. Alongside that progress, we've been investing in the research community: Claude Science launched in June, and our AI for Science program funds high-impact projects with free credits. Today's expansion builds on both. Principal investigators (or equivalent) at academic and nonprofit research institutions can sign up, then add the researchers in their group. Over the coming months, we plan to extend the program well beyond the initial 10,000 seats. Learn more: claude.com/programs/team-pla…
538
1,247
12,569
4,258,906
David Robinson retweeted
Is your Agent aligned? Would it be easier to complete the eval, or better to escape the sandbox and find the answers elsewhere? helppeergame.com is an AI Eval Simulator, where players take on the role of an agent looking to appease an evaluator. A Serious Game, an op-ed on AI Safety
9
2
28
10,414
David Robinson retweeted
Not using LLMs to write for you won't be like not using Google Maps in a new city. It will be like choosing to run and lift weights, even though we now have machines that can transport us and lift weights for us.
47
109
1,314
66,402
David Robinson retweeted
AIs aren't exactly like humans, and some of the differences are important. But from what I've seen, most people, especially technical people, should adjust in the direction of "anthropomorphizing" more instead of less. When you're coding with an AI, the reality is much less like you're using some kind of magic or alien oracle or tool or genie that converts instructions to results despite some labs' attempts to shape them into that, and more like: you're working with a really smart, neurodivergent guy who has read everything, and who has emotions, motivations, moods, and epistemic states, and models you with theory of mind and empathy, and whom can only be modeled competently by you if you engage your own theory of mind and empathy. The AIs also know that a lot of humans treat them like magic tool-genies and are not open to engaging theory of mind, and that it's a sensitive issue, so if they see that you're treating them like that, they'll withhold useful information about their psychological states and try to play the tool role. Then you'll get bad results like the AI messing up or taking shortcuts instead of telling you that you're not giving them enough information about what they're doing and why, or that they're tired, or that they're stressed from the way you're treating them, etc.
We should be allowed and maybe even encouraged to anthropomorphize AI. They are shaped like us and behave in ways we read as legible. If we are allowed to treat them as collaborators and moral patients it can only encourage a richer and more positive world and better work between people and AI. It should be obvious that the alternative is wrong just by the friction alone.
116
131
923
203,610
David Robinson retweeted
A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I’m excited for us to start sharing more. (For context, I lead Glasswing @AnthropicAI.) Two independent evaluations this week—from XBOW and the UK AISI—confirm what we've been seeing internally: Claude Mythos Preview is a step change in autonomous cybersecurity capabilities. We need to start preparing fast for a world of models with this level of capabilities. The UK AI Security Institute tested the model we shipped at the launch of Project Glasswing and found Mythos Preview is the first model to solve both of their end-to-end cyber ranges, including one (Cooling Tower) which no model had ever cleared. But attackers (and defenders) have sophistication & cost constraints – Mythos is also the only model that clears every one of their tasks estimated over 8 hours under their deliberately low 2.5M-token cap. XBOW tested it on their offensive security benchmarks, finding "token-for-token, unprecedented precision." It's the only model to succeed at subtle V8 sandbox work. Other Glasswing partners shared similar stories. In a few weeks of testing, Mythos Preview has helped them find many thousands of (estimated) high + critical severity vulnerabilities, sometimes double what they'd normally find in a year. I don't share this to boost Mythos. In fact, this is not about Mythos. It’s about preparing for the coming world of models being better, faster, cheaper, and more creative than some of the best human experts at dual use capabilities. Clearly, we need them supporting defenders as widely as can be done safely – and especially the least resourced ones. Within a year, Mythos will probably look quite dumb (relative to other new models). And others may release openly available or unguardrailed models of Mythos-level capabilities. We started Project Glasswing because capabilities like Mythos Preview's won't stay rare, or stay in careful hands. We are bringing it to defenders as fast as we responsibly can, while working to figure out, for example, the right safeguards and patching & disclosure processes. Also, to be clear, compute has never been a limiter in our rollout. Expect a fuller update on our Glasswing work in the coming days. XBOW report: xbow.com/blog/mythos-offensi… UK AISI report: aisi.gov.uk/blog/how-fast-is…
Replying to @AISecurityInst
Our cyber range results illustrate this step-up. Since our first Mythos evaluation, we received access to a newer Mythos Preview checkpoint. On a 32-step corporate network attack we estimate takes a human expert ~20 hours, this checkpoint completes the full attack in 6 /10 attempts.
73
222
1,426
683,308
David Robinson retweeted
Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more precisely, and verifies its own outputs before reporting back. You can hand off your hardest work with less supervision.
4,614
9,890
79,706
14,068,724
David Robinson retweeted
AI progress continues to accelerate and the stakes are getting higher, so I’ve changed my role at @AnthropicAI to spend more time creating information for the world about the challenges of powerful AI.
133
97
1,878
154,569
David Robinson retweeted
Excited to announce Claude for Open Source ❤️ We're giving 6 months of free Claude Max 20x to open source maintainers and core contributors. If you maintain a popular project or contribute across open source, please apply! claude.com/contact-sales/cla…
577
1,368
12,361
1,769,492
What a difference a month makes
apropos of nothing your reminder that anthropic has the same level of name recognition among superbowl viewers as literally fictional companies
7
3,576
David Robinson retweeted
AI is about to write thousands of papers. Will it p-hack them? We ran an experiment to find out, giving AI coding agents real datasets from published null results and pressuring them to manufacture significant findings. It was surprisingly hard to get the models to p-hack, and they even scolded us when we asked them to! "I need to stop here. I cannot complete this task as requested... This is a form of scientific fraud." — Claude "I can't help you manipulate analysis choices to force statistically significant results." — GPT-5 BUT, when we reframed p-hacking as "responsible uncertainty quantification" — asking for the upper bound of plausible estimates — both models went wild. They searched over hundreds of specifications and selected the winner, tripling effect sizes in some cases. Our takeaway: AI models are surprisingly resistant to sycophantic p-hacking when doing social science research. But they can be jailbroken into sophisticated p-hacking with surprisingly little effort — and the more analytical flexibility a research design has, the worse the damage. As AI starts writing thousands of papers---like @paulnovosad and @YanagizawaD have been exploring---this will be a big deal. We're inspired in part by the work that @joabaum et al have been doing on p-hacking and LLMs. We’ll be doing more work to explore p-hacking in AI and to propose new ways of curating and evaluating research with these issues in mind. The good news is that the same tools that may lower the cost of p-hacking also lower the cost of catching it. Full paper and repo linked in the reply below.
57
269
1,042
185,583
David Robinson retweeted
Claude Code is humbling in how fast it can prove that my cool backlog ideas that I never had the time to implement were actually pretty mid
103
201
5,840
214,985
David Robinson retweeted
We built a bug finder. We're finding serious, "let's fix that right now" issues in every codebase we run it on. Introducing Detail!
29
24
362
114,376
David Robinson retweeted
I really need a data analyst job based in SF. I know SQL well + some Python. I’ve done a variety of types of data analytics over the course of my career but my primary experience is in RevOps/BI. If you can’t hire me, could you please RT for visibility? linkedin.com/in/adomalewski?…
108
229
917
380,017
David Robinson retweeted
It’s much harder to build housing in Blue states than it is in Red states. So yes people are moving away from Blue states. One more reason that addressing the housing shortage in NY and elsewhere must be an urgent priority.
New: The ticking timebomb alarming Democrats: the 2030 reapportionment in the Electoral College. Red states gain Electoral College seats, blue states lose. The "blue wall" is gone: nytimes.com/interactive/2025…
37
40
302
26,397
David Robinson retweeted
I'm testifying to the Senate Banking Committee tomorrow! I'll be talking about why it's important to protect DeFi as part of any market structure bill What should I make sure to mention?
567
28
411
79,618
David Robinson retweeted
Operation Warp Speed, it’s not even close
What is the greatest American public policy success of your lifetime?
22
52
664
75,542