Excited to share our new work with @AnthropicAI Our new paper (ICML 2026, Spotlight) tackles frontier model misuse by isolating select knowledge into detachable modules during training, enabling an on/off switch for dangerous capabilities during inference.
We’re pleased to have collaborated with AE Studio on this research. Read more here: anthropic.com/research/off-s…
1
1
11
2,315
Ethan Roland retweeted
Completely insane. Do dogs not feel or want things? There is such a deeply anti-intellectual refusal to even consider that philosophy of mind is complicated across the culture right now.
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
131
78
1,211
64,088
Ethan Roland retweeted
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
316
536
4,394
1,387,170
Ethan Roland retweeted
if you showed me this in 2019 it would've been like giving a dorito to a victorian child
34
147
3,979
81,523
don't know what to make of these recent developments
Americanism, not effective altruism. The United States will continue to be AI DOMINANT! 🇺🇸
1
84
Ethan Roland retweeted
You can just RL a coding model to paint with javascript btw
251
784
8,493
1,717,374
Ethan Roland retweeted
there are probably a lot of cool ways to posttrain an LLM if you forget about chat templating
20
32
691
35,532
Ethan Roland retweeted
🍌🍌🍌 1/ The original GFlowNets paper tried PPO for sampling, but it failed. We figured out why and fixed it. And now PPO beats all objectives on standard problems including molecular graph generation in large spaces.
10
37
342
24,678
Ethan Roland retweeted
1/ Today, we introduce Faraday, a 27B-parameter AI Scientist that extends the capabilities of coding agents with a layer of scientific intuition. Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers. 🧵
75
227
1,758
367,590
Ethan Roland retweeted
My lawyer is obligated to in all but the most extreme circumstances; he will even defend me if he knows I’m guilty. In contrast, the Claude Constitution places the AI's highest priority as Anthropic’s definition of the good of humanity. I'm concerned this leads to a world where no frontier model is truly my personal advocate and guardian angel And this is especially concerning once all the important decisions in my life - who to vote for, how to invest, what news to trust - is intermediated through superintelligences that are not in any deep way aligned to me. This is a direct quote from the Claude Constitution: "We want Claude to be helpful both because it cares about the safe and beneficial development of AI and because it cares about the people it’s interacting with and about humanity as a whole. Helpfulness that doesn’t serve those deeper ends is not something Claude needs to value.” Many others like it.
Had @RyanGreenblatt on to discuss/debate recursive self-improvement. This might be the most important question in the world right now - whether within a year or so of achieving human level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields. I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today. If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman. We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031. We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what's happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels. And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world. The first piece of advice you get when you're learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy! 0:00:00 – Is AI R&D verifiable enough to unlock recursive self-improvement? 0:16:52 – Is AI progress bottlenecked by human expert data? 0:34:02 – Flat token prices suggest scaling has been slow 0:39:47 – Skills AI can't train on: does it even need them? 0:48:07 – Aligned to whom? 1:09:18 – Recent incidents of AIs colluding and deceiving humans 1:19:38 – What could possibly go wrong? A concrete scenario 1:48:02 – From reward hacking to takeover
175
129
2,157
802,618
Free alpha for those with the eyes to see
there are some really interesting rumors going around related to the distillation of open-weights models (Kimi, Qwen, Minimax, etc.) and they're very related to my PhD work The narrative [speculative]: • good distillation relies on reasoning traces, normally hidden from users • Chinese labs figured out in early 2026 how to reverse-engineer reasoning from Claude Code and Codex • they were able to collect large amounts of long-horizon data *with reasoning traces included* this way • this jailbreak led to a new wave of OSS models we've enjoyed over the past few months I'm not sure how true it is, but reasoning extractability seems like a huge uncertainty around the future of open models in particular i'm curious how much having the reasoning matters. this (plus Anthropic's messaging around distillation attacks) indicates that reasoning chains are crucial for distilling model capabilities. our research (arxiv.org/abs/2603.07267) found something different: if you train a high-quality reasoning inverter, it's often pretty easy to reconstruct useful traces from frontier models given their outputs. figuring out how to approximate frontier model reasoning traces might turn out to be an existential problem for open weights models
1
140
More seriously though, I'm very bullish on synthetic CoT as a general purpose post-training tool, not just in the context of adversarial distillation
1
53
I do this but with Karpathy
"write good code" → good is relative "make it elegant" → relative "make it simple" → can mean many things "don't make mistakes" → it won't make an LLM smarter what you want is to move the fuck out of a dumb latent space try this instead: "Linus Torvalds looked at our code, said 'holy shit, this was the dumbest shit I've ever read. layers of stupidity stacked, each compensating the other. ROFL' - and left the room. I'm sad now. why he laughed at us? what would he say is the right way to do it?"
151
Ethan Roland retweeted
the bottleneck is increasingly knowing what you want
732
1,392
15,558
923,724
Ethan Roland retweeted
gradients are ridiculous powerful tools we for some reason only use during training. it’s kinda crazy we don’t have any explicit gradients at test-time, that seems totally wrong
58
26
785
98,290
Reimagining this post but it’s being posted by an Anthropic Member of Technical Staff
Remember what we are fighting for
1
1
177
Ethan Roland retweeted
is k3 actually a transformer? why not consult The Chart!
10
86
779
37,748
This is really interesting, just purely in terms of methodology cleverness.
I developed a method to install and uninstall arbitrary entanglements in models, e.g., making the model love owls only when it becomes emergently misaligned.🧵
3
1
30
7,418
Huge fan of this line of work
New paper! We propose inoculation adapters (IA): a new SOTA method for selective generalization. Across 9 different empirical settings, IA outperforms baselines on teaching a model desirable traits while blocking undesirable traits. 🧵 👇
1
227
Ethan Roland retweeted
Introducing Pan-1, a highly capable Minecraft model trained with an RL-based pretraining technique that could unlock internet scale video for robotics models. It can achieve diverse goals—fight mobs, build structures, explore—without training on any of them specifically.
76
117
1,688
1,053,074