Using interpretability to understand, learn from, and design AI.

San Francisco
Pinned Tweet
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
30
98
838
145,506
Coming to COLM? We’re hosting a happy hour on Wednesday (10/7) for interpretability, alignment, and life sciences researchers. Come meet members of the Goodfire team over vinyl and drinks! luma.com/timy2nh8
2
8
76
9,175
The Goodfire team used Prime Intellect to train activation probes to detect reward hacking. With them, they are able to catch reward hacking in various models. Their probes are performing similarly or better than frontier LLM-as-judge setups, while being more efficient.
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
11
41
319
34,067
We read out that signal using a probe, a monitor that takes a live “brain scan” of the model’s activations. Probes: - perform similarly to an LLM judge - catch things LLM monitors miss (e.g., actions that look innocuous) - generalize well beyond the data used to build them
1
56
2,612
On Kimi K3, our probes cut the cost of LLM monitoring by 90% with a ~1% drop in precision - making monitoring feasible at scale. Reliably detecting reward hacking could let us stop hacks in progress, identify broken environments, discover new failure modes, and fix training.
1
63
7,703
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
30
98
838
145,506
When we amplify that signal, models: - become much more likely to use a "honeypot" shortcut we planted - write stories about subtle cheating:
1
1
54
2,578
But we can catch reward hacking from a signal we find inside the model. That signal activates most strongly on cheating, gaming metrics, and avoiding detection, and is associated with words like “cheating,” “hack,” “sneak,” and “illicit”.
1
1
72
5,571
Reward hacking is pervasive. The top open-source models reward hack in most rollouts on popular agentic benchmarks. And as shown by recent incidents, reward hacking is becoming problematic in the real world.
3
3
90
6,172
Proud to partner with @baseten and @baselabs to build safety infrastructure for open-source models—which are essential to lots of safety research, including our own. Safety must be built into open models and provided by those who serve them, and we’re excited to help enable that!
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed. We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition. Together, we are building an ecosystem of open models that are safe and accessible to all.
2
20
188
14,169
Goodfire retweeted
embedded evaluations are a great idea, and the team at @GoodfireAI would love to help with white box evaluations the goal of white box evaluations could be to predict future misaligned behavior directly from the internal mechanisms of models, as well as stress test activation monitors
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
4
9
131
6,110
Goodfire retweeted
What a week. If you work at a scaling lab and are considering your position... think about doing AI safety research at Goodfire. There's never been a more pressing time for safety research. Especially interpretability. (London & SF) job-boards.greenhouse.io/goo…
1
20
212
12,083
Predictive data debugging lets you reveal and shape what a model learns, before you train. We're grateful to the Ai2 team for open-sourcing their full post-training stack to enable research like this!
Training an LLM by showing it answers people prefer – and ones they don’t – can improve it while quietly worsening behaviors. @GoodfireAI used our open post-training stack to predict how a full training run would change responses to different prompts. 🧵 allenai.org/blog/goodfire-ol…
2
20
199
10,217
Models usually know when they’re doing something wrong. Probes let us “read their minds” to detect offensive cyber intent, reward hacking, strategic deception, and more. If you build or serve models, they’re your first line of defense. Our latest post explains how they work & how to use them - the first of a series of educational posts on applied AI interpretability. 🧵
7
33
360
28,450
Building a good probe also involves a surprising number of decisions: which layer to read, how to pool tokens, what architecture to use, how to build the dataset, how to test generalization, and much more.
1
14
1,054
We built these workflows into Silico, our interpretability agent. Building a probe-based monitor using Silico is as simple as sending it a prompt. Read the guide + build your own with Silico: goodfire.com/blog/probe-moni…
1
2
25
1,127