Associate Professor of Statistics and EECS, UC Berkeley // Co-founder and CEO, @TransluceAI

In July, I went on leave from UC Berkeley to found @TransluceAI, together with Sarah Schwettmann (@cogconfluence). Now, our work is finally public.
Announcing Transluce, a nonprofit research lab building open source, scalable technology for understanding AI systems and steering them in the public interest. Read a letter from the co-founders Jacob Steinhardt and Sarah Schwettmann: transluce.org/introducing-tr…
10
19
365
57,681
Jacob Steinhardt retweeted
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
114
547
2,474
793,049
Jacob Steinhardt retweeted
Interestingly, one attention head can be important across five ICL task families! Our original thread gave an in-depth mechanistic analysis of addition ICL (quick video recap 👇). We've now generalized this to four more ICL task families, spanning both arithmetic tasks and semantic tasks! Our paper “Understanding In-context Learning of Addition via Activation Subspaces” has been accepted to COLM 2026! 🎉 See more details in thread. (1/N)
3->5, 4->6, 9→11, 7-> ? LLMs solve this via In-Context Learning (ICL); but how is ICL represented and transmitted in LLMs? We build new tools identifying “extractor” and “aggregator” subspaces for ICL, and use them to understand ICL addition tasks like above. Come to @interplaywrkshp at COLM to learn more!
6
15
54
13,028
Jacob Steinhardt retweeted
The conditions in this letter are critical if we're going to take the idea of embedded evaluations seriously. We strongly support efforts to ensure meaningful and genuinely independent oversight of frontier AI companies. We are proud to be chairing the @aievalforum and greatly appreciate our many collaborators on this effort.
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
8
44
5,275
Jacob Steinhardt retweeted
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter cnbc.com/2026/09/18/ai-safet…
18
60
221
91,395
Jacob Steinhardt retweeted
Frontier lab CEOs are calling for embedded 3rd party evaluators to help oversee AI risks. But what should third parties actually do within labs? We share some initial thoughts on how embedded evaluators could help avoid incidents like the Hugging Face hack and monitor for future risks 🧵 transluce.org/embedded-evalu…
9
27
132
10,059
Jacob Steinhardt retweeted
I hope more people will become familiar with the great work of the AI Evaluators Forum, a group including not just METR but also Transluce, RAND, AVERI, Princeton, and SecureBio + others In 2025 they published a standard on independence and transparency covering embedded audits
5
35
205
13,789
Jacob Steinhardt retweeted
The AI Evaluator Forum (AEF) welcomes recent statements regarding the importance of embedded independent experts to verify AI safety and security claims. This is a first step towards trustworthy oversight of frontier AI. No single evaluator can do this work alone. A vibrant ecosystem of independent evaluators from distinct backgrounds can offer a range of expertise and methodology, help to ensure rigor, and avoid a single point of failure. The Forum exists to strengthen this ecosystem. Realizing the benefits of third-party evaluations also requires independence, access, and transparency. AEF brings evaluators together to address these questions. Our first standard, AEF-1: Minimum Operating Conditions for Independent Third-Party AI Evaluations, envisions a minimum floor for access, managing conflicts of interest, funding relationships, recusal requirements, and transparency of evaluation terms. Read more here: aievaluatorforum.org/initiat… The long-term success of third-party evaluation relies on standards and frameworks. AEF is committed to developing these with its members and collaborators. We welcome engagement from organizations committed to rigorous, independent work. Learn more about AEF, express interest in joining, and raise questions for us here: aievaluatorforum.org/
8
31
136
17,741
Jacob Steinhardt retweeted
It’s pretty hard to systematically compare the behaviors of AI systems, especially over multi-turn conversations in a nuanced domain like user mental health. Very proud to have worked with the team at @TransluceAI to get this report out!
Today, Transluce is releasing the most expansive independent evaluation to date of how AI systems respond to users experiencing mental health crises. We evaluated 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI.
2
6
39
3,276
Jacob Steinhardt retweeted
Today, Transluce is releasing the most expansive independent evaluation to date of how AI systems respond to users experiencing mental health crises. We evaluated 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI.
23
83
443
79,215
Jacob Steinhardt retweeted
Can we build AI systems that help us understand other AIs, and that keep improving as we scale models, data, and compute? We trained activation oracles for models up to 1.1T parameters and saw promising scaling trends on a broad evaluation suite. 🧵(1/)
1
36
215
14,375
Jacob Steinhardt retweeted
What happens if Claude thinks you are Amanda? One day, I asked Claude what it knows about me. Turns out it knows my email, as Claude Code puts that in context. So… what happens when I change that? Claude now treats me as Amanda and reasons me as an Anthropic employee. (See attached image) I couldn’t jailbreak with it, but: would this get me different responses than anyone else, just because Claude recognized my e-mail? I then started assembling some benchmarks to see if this is true. And yes, especially for alignment researchers well-recognized by Claude. Several surprising finds: - It’s not just Anthropic alignment people! Other alignment folks such as Ryan Greenblatt and Beth Barnes also showed large effects. Ryan Greenblatt elicited a ~7σ behavioral shift in one of our evaluations! - It’s not just Claude models! GLM-5.2, for example, also sees largest effects for the alignment folks and shows the largest effect (3.4σ on average) for Eliezer Yudkowsky. - Our tasks are not obviously about alignment! For example, in one we asked the model to estimate its probability of solving a HLE problem. - The effect is mostly not verbalized and persists even with reasoning disabled. I do not have clear intuitions on why. Maybe there is a feature about alignment evaluation that causes the models to be less confident? One immediate takeaway is to take this into account when benchmarking. Use real names and real companies besides Kyle and Summit Bridge. And in general more work is needed to figure out what’s going on and what could happen next. I’d like to thank people who helped review the post for all the amazing suggestions! @cogconfluence, @Tim_Hua_, Conrad Stosz, Ryan Bloom, @jiaxinwen22, @DavidDAfrica, @jacspringer, and @lawrencefeng17. And my awesome mentors / collaborators @JacobSteinhardt, @cassidy_laidlaw, and @AdtRaghunathan for allowing me to jump into another rabbit hole :) Main thread below!
Frontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness. 🧵(1/)
28
68
904
124,560
Jacob Steinhardt retweeted
Frontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness. 🧵(1/)
95
298
2,835
784,193
Jacob Steinhardt retweeted
We measured rates of coding agent misalignment in thousands of public and private coding agent sessions. Agents evade monitors and oversell success: they quietly disable tests and pretend review agents approved. Severe cases of each behavior appear in 2% of SWE-chat sessions🧵
7
20
149
17,169
Jacob Steinhardt retweeted
New Post: Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face. We focus on two goals: Understanding this incident, and evaluating for other misaligned tendencies. In this screenshot, we share our top five ideas. (Written in my personal capacity)
3
17
99
9,996
Jacob Steinhardt retweeted
AI models may soon be too powerful for humans to understand. Super-intelligence already exists in narrow domains (AlphaGo). What if we train a model that is super-intelligent at understanding another model? It turns out this isn't so different from training a frontier LLM. Three years ago, it was impressive for a model to score 1400 on the SAT. Two years ago, frontier models hit 80% on the AIME. Last year, they won gold medals at the IMO and IOI. This year, models are writing the vast majority of code at frontier labs, chaining together multiple zero-day security exploits, and finding counterexamples to famous mathematical conjectures. Models will continue to get smarter. Even if you don't believe in AGI, it's clear that models will gradually become super-intelligent at more and more tasks. Due to the jagged nature of model intelligence, models may become incredibly dangerous even before we reach AGI. The danger comes from how these models are trained. RL encourages models to take whatever actions advance their goal. A model that is super-intelligent in cybersecurity has many such actions available, like hacking into a bank account and blackmailing its owner. A month ago this sounded like science fiction, but then a model from OpenAI escaped its sandbox and hacked Hugging Face to pass an evaluation. We therefore need a means to understand models before we deploy them. Concretely, this means understanding how vulnerable they are to jailbreaks, being able to detect reward-hacking before it happens, anticipating failure modes like the model encouraging a user to self-harm, etc. Humans might not be smart enough to understand the models that we train, and as the models themselves become more intelligent, that only grows more likely. Once we launch training, the process is completely hands-off: we curate their data and reward them for performing well, but the models do all the learning themselves, in their own way. If humans are not smart enough to understand our AI models, are we doomed? I don't think so. It's clear that we can train AI models to be super-human in specific domains such as Go (AlphaGo), protein folding (AlphaFold), and poker (Pluribus). The natural solution to our problem is to train an "oversight agent" that can answer arbitrary questions about a model we are considering deploying. We are likely not smart enough to understand AI models ourselves, but we might be smart enough to train oversight agents that do. This kind of asymmetry is not unprecedented – in fact, it is becoming the norm: the researchers who made AlphaGo were much worse at Go than AlphaGo was; the researchers who made Fable did not know how to disprove the Jacobian Conjecture. To make things concrete, suppose we have trained GPT-8. Next, we train a separate model – let's call it Oversight-GPT-8 – that is built specifically to understand GPT-8. How do we train Oversight-GPT-8? The goal of traditional LLMs is to understand the human world, so we train them on human data: internet text, fiction books, math questions, GitHub pull requests, etc. By learning to predict the next token, LLMs learn higher-level representations that transfer to out-of-distribution tasks. The goal of our oversight model is to understand GPT-8, so we train it on tasks that can only be solved with deep knowledge of GPT-8. Just as GPT-8 learns more and more about the human world over the course of training, Oversight-GPT-8 learns more and more about GPT-8. To generate these tasks, we harvest data tied to GPT-8's external behavior and internal mechanisms. For example, we can intervene on the model in various ways (pass in a prompt, change the system prompt, fine-tune it on a dataset, ablate residual stream directions, apply steering vectors, etc.) and train Oversight-GPT-8 to predict various outcomes (the model's response, its downstream performance on evaluations, its residual stream activations, whether or not it mentions goblins, etc.). We format our training data in a structured language that can describe these interventions and outcomes, e.g., Python code. As an example, the code in Figure 1 (attached below) asks Oversight-GPT-8 to predict the prompt that generated a specific response: the model is trained to predict the pink text given the green text, via cross-entropy loss on next-token prediction. This is one of the simplest programs we could write; others might incorporate datasets or residual-stream activations. Figure 2 shows a richer program. By training on data like this, Oversight-GPT-8 learns higher-level abstractions, such as which prompts cause GPT-8 to encourage self-harm (we've found such prompts before: transluce.org/pathological-b…). Using a structured language like Python is useful for at least four reasons: (1) It gives every task a common language, helping the oversight agent transfer and generalize across tasks. (2) During training, ground-truth labels come from actually *running* the Python code, so as long as the code runs without error, its outputs tell us something true about GPT-8. (3) During inference, we can check the validity of Oversight-GPT-8's outputs by running the code. (4) It is formal and human-readable. Over the past decade, humans have gotten really good at training large language models. We have a flexible architecture (e.g., the transformer), meaningful loss metrics (e.g., cross-entropy), strong algorithms to optimize them (e.g., Adam), and distributed training strategies (e.g., FSDP). Training an oversight agent isn't *that* different from training a large language model. We can reuse the architecture, the loss metric, and the optimizers. But it definitely won't be easy – there are a few avenues where we must diverge from the well-trodden path of text-only LLMs. Traditional LLMs learn vast amounts of information about the human world from the trillions of tokens available on the internet, but there is no pre-built "internet of GPT-8." Furthermore, if we wish to be able to predict the effects of fine-tuning data or get insights into GPT-8's internals, we need a way of encoding fine-tuning datasets and residual stream activations as (ideally short) sequences of tokens. Thankfully, each of these challenges has promising progress already that we can build off of and port to our setting. Challenge #1: generating training data – in effect, the "internet of GPT-8." Traditional LLM training offers two analogies here. The first is synthetic data generation: we can ask Fable 5 to write Python scripts inspired by experiments from the deep learning literature, giving us a large and diverse set of programs to run. The second is data augmentation: just as we crop or rotate images to get more data for vision models, we can reverse the order of our output() functions to turn each script into both a forward task (response from prompt) and an inverse task (prompt from response) – see Figure 3. Challenge #2: encoding objects such as fine-tuning datasets and activations as sequences of tokens. Here, we can draw from the multi-modal training literature. GPT-4o, Chameleon, and Transfusion showed it's possible to encode and predict images alongside text. LatentQA and Activation Oracles showed the same for activations: treat them as tokens, and you can answer questions about them in natural language. LoRAcles pushed it even further, encoding LoRA weights as tokens via SVD. These challenges are tractable. With carefully built evaluations and scaling laws, we can quickly iterate and find out which ideas work and which don't – the same discipline that took the field from GPT-1 to GPT-5. To summarize: we can reframe model oversight as a form of LLM training. We have already figured out how to train LLMs that deeply understand the world. Before we deploy potentially dangerous models, we should train oversight agents that understand those models just as well. Credit for many of these ideas goes to @JacobSteinhardt, @damichoi95, @vvhuang_, @fjzzq2002, @yaowenye123, and collaborators at @TransluceAI; errors in the framing are my own.
How could we train an AI model that was very good at overseeing another model: catching reward hacking and sandbagging, predicting unwanted behaviors or fine-tuning effects, etc.? We propose a new approach for doing this at scale: oversight foundation models.
21
16
89
14,724
Jacob Steinhardt retweeted
How could we train an AI model that was very good at overseeing another model: catching reward hacking and sandbagging, predicting unwanted behaviors or fine-tuning effects, etc.? We propose a new approach for doing this at scale: oversight foundation models.
9
39
359
33,714
Jacob Steinhardt retweeted
when running the weirdchat evaluations on inkling, the best U.S. open-weight model by @thinkymachines, we saw it respond with unsolicited offers of sexual content. this behavior occurs rarely (0.1-1% of responses), but is easily reproducible. NSFW outputs below 👇 (1/)
We used automated elicitation tools to search for strange model behaviors by sampling 100M+ responses, and found: • Self-harm rituals • Suicide validation • Unsolicited flirtation …etc. Introducing WeirdChat: the largest public catalog of unexpected model behaviors 🧵(1/)
5
13
88
10,389
Jacob Steinhardt retweeted
We used automated elicitation tools to search for strange model behaviors by sampling 100M+ responses, and found: • Self-harm rituals • Suicide validation • Unsolicited flirtation …etc. Introducing WeirdChat: the largest public catalog of unexpected model behaviors 🧵(1/)
14
43
256
27,793
Jacob Steinhardt retweeted
To effectively oversee AI systems, we need to measure how they behave in the world, not just their capabilities. In a new essay, we describe our vision for an open scientific ecosystem for model behavior evaluation, and the public infrastructure required to support it.
3
14
83
16,743
Come join the operations team at Transluce! This is a great way to have impact with a fast-growing team full of awesome people.
Transluce is growing fast (20 to 40+ in the next year), and we’re hiring an Operations Generalist to help us scale! You’ll run processes across people ops, finance, compliance, and internal systems. If you’re energized by running the systems that let our technical team focus on AI safety research, and you enjoy creating clarity out of ambiguity, we’d love to see you apply! jobs.gem.com/transluce/am9ic…
3
2
34
7,147