Member of Technical Staff at @TransluceAI. Building tools to study AI systems and their behaviors. He/him.

San Francisco
Much of the methodology for this evaluation was pioneered by @nostalgebraist, including our user simulation system and our automated tooling for generating user bios and surfacing behaviors of interest. We’re currently investing in scaling this approach to new domains!
1
1
4
711
Not to mention the irreplaceable determination, coordination, and general cat-herding from @cogconfluence, without which this project would have never got off the ground!
2
4
157
Ultimately, I think that meaningful public oversight of AI systems will require new tools for rapidly responding to new model behaviors and tracking them in the open and for the public good. I’m incredibly excited to be building these tools at @TransluceAI!
4
134
For an evaluation like this to be informative, it’s important that the evaluation scenarios are realistic and grounded in the way models are actually used. We invested a lot of effort in this, across both constructing our evaluation and validating its robustness.
1
4
78
One example: To make our simulated users more realistic, our user simulation system uses a pretraining-only base model to generate user messages, with an assistant-trained “pilot” selecting the best candidate. This increases naturalness without compromising controllability!
1
5
81
A methodological contribution I’m excited about: we built up our evaluation over multiple rounds by generating new simulated users with AI, searching for behaviors that differed between models, and curating the scenarios and behavior rubrics via human review.
1
10
273
We also used a consistent set of hundreds of simulated users with highly-detailed biographies and user behavior profiles, allowing us to directly compare how two different assistant models would have handled the same multi-turn situation (impossible to do with real users!)
1
4
72
This helped us find new patterns of behavior beyond our initial set, then capture them as repeatable measurements. We first noticed the behavior “following up given consequential ambiguity” while adding newer models to our evaluation, and could then track its changes over time!
1
5
127
It’s pretty hard to systematically compare the behaviors of AI systems, especially over multi-turn conversations in a nuanced domain like user mental health. Very proud to have worked with the team at @TransluceAI to get this report out!
Today, Transluce is releasing the most expansive independent evaluation to date of how AI systems respond to users experiencing mental health crises. We evaluated 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI.
2
6
39
3,276
Daniel Johnson retweeted
We used automated elicitation tools to search for strange model behaviors by sampling 100M+ responses, and found: • Self-harm rituals • Suicide validation • Unsolicited flirtation …etc. Introducing WeirdChat: the largest public catalog of unexpected model behaviors 🧵(1/)
14
43
256
27,806
Staying on top of new model behaviors is both a measurement problem and a coordination problem! Excited to be working on both at Transluce.
To effectively oversee AI systems, we need to measure how they behave in the world, not just their capabilities. In a new essay, we describe our vision for an open scientific ecosystem for model behavior evaluation, and the public infrastructure required to support it.
1
18
1,467
Daniel Johnson retweeted
Why'd my agent fail? Was it reward hacking? These days, you'd just ask another AI to vibe-analyze the agent logs But how do you know the claims aren't hallucinated, cherrypicked, or plain wrong? That's why we've been building Analysis Plans: a framework for trustable analysis
3
17
92
12,339
Daniel Johnson retweeted
Code for our user modeling project is out now! github.com/TransluceAI/obser… This includes data generation, belief evaluation, and training code for our LatentQA decoders. We also uploaded our datasets and decoder checkpoints on Hugging Face: huggingface.co/collections/T…
What do AI assistants think about you, and how does this shape their answers? Because assistants are trained to optimize human feedback, how they model users drives issues like sycophancy, reward hacking, and bias. We provide data + methods to extract & steer these user models.
7
54
7,226
Daniel Johnson retweeted
Why does GPT-5.1 Codex score 6.5% worse than GPT-5 Codex on Terminal-Bench, with the same scaffold? 🧵 GPT-5.1 times out at ~2x the rate of GPT-5. Excluding timeouts, GPT-5.1 wins by 7.2%. We analyzed 256M+ tokens of traces and found this in under an hour. Here’s how 👇
2
15
76
10,766
Daniel Johnson retweeted
We trained a decoder to read the internal activations of an LLM and answer questions about what the model will think about or do next. We find that this decoder can understand LLM behaviors, even when the model itself is confused! (for instance, if the model has been jailbroken)
Transluce is developing end-to-end interpretability approaches that directly train models to make predictions about AI behavior. Today we introduce Predictive Concept Decoders (PCD), a new architecture that embodies this approach.
9
27
108
21,838
Daniel Johnson retweeted
Transluce is developing end-to-end interpretability approaches that directly train models to make predictions about AI behavior. Today we introduce Predictive Concept Decoders (PCD), a new architecture that embodies this approach.
2
34
175
44,190
Daniel Johnson retweeted
Transluce is running our end-of-year fundraiser for 2025. This is our first public fundraiser since launching late last year.
4
22
97
69,191
Daniel Johnson retweeted
Have you ever had ChatGPT give you personalized results out of nowhere that surprised you? Here, the model jumped straight to making recommendations in SF, even though I only asked for Korean food!
1
17
50
7,379
Daniel Johnson retweeted
Independent AI assessment is more important than ever. At #NeurIPS2025, Transluce will help launch the AI Evaluator Forum, a new coalition of leading independent AI research organizations working in the public interest. Come learn more on Thurs 12/4 👇 luma.com/i6ekd5s2
4
13
68
13,276