Working on interpretability @MIT_CSAIL CS @MIT, ex HRT intern

Chris Ge retweeted
New Paper 📄: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of explanations from a base model, the SFT-ed model explains its own current behaviors better than the base model’s behaviors, tracking behavioral drift even when we don’t explicitly train it to. We call this introspective coupling: self-explanations track a model’s own behavior as that behavior changes, and it shows promise in making introspection training a part of scalable post-training pipelines. 🧵
6
36
195
37,363
Chris Ge retweeted
New @fulcrum_inc research: Inverse Rubric Optimization (IRO), a testbed for agent science. Long-horizon tasks are often noisy, making them hard to study. In IRO, an agent learns a hidden judge's preferences under a label budget. We observe rich agent behavior and smooth scaling.
3
7
38
5,932
Chris Ge retweeted
How can we tell whether a brain region causally represents a visual concept, rather than merely correlating with it? Introducing BrainCause, a framework combining generative and brain models to create controlled stimuli and causally test neural representations. More below 🧠👇
2
10
35
3,636
FLUX.2's @bfl_ml text tokens aren't just holding your prompt. During image editing, they absorb reference image content, and some of that absorbed content, like color and style, causally drives the output appearance. New paper 🧵👇
7
37
207
30,117
Our findings suggest an efficiency opportunity: for some image editing tasks, once the text tokens have absorbed the reference content, the reference image no longer needs to participate in the rest of the computation.
1
8
1,087
For more qualitative and quantitative results, check out our paper and project page. Project: chrisg777.github.io/i2i-inte… Code: github.com/ChrisG777/i2i-int… Paper: arxiv.org/abs/2605.24624 This work was done in collaboration with @rohitgandikota, Antonio Torralba, and @TamarRottShaham
11
942
Come check out our work on better utilizing agentic coding benchmark evaluation data, presented at #ICLR2026 Agents in the Wild Workshop!
Today’s coding agent evals = single-number benchmark accuracies. But this obscures important details: which tasks in a benchmark are harder, and why? We study agent performance at the task level, and predict how new agents perform on new tasks. 📃To appear at ICLR 2026 AIWILD!
6
342
Chris Ge retweeted
🚨 We're open-sourcing Druids, a library for coordinating and deploying coding agents across machines. Our beta users have used Druids to work on open math problems, conduct ML "autoresearch," and make software faster.
4
30
232
28,476