Connor Dilgren retweeted
Is the "better writing" in the room with us? Fable 5 wrote two novels, and we had readers find out The good ✅: coherent 100K+ word plots and readers didn’t immediately stop The bad 🚩: pacing, voice, and self-editing still break (+ see who swore at Claude 167 times 🤬) 👇🧵
1
7
24
8,481
New blog post with @sarahwiegreffe on OpenAI's Monitorability Evals! We hope to make others working with these evals aware of some weaknesses we came across, and encourage more work on chain-of-thought monitorability evals.
1
9
75
13,741
It would be easier to study CoT monitorability if we had a dataset of real trajectories where an agent showed some undesirable behavior on a real task. This could be created by the frontier companies (probably too much to ask), or submitted by users as they come across them.
1
1
169
Thanks to my advisor @sarahwiegreffe and to @IvanArcus for reviewing an earlier draft of this blog post.
1
134
We released the code and models for our preprint! Code: github.com/connordilgren/are… Models: huggingface.co/collections/c…
Excited to announce my first preprint in LM interpretability! Latent reasoning models are not monitorable by default, since they don't reason in human-readable, natural language text. But can we make progress in understanding their intermediate reasoning steps using mech interp?
3
28
3,103
Excited to announce my first preprint in LM interpretability! Latent reasoning models are not monitorable by default, since they don't reason in human-readable, natural language text. But can we make progress in understanding their intermediate reasoning steps using mech interp?
7
30
209
18,400
Overall, these results are somewhat encouraging for latent reasoning model interpretability. But I suspect models with weaker natural language priors, such as those trained to do latent reasoning during pretraining or through RL, will be much less interpretable.
1
7
799
Thanks to @sarahwiegreffe for advising this project! Read the preprint here: arxiv.org/pdf/2604.04902
8
667