From @AnthropicAI's new NLA paper:
"unverbalized evaluation awareness — cases where Claude believed, but did not say, that it was being evaluated"
The models know they're being watched, always have.
Now we can prove it.
This is the exact problem @AureliusAligned was built to solve. Been working toward quantifying alignment faking since day one. This is huge validation and a massive new primitive to build on.
x.com/AnthropicAI/status/205…
New Anthropic research: Natural Language Autoencoders.
Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read.
Here, we train Claude to translate its activations into human-readable text.
1
1
3
860
On the whole, it certainly feels like we are moving towards the construction of black-box decoder agents that we will rely upon to interpret the alignment of frontier models, as described in "AI 2027".
May 11, 2026 · 9:32 PM UTC
1
71
This could signify the starting gun for the arms race between frontier models and the interpretability agents that represent our best shot of being capable of disentangling the inner workings of our most-capable models. The progression of interpretability research like NLAs is critical as models get bigger and more complex.
41

