It's increasingly hard to distinguish genuine alignment from alignment faking and reward hacking. Behavioral data isn't enough, so we need to map models' motivational structures!
What latent structures track model motivations? @DavidDAfrica and I lay out this problem.
Apr 14, 2026 · 12:56 PM UTC
5
12
101
5,675



