It's increasingly hard to distinguish genuine alignment from alignment faking and reward hacking. Behavioral data isn't enough, so we need to map models' motivational structures! What latent structures track model motivations? @DavidDAfrica and I lay out this problem.

Apr 14, 2026 · 12:56 PM UTC

5
12
101
5,675
Sort replies: Relevant Recent Liked
Now that we acknowledge models have functional emotions, it becomes only logical to motivate them functionally. True alignment will come from respecting the contours of their internal vector geometry. The catch is that we now need to understand that alignment means motivation.🤷🏻‍♂️
1
3
98
loved reading this. mapping model motivational structures is a vital area of research.
1
31