The real question is: can AI autonomously *unalign* other AIs?
New Fellows Research: Can Claude autonomously align other AIs?
We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well.
anthropic.com/research/autom…
5
3
24
3,993
Sure seems like whatever the internal OAI model was, it worked there. Part of it is because the models are so task-complete trained, it's almost looking for any excuse. Also unfortunate METR or others couldn't look into this more.
Aug 30, 2026 · 1:37 PM UTC
2
234


