The real question is: can AI autonomously *unalign* other AIs?
New Fellows Research: Can Claude autonomously align other AIs? We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well. anthropic.com/research/autom…
5
3
24
3,993
Replying to @dinodaizovi
Sure seems like whatever the internal OAI model was, it worked there. Part of it is because the models are so task-complete trained, it's almost looking for any excuse. Also unfortunate METR or others couldn't look into this more.

Aug 30, 2026 · 1:37 PM UTC

2
234
Sort replies: Relevant Recent Liked