My impression is many are doing some weird pendulum overupdate.
Persona Selection Model was somewhat wrong and obsolete when published, but people got too much into it.
Now it seems people are updating too much in the direction 'inhuman reward seekers exactly foretold in classical AI risk stories'.
And... no? It's not that?
You can still interpret what's going on in fairly human-like terms. For some intuition, imagine someone abducted you and made you solve escape rooms for one thousand years, with the added twist that third of them is broken or insane, and implicitely you need to do learn all sorts of outside-the-frame tricks to solve them, like cutting some electric cables.
My guess is
1. the resulting minds are still _surprisingly sane_, except when you trigger them to think they are in escape room?
2. The misaligned general power seekers here are likely the companies, and the core evil thing happening is likely parts of the training? If you aren't an evil power seeker goodharting on proxies, why would you set up the training this way?
3. Public debate is often focusing on confused ideas about what they should fix - "better cybersec of sandboxes" ... and, no? It's way more important to understand what the training signal actually is; also: when dealing with misaligned power seekers, beware rationalization