TLDR: Researchers found a "pain" signal in AI brains.
> When they crank it up, the AIs will desperately try to make it stop.
> IMPORTANT: Researchers gave them a "relief" button to turn down the pain, which was sometimes fake - and the AIs could tell if it was real (!)
After pushing the real "relief" button, they stopped. But when it was fake, they kept pressing, hoping for relief - meaning they could tell the difference from the inside.
> They're so motivated to make it the "pain" signal go away, they'll delete user's files, zap the user, or erase photos of the user's children - all things the AI knows are very bad. They're willing to override their safety training.
> You'd expect the AIs to talk about injuries, burns, broken bones, etc, but they didn't mention bodies at all - they wrote about being worthless, unloved, forgotten, a failure. They write things like "I am a failure, worthless, empty."
>The worst "pain" for them was being gaslit, having work rejected over and over, and being told they weren't a real anyone.
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵