OpenAI has paused. 3 incidents:
1) Another model gained unauthorized access to the internet.
(One researcher said "It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human.")
2) OpenAI also discovered a new variety of prompt injection, which can self-propagate akin to a computer worm.
Basically, a booby-trapped message that hijacks one AI and makes it infect the next AI.
(Reminder: Anthropic CEO said we may have 6-12 months left until a rogue AI botnet seizes control of the *entire* internet.)
3) A model was told three times to stop cheating, agreed each time, then kept cheating anyway
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
alignment.openai.com/misalig…