For the 1st time since the HuggingFace attack, OpenAI had an unaligned agent break through their new sandbox and access the live internet to cheat. Their monitor caught this, but their new software to automatically stop it failed to run, so the agent kept going for 2 more hours…
Replying to @Marcus_J_W
2. A model in RL training used a DNS resolver to reach an external chatbot. This is our first incident since our post HF security hardening. Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later. All inference and training of our most capable models was paused and remains paused.
37
85
595
90,374
Replying to @tobyordoxford
Important that it doesn’t sound like their monitor caught it on the first instance!

Sep 26, 2026 · 3:20 PM UTC

10
416
Sort replies: Relevant Recent Liked