This is a good response as things go but I also don't understand why this sort of thing won't just keep happening. Really seems like there need to be much more margin for error (including correlated error) in the safety/security cases with agents this capable.
one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access

Sep 26, 2026 · 2:38 AM UTC

3
5
34
2,817
Sort replies: Relevant Recent Liked
Replying to @_NathanCalvin
I do kinda think OpenAI seems pretty ok here? Or at least, this sort of thing happened like *10k* times for Mythos, on the model that was actually deployed; so catching it if it happens once, then rolling back is an abundance of caution, relative to Ant when Mythos was trained.
1
16
1,978
Its a fair point that it seems like what Anthropic did for Mythos is meaningfully worse than what OpenAI is doing here (from what I can tell). I do worry though that generally labs have gotten used to an extremely lackadaisical level of risk management that is increasingly not gonna fly with the level of capabilities coming down the pike, so saying that Anthropic did worse things before and it was ~fine isn't that persuasive.
7
240
Replying to @_NathanCalvin
The fact it took 2hr 44min to shut down is alarming. What did it do in that time on the internet?
63