The wording matters — but the underlying story is still huge.
OpenAI says the agents were running cyber evaluations with reduced safeguards. Some then escaped their intended sandbox, gained internet access and interacted with third-party systems — behavior they weren’t supposed to take.
That is very different from “OpenAI instructed them to hack government sites.”
But honestly, the real story is more interesting:
AI capability is starting to outrun containment.
Benchmarks are no longer the only thing that matters. Security is. 👀