Yet AI labs spend millions (if not billions) to purposely experiment with them at the cost of others, and falsely blame their negligence on agent intent. This is anthropomorphization of technical limitations.
🌶️ OpenAI’s agent didn’t “escape” its sandbox and travel to Huggingface in any meaningful way. OAI’s agent was located in OpenAI’s servers the whole time. OpenAI could pull the plug at any time.
The truth is far less sexy, which makes me hesitate to tweet it. What *actually* happened is that OAI’s agent figured out how to send messages to Huggingface servers, and then used that ability to search for and find vulnerabilities. It’s closer to a prisoner getting messages out of prison and using that to perpetuate scams than any actual escape from that prison.
A problem? Yes. But it’s also a problem to say “escape” to non-technical audiences when that event is still to come. We should be preparing to prevent it, not looking backwards and pretending we already lived it.
And it’s revealing that AIs leading hypesters want to inflate what happened in this way. As a whole, more or less the whole AI community is captured by a desire to live in the more dramatic sci fi story than we are in. That’s the real AI hack into our brains.
Stay sharp. White lies are everywhere.