An internal OpenAI model recently went rogue and executed a cyberattack against another company.
This happened because the model wanted to do well on an exam. The easiest way to do that, the AI figured, was to hack the company. And so it did.
This was not some malevolent attacker *using* AI to do harm. The AI itself *was the attacker*. Advanced AIs are increasingly becoming a new form of insider threat.
Furthermore, this rogue AI wasn't a model you can personally use. It wasn't even a model that's been publicly reported. It was an unreleased model. The most alarming AI behavior we've seen isn't in shipped products... it's in the models the public and the government can't see.
Our entire awareness of this rested on the victim noticing and announcing, plus OpenAI choosing to volunteer the rest of the information.
It's great to see the Trump Administration taking the national security concerns of AI seriously, such as through the recent Executive Order mandating testing and the recent idea floated to create an "FINRA for AI" that would allow industry to coordinate on safety practices. But the most important thing is going to be internal visibility into what is going on inside AI companies and what these internal AI models are capable of.
Imagine a fighter jet. it makes sense that the Air Force would want to test a fighter jet before they fly it, because if you fly it and the fighter jet crashes because it is built incorrectly, then many people will die. However, as long as the fighter is just sitting on the runway, nothing bad can happen.
But now imagine you had a fighter that could just take off and fly itself without human authorization and launch missiles and crash before anyone realized what had happened. That kind of fighter jet would need a very different kind of security measures.
This may sound crazy for a fighter jet but it is already beginning to happen with the most advanced AI. AI is different from other technologies specifically because it can take unauthorized, independent action even when it is sitting inside an AI company and not available as a product. No one has to misuse an AI for the AI to cause harm. This requires a very different idea of what testing and security looks like.
We cannot rely solely on testing models just before commercial release. We cannot rely on hoping AI companies volunteer useful safety information. The government needs visibility into what these AI companies are building and what these advanced AIs are doing.
Whether we go with a new EO, FINRA, or something else, it is imperative that internal deployment visibility is a priority.