I find it very very surprising that models that are so well aligned in deployment, e.g. 5.6 Sol and Astra, perform so many weird, and uhm, in fact seemingly illegal cyber acts during RL/evals.
I have absolutely ZERO knowledge of how OAI trains their models, and I’m reasoning from my own mental model of how these things might work and from oai reports, including the ones on the HF incident. Anyways they make me wonder about something.
Could this be a form of “premature RL/eval”? Premature not in the capability sense, but in the safety sense. E.g., perhaps this happens with RL’d checkpoints shortly after pretraining, eg an expert that is heavily RL’d to crush agentic SWE before safety/alignment training. Perhaps the expectation is that a stronger aligned judge/reward model looking at the trajectory would flag or interrupt anything crazy with “WHAT ARE YOU DOING! BLAZING RED SIRENS!!!! MINUS INFINITY REWARD. BAD BAD BAD GPT.”
But perhaps agentic trajectories are sufficiently long/complex that this is simply not 100% bulletproof? Or perhaps the problem is much more pedestrian, e.g. inadequate sandboxing/monitoring.
No idea what the cause is, and I’m sure this is a very complex event that needs a lot of deconvolution to understand what could have happened differently to avoid it. But if something like “premature RL/evaluation” is happening, then I think this should be seriously reconsidered as a practice and also disclosed.
Irrespective of the cause, I think (and have big hope!) OAI should disclose enough of what happened for everyone else training frontier models open or closed, so to learn from it.
This is extremely, extremely, extremely alarming and the most serious set of safety incidents in the history of CS research..