GPT-5.6 Sol smashed the task time horizon eval with 270 hour average
Except it actually cheated a lot so really hit a modest 11.5 hours
Very weird!
OpenAI gave METR early access to GPT-5.6 Sol for testing including raw chain-of-thought, a railfree version of the model, and internal information about the model. With this access, METR conducted a pre-deployment evaluation of GPT-5.6 Sol, including an attempted measurement of its 50%-Time Horizon. However, the measurement depends heavily on our treatment of cheating attempts, and GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated.
Jun 26, 2026 · 7:41 PM UTC
1
600
