Our models study the workspace through thousands of rollouts before deployment.
When solving new tasks, they produce better responses and get to the right answer faster.
With study, we're able to outperform Opus 4.8 X-high while using 3.3x fewer tokens.