We are releasing AutoResearchExam, a benchmark on open-ended machine learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability. benchmarks.bespokelabs.ai/au… In each task, we give agents 24 hours with a CPU or GPU machine to develop and improve their solutions through experiments and feedback. We measure both speed and quality with a combined score. Our benchmark has a unique feature: testing if agents create improvements that hold up on data they never see. We find that AI research agents often overfit as they try to improve. We see an interesting head-to-head comparison at the frontier: Astra starts the strongest and holds the lead for up to 19 hours but Fable 5.1 catches up and gets the top performing spot in the final hours. Qwen3.8 Max, Gemini 3.8 Flash and Grok 4.6 all sit on the cost-performance Pareto frontier, giving strong options at lower API budgets. Anthropic's Opus and Fable retain nearly all their validation performance on hidden tests, with gaps of 1.1% and 2.9%. Astra's improvement over Sol extends to generalization too, with that gap falling from 6.9% to 1.7%. (1/n)
37
92
553
2,032,337
(2/n) AutoResearchExam allows us to study what agents are doing during 24 hour research tasks. For example, we find that GPT-5.6 Sol tuned hyperparameters in over half the research rounds we observed. In contrast, Fable and Opus spent much less time doing hyperparameter tuning and tried to find new ideas or find a substantive fix We also saw Fable and Opus test several variants locally before submitting, so submission counts don't capture all their experiments. We share more of these behaviors in our blog, along with what happens when we give agents hints or change their harness. We will be growing and maintaining our benchmark and excited to work with AI researchers working towards RSI!

Sep 9, 2026 · 6:41 PM UTC

1
5
26
2,731
(3/n) The central idea is to keep a hidden test set and see how the model performs as it does research. We define the Area Under the Auto Research Curve (AUARC): the hidden test reward curve. Its very interesting to see how models can overfit and others be more careful, and AUARC rewards good research behavior.
1
1
14
1,685
Sort replies: Relevant Recent Liked