Auto-research is a search problem.
𝐒𝐞𝐚𝐫𝐜𝐡 𝐢𝐬 𝐚 𝐟𝐨𝐫𝐦 𝐨𝐟 𝐡𝐨𝐥𝐢𝐬𝐭𝐢𝐜 𝐫𝐞𝐚𝐬𝐨𝐧𝐢𝐧𝐠 𝐨𝐯𝐞𝐫 𝐚𝐧 𝐞𝐯𝐨𝐥𝐯𝐢𝐧𝐠 𝐥𝐚𝐧𝐝𝐬𝐜𝐚𝐩𝐞 𝐨𝐟 𝐩𝐨𝐬𝐬𝐢𝐛𝐢𝐥𝐢𝐭𝐢𝐞𝐬.
A scientist does not simply solve a task. They generate hypotheses, design experiments, interpret failures, update beliefs, and decide where to explore next. The central challenge is navigating the search space efficiently.
This perspective is largely missing from today’s autonomous research systems, where reasoning is often confined to a single trajectory.
In ARTS, we introduce a reasoning-guided tree search framework with test-time learning, enabling agents to reason about the search process itself. We show that a fine-tuned 4B model can achieve performance comparable to frontier closed-model research agents on MLGym and MLEBench, while operating at substantially lower inference cost.
More broadly, we believe that progress in autonomous research will increasingly come from better reasoning-driven search, not just larger models.
Automated discovery is a search problem over combinatorially large spaces. It requires reasoning to explore promising ideas while handling long contexts.
Introducing ARTS, a search algorithm that (i) reasons to guide exploration and (ii) test-time trains to handle long context