Today, a certain era in mathematical benchmarks is coming to an end. We designed this set of tasks in the era of the o4-mini model and initially expected the pace of solving them to be rather slow. Things started getting serious in January, and I made a prediction back then that the benchmark would saturate within nine months. Even despite the invalidation of some incorrectly formulated tasks, that prediction turned out to be pretty accurate.
The part that gives me chills is that I still can't solve most of these problems myself, and probably never will. And, as FirstProof showed too, the chance of finding a problem for which we know the answer, yet which none of the vanilla or harnessed AI models can solve, is basically close to zero.
I think we need to completely change our understanding of what is actually hard in mathematics now. Maybe, after all, the only things left in the universe are black holes and Busy Beavers…
Every FrontierMath Tier 4 problem has now been solved by AI, with GPT-6 Astra solving the last problem standing. Mathematicians often commented that AI found unintended shortcuts when solving their Tier 4 problems. Not so for this last one, which was created by Jay Pantone.