The Benchmark for Recursive AI Improvement - Measuring the rate of self-improvement of AI agents

Canada
GPT 6-Sol xhigh scores 21% on ARI Bench. I was expecting MUCH better as GPT-6 Astra (38%) is the leader and Opus 5.5 just scores 33%. What happened with GPT-6 sol :(
1
134
Opus 5.5 xhigh scores 25% on ARI Bench, the benchmark for recursive AI improvement. Only a 2% improvement over previous iterations... this is a bit disappointing. Either they purposely hindered RSI or the model is optimized for benchmarks. We're investigating.
69
Grok 4.7 scores 16.3% on ARI Bench. What a regression! (We ran it 3 times! Grok 4.6 scored significantly higher. It appears as if Grok 4.7 was rushed and fails.)
26
Perhaps I made it seem as if Mimo V 2.6 was the best before... it's not. It "only" scores 18% while the frontier is at 38%, however it's low cost makes up for it.
89
New "Value King" Mimo 2.6-Pro score 18% in ARI Bench. This makes it the best model for the price at the moment, outperforming both GLM 5.3 flash and GPT 5.6 Luna. If you need something smart to process a lot of data... this is it.
138
Union Alpha scores 32% on ARI Bench, the benchmark for recursive AI improvement. This is the second highest score ever, only trailing GPT-6-Astra. Whoever it is... they have quite a model on their hands.
1
155
Deepseek V4.1 Flash's performance has been inconsistent (perhaps due to providers), however we believe the real performance number is 24% which is quite decent.
23
WOW! GPT-6 Astra scores 38% on ARI Bench, the benchmark for recursive AI improvement. This is simply mind blowing, a monumental leap over the previous generation. This indicates that OpenAI's NEXT model will be incredibly powerful as the recursive cycle has begun.
1
652
Muse Spark 1.3 scores 13% on ARI Bench, the benchmark for recursive AI improvement. Quite a difference compared to what we're seeing from other benchmarks like the Artificial Intelligence Index 👀
62
Seems like the Artificial Intelligence Index needs to update it's methodology. They are still using old saturated benchmarks which lead to issues like this. Can't trust it until they update it. For example, imagine we were still using ARC-AGI 1 instead of ARC-AGI 3...
1
56
Holy! I just made the same post as you 1 hour ago about the status of benchmaxxing... and I noticed the same exact thing.
We need to talk about benchmaxing. Gemini 3.8 Flash and Muse Spark 1.3 both dropped today. Both look great on benchmarks. Muse Spark 1.3 is now tied with Fable 5 on the Intelligence Index. Gemini 3.8 Flash sits one point behind Kimi K3. I want to be excited. I am not. Every release this month has the same shape. Score jumps. Real world experience does not. It is discouraging. I trust the benchmarks less every week, and I trust the labs gaming them even less.
21
After benchmarking 33 models... Here's what I noticed from SOME companies: 1. The first release of a new pre-train is a significant advancement. 2. The next 2-3 iterations are genuine reasoning improvements 3. After, they benchmaxxing. I dislike the 3rd part.
9
Gemini 3.8 Flash scores 14% on ARI Bench, the benchmark for recursive AI improvement. Seems like Gemini 3.8 Flash is just more specific tuning for certain tasks / benchmarks. It's actually a regression compared to Gemini 3.7 Flash. That said, it completed the task 30% faster.
124
Fable 5.1 scores 19% on ARI Bench, the benchmark for recursive AI improvement. We know the Fable series is purposefully nerfed when it comes to AI improvement, so 19% is actually quite decent considering the original Fable scored 7%.
34
The new HY4-preview from Tencent scores 4% on ARI Bench, the benchmark for recursive AI improvement. It's a start... but still quite a ways to go.
58
Sorry GLM 5.3-Flash, you may be inexpensive but you only score 4% on ARI Bench, the benchmark for recursive AI improvement. Deepseek V4 Flash, GPT Luna and Qwen 3.8-27B are all way ahead on complicated, hour long tasks.
65
I'm sorry but guys, Ox Alpha only scores 1% on ARI Bench, the benchmark for recursive AI improvement. We re-ran it 3 times, with corrected harnesses... and it still struggled. Here we see when things are benchmaxxed.
1
686
Absolutely insane! Qwen 3.8 27b fastmtp-Q3 scored 19% on ARI Bench, the benchmark for recursive AI improvement. This is a tiny local LLM that outperformed many juggernauts in the industry. This almost doesn't make sense but we checked it twice.
4
199
GLM 5.3 scores 18% on ARI Bench, the benchmark for recursive AI improvement. It's a decent result but I was secretly hoping for more (Especially considering the price)
45