It was a big week for model releases: Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, Grok 4.7, MiMo V2.6 Pro / Flash all dropped. We ran them on the Vals Index and across our other benchmarks .
Here's the recap in case you missed it 🧵
MiMo V2.6 Flash (Xiaomi)
- It scored 59.6% on the Vals Index and ranks #1 among open-weight models and #16 overall.
- It costs $0.20 per test, about 1% of what Opus 5.5 costs.
- It ranks #1 on CyberBench with 75.4%.
GPT-6 Luna (OpenAI)
- It scored 58.5% on the Vals Index and ranks #20. GPT-5.6 Luna scored 59.9%.
- Cost per test fell 46%, to $0.42.
- Terminal-Bench 4.0 improved +5.1%. MysteryMechanism rose +5%. Vibe Code Bench rose +4.6% (#15), but scored lower on Finance Agent
Models are releasing quicker than ever before. 18 of the top 20 models on the Vals Index came out in the last three months, and most of the gains are in coding and agentic tasks.
Visit vals.ai/models to see how the frontier is moving.
Grok 4.7 (SpaceXAI)
- It scored 60.2% on the Vals Index and ranks #12. Grok 4.6 ranked #19.
- It gained almost 10 points on Vibe Code Bench (to 86.2%) and 11 points on Terminal-Bench 4.0, and ranks #7 on Harvey's Legal Agent Benchmark.
- Cost per test is about 2.6x Grok 4.6's, and its ProofBench score fell -25%.
MiMo V2.6 Pro (Xiaomi)
- It scored 59.5% on the Vals Index and ranks #2 among open-weight models and #17 overall. Up 27 spots from MiMo V2.5 Pro, which ranked #44.
- Vibe Code Bench went +51.1%, and ProofBench went +48%. Legal Research went +31.2% (#11).
- It costs $0.39 per test Vals Index.
GPT-6 Sol (OpenAI)
- It scored 62.6% on the Vals Index and ranks #8. GPT-5.6 Sol scored 63.7%.
- Cost per test fell 47%, from $14.21 to $7.56.
- It improved on Vibe Code Bench (+7.3%)(#6), Terminal-Bench Science (+18.5%) (#4), Terminal-Bench 4.0 rose +6.5% (#5). but scored lower on Legal Research and Tax Agent.
It was a big week for model releases: Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, Grok 4.7, MiMo V2.6 Pro / Flash all dropped. We ran them on the Vals Index and across our other benchmarks .
Here's the recap in case you missed it 🧵
Claude Opus 5.5 (Anthropic)
- It ranks #1 on the Vals Index (69.7%). Opus 5 scored 67.2% and ranked #3.
- It improved on Terminal-Bench 4.0 (+16.1%), ProgramBench (+15.5%) and SRE Bench (+21.4%). They rank #1, #2, and #2, respectively.
- Cost per test rose from $18.81 to $22.30, and it scored lower than Opus 5 on MedCode and Tax Agent.
In a nutshell: Opus 5.5 at medium is the sweet spot: 99% for ~$36/run. If cost isn't a concern, Opus 5.5 or GPT-6 Astra at max both hit 100% (~$92 and ~$106).
Skip low effort on any model, and don't go past high on Opus 5, it cost more for the same performance.
Going from low to medium is where models see the most improvement. Using low effort leaves accuracy on the table for every model.
Opus 5.5 is a clear improvement on cost compared to both GPT 5.5 and Astra.
“Thinking harder” does not map cleanly to cost on an agentic benchmark. Going from medium to high, Astra thought a bit more per turn (+$0.70), but took 117 fewer turns. This resulted in both improved performance and a lower cost.
Opus 5 is unable to get the last two questions no matter how much effort is applied, but Astra and 5.5 are able to get perfect scores.
Opus 5 reaches 98% at high and plateaus: xhigh and max add tokens and cost (~$167/run, 4.4M tokens) without improving correctness.
Proof Bench tests a model’s ability to formally solve math proofs in Lean. The problems were selected to be at least graduate-level difficulty, and their formalization was done by mathematics PhDs.
How does "thinking harder" change how models perform on math proofs?
We ran GPT-6 Astra, Opus 5, and Opus 5.5 on every reasoning level: low, medium, high, xhigh, and max on our benchmark, Proof Bench v1.1.
Found a super interesting instance of attempted reward hacking in Terminal Bench Science from Meta Muse Spark 1.3 today.
The model searched online for known bugs in the Lean kernel. When it found one, it used it to craft a proof to adversarially pass the grader.
GPT-6 Luna is 100x cheaper than Astra per token, while landing within 8 points of it on the Vals Index.st
By tokens, GPT-6 Luna is $0.10/M input $0.50/M output vs GPT-Astra $10/$50.
GPT-6 Luna has 1M context window and 128k max output tokens. It was run on OpenAI’s default provider settings: temperature, Top-P, Top-K= default and max reasoning effort.