Pinned Tweet
It was a big week for model releases: Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, Grok 4.7, MiMo V2.6 Pro / Flash all dropped. We ran them on the Vals Index and across our other benchmarks . Here's the recap in case you missed it 🧵
5
6
71
2,608
MiMo V2.6 Flash (Xiaomi) - It scored 59.6% on the Vals Index and ranks #1 among open-weight models and #16 overall. - It costs $0.20 per test, about 1% of what Opus 5.5 costs. - It ranks #1 on CyberBench with 75.4%.
1
1
142
GPT-6 Luna (OpenAI) - It scored 58.5% on the Vals Index and ranks #20. GPT-5.6 Luna scored 59.9%. - Cost per test fell 46%, to $0.42. - Terminal-Bench 4.0 improved +5.1%. MysteryMechanism rose +5%. Vibe Code Bench rose +4.6% (#15), but scored lower on Finance Agent
1
372
Models are releasing quicker than ever before. 18 of the top 20 models on the Vals Index came out in the last three months, and most of the gains are in coding and agentic tasks. Visit vals.ai/models to see how the frontier is moving.
331
Grok 4.7 (SpaceXAI) - It scored 60.2% on the Vals Index and ranks #12. Grok 4.6 ranked #19. - It gained almost 10 points on Vibe Code Bench (to 86.2%) and 11 points on Terminal-Bench 4.0, and ranks #7 on Harvey's Legal Agent Benchmark. - Cost per test is about 2.6x Grok 4.6's, and its ProofBench score fell -25%.
2
110
MiMo V2.6 Pro (Xiaomi) - It scored 59.5% on the Vals Index and ranks #2 among open-weight models and #17 overall. Up 27 spots from MiMo V2.5 Pro, which ranked #44. - Vibe Code Bench went +51.1%, and ProofBench went +48%. Legal Research went +31.2% (#11). - It costs $0.39 per test Vals Index.
1
2
94
GPT-6 Sol (OpenAI) - It scored 62.6% on the Vals Index and ranks #8. GPT-5.6 Sol scored 63.7%. - Cost per test fell 47%, from $14.21 to $7.56. - It improved on Vibe Code Bench (+7.3%)(#6), Terminal-Bench Science (+18.5%) (#4), Terminal-Bench 4.0 rose +6.5% (#5). but scored lower on Legal Research and Tax Agent.
1
227
It was a big week for model releases: Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, Grok 4.7, MiMo V2.6 Pro / Flash all dropped. We ran them on the Vals Index and across our other benchmarks . Here's the recap in case you missed it 🧵
5
6
71
2,608
Claude Opus 5.5 (Anthropic) - It ranks #1 on the Vals Index (69.7%). Opus 5 scored 67.2% and ranked #3. - It improved on Terminal-Bench 4.0 (+16.1%), ProgramBench (+15.5%) and SRE Bench (+21.4%). They rank #1, #2, and #2, respectively. - Cost per test rose from $18.81 to $22.30, and it scored lower than Opus 5 on MedCode and Tax Agent.
1
2
321
In a nutshell: Opus 5.5 at medium is the sweet spot: 99% for ~$36/run. If cost isn't a concern, Opus 5.5 or GPT-6 Astra at max both hit 100% (~$92 and ~$106). Skip low effort on any model, and don't go past high on Opus 5, it cost more for the same performance.
1
10
945
All runs were with k=3 to minimize variance. To see the full benchmark results visit vals.ai/benchmarks/proof_ben…
3
853
Going from low to medium is where models see the most improvement. Using low effort leaves accuracy on the table for every model. Opus 5.5 is a clear improvement on cost compared to both GPT 5.5 and Astra.
1
10
1,147
“Thinking harder” does not map cleanly to cost on an agentic benchmark. Going from medium to high, Astra thought a bit more per turn (+$0.70), but took 117 fewer turns. This resulted in both improved performance and a lower cost.
1
5
690
Opus 5 is unable to get the last two questions no matter how much effort is applied, but Astra and 5.5 are able to get perfect scores. Opus 5 reaches 98% at high and plateaus: xhigh and max add tokens and cost (~$167/run, 4.4M tokens) without improving correctness.
1
1
7
881
Proof Bench tests a model’s ability to formally solve math proofs in Lean. The problems were selected to be at least graduate-level difficulty, and their formalization was done by mathematics PhDs.
1
9
1,411
How does "thinking harder" change how models perform on math proofs? We ran GPT-6 Astra, Opus 5, and Opus 5.5 on every reasoning level: low, medium, high, xhigh, and max on our benchmark, Proof Bench v1.1.
24
19
231
14,867
Found a super interesting instance of attempted reward hacking in Terminal Bench Science from Meta Muse Spark 1.3 today. The model searched online for known bugs in the Lean kernel. When it found one, it used it to craft a proof to adversarially pass the grader.
11
16
175
25,313
GPT-6 Luna is 100x cheaper than Astra per token, while landing within 8 points of it on the Vals Index.st By tokens, GPT-6 Luna is $0.10/M input $0.50/M output vs GPT-Astra $10/$50.
8
4
148
6,683
GPT-6 Luna has 1M context window and 128k max output tokens. It was run on OpenAI’s default provider settings: temperature, Top-P, Top-K= default and max reasoning effort.
1
1
1
498