Is computer vision “solved”? Not yet Current models score 0% on ZeroBench 🧵1/6
57
239
2,546
1,449,880
Jonathan Roberts retweeted
Mistral Large 4 is one of the world's strongest AI models for cybersecurity.
96
201
2,745
122,536
Jonathan Roberts retweeted
Meet Mistral Large 4, aka Le Chonk. • 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. • State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. • Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. • Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.
Made with AI
1,947
4,471
39,741
5,085,761
Opus 5.5 (max) is no.2 on ZeroBench 43% pass@5 25% pass^5 Now, 77% of questions have been solved at least once
2
2
6
266
More details on the project page: zerobench.github.io/
24
All three GPT-6 models sit on the ZeroBench performance-cost pareto frontier pass@5 | pass^5 GPT-6 Astra (max) 51% | 35% GPT-6 Sol (max) 41% | 17% GPT-6 Luna (max) 20% | 7%
4
1
20
1,154
GPT-6 Astra (max) on ZeroBench: pass@5: 51% (prev. SOTA: 30%) pass^5: 35% (prev. SOTA: 13%) If this reflects a genuine capability improvement, it is a significant step forward in both performance and consistency
6
1
65
7,047
Jonathan Roberts retweeted
Now let’s take a look at how Muse Spark 1.2 performs across visual reasoning, chart understanding, and knowledge-intensive tasks.
4
6
67
10,037
Impressive! Muse Spark 1.2 w/ tools surpasses 50% pass@5 on ZeroBench!
Replying to @alexandr_wang
2/ muse spark 1.2 performs quite strongly across a wide variety of multimodal capabilities and evals.
3
340
Grok 4.6 (xhigh) on ZeroBench (no tool use): pass@5 / pass^5 Grok 4.6: 17% / 7% GPT-5.6 Sol (SOTA): 30% / 13% More evaluation details available on the project page
2
3
267
Jonathan Roberts retweeted
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
383
1,962
13,518
3,517,232
It’s been great seeing ZeroBench evaluated as part of several recent model releases from @Alibaba_Qwen, @Kimi_Moonshot and ByteDance We’ve added an externally reported results leaderboard and included these scores in our visualisations
1
4
336
As evaluation settings can differ, these reported results are shown for comparison but kept distinct from our official evals For each result, we include the available evaluation details reported in the original source zerobench.github.io/
54
We’ve updated the ZeroBench evaluation protocol after red-teaming with Fable 5 All ZeroBench results now use this updated protocol The benchmark itself hasn’t changed, but 0.71% of the grading decisions have Overall, pass@5 scores increased by an average of 0.79 pp
After our latest evals, these are the top 3 models on ZeroBench (no tool use): pass@5 / pass^5 GPT-5.6 Sol (max): 30% / 13% Claude Opus 5 (max): 26% / 11% Claude Fable 5 (max): 24% / 9% GPT-5.6 Sol is the first to reach the 30% human baseline on pass@5
1
8
380
After our latest evals, these are the top 3 models on ZeroBench (no tool use): pass@5 / pass^5 GPT-5.6 Sol (max): 30% / 13% Claude Opus 5 (max): 26% / 11% Claude Fable 5 (max): 24% / 9% GPT-5.6 Sol is the first to reach the 30% human baseline on pass@5
3
6
22
23,064
Across all our evals, 70% of ZeroBench questions have been answered correctly at least once Here’s the distribution for the top 5 models zerobench.github.io/
1
6
407
🎉Strong Qwen3.8-Max scores on ZeroBench: 24% pass@5 49% pass@5 w/ code interpreter
📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: - Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:github.com/qwen-code-dev-bot… - Real work, real results: Production-quality deliverables across hundreds of professions. - Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy. - Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction. 💰Pricing: Input: $2.0 / M tokens Output: $6.0 / M tokens Implicit Caching: $0.25 / M tokens Start building with Qwen3.8-Max! 🚀 📖 Blog: qwen.ai/blog?id=qwen3.8 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.8… ⚡ API: qwencloud.com/models/qwen3.8…
1
1
9
683
I’m at #ICML2026 in Seoul presenting our poster on ZeroBench later today! "ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models" Stop by to chat about benchmarking and frontier evals 📍 HALL A #4305 🕰️ 2:30 PM – 4:15 PM
1
3
14
1,705