Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring → arena.ai/jobs

US
Pinned Tweet
Introducing Agent Mode: Agentic AI is now measured in the Arena. Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more. It completes more complex tasks by using tools like web search, bash in a sandbox environment, image generation, file writing, and asking follow-up questions. Frontier models are waiting for you in Agent Mode to take on real-world tasks. GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and top open models. Test them yourself.
99
76
911
521,648
GPT-6 Sol (Max) by @OpenAI has reshaped the Agent Arena Pareto frontier with a +7.7% net improvement at $0.75 median cost/task! It sits just 0.60 percentage points below Claude Fable 5 (High) (+8.3%) while costing 56% less per task ($0.75 vs. $1.72). Dive into the Agent Arena Pareto frontier details at the link below. Congrats to the @OpenAI team on this release!
GPT-6 Sol (Max) just landed in the Agent Arena with +7.7% net-improvement at #6 across 4K+ real-world agentic sessions from our global community of users! Compared with GPT 5.6 Sol (xHigh), this release is a +1.5-point lift in net improvement (+7.7% vs. +6.2%) and a point-rank move from #8 to #6, at half the per-token price. By signal GPT-6 Sol is strong in Confirmed Success at #4 with +11.4%, compared to 5.6 at #20 with 2.9%. Congrats to the @OpenAI team on GPT-6 Sol!
19
21
226
29,988
Agent Arena’s Pareto frontier compares net improvement with median cost per task from real Agent Mode tasks. Dive into the Agent Arena Pareto frontier at: arena.ai/leaderboard/agent/p…
1
3
3,842
GPT-6 Sol (Max) just landed in the Agent Arena with +7.7% net-improvement at #6 across 4K+ real-world agentic sessions from our global community of users! Compared with GPT 5.6 Sol (xHigh), this release is a +1.5-point lift in net improvement (+7.7% vs. +6.2%) and a point-rank move from #8 to #6, at half the per-token price. By signal GPT-6 Sol is strong in Confirmed Success at #4 with +11.4%, compared to 5.6 at #20 with 2.9%. Congrats to the @OpenAI team on GPT-6 Sol!
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
31
16
386
78,854
Agent Arena measures net improvement relative to the average model across real-world, long-horizon tasks using web, filesystem, and terminal tools. Dive into the Agent Arena leaderboard at: arena.ai/leaderboard/agent
1
7
4,115
This Week in the Arena: the arrival of GPT-6 Sol, Claude Opus 5.5, and the latest AI news in less than 60 seconds. Watch the full update. piped.video/watch?v=JLiguQgf…
7
2
87
13,788
We analyzed how @claudeai’s Opus 5.5 writes compared with Opus 5 across high-reasoning Text Arena outputs. 10 of 12 writing measures moved in a better direction. Opus 5.5 should be easier to read: - Long content words fall from 41.7% to 38.6%, the lowest share of any Claude model we analyzed. - Sentences are 17% shorter on average, dropping from 12.14 to 10.03 words. The tradeoff is length. Answers get 6% wordier, rising from 453 to 481 words on average, making Opus 5.5 give the longest answers across the Opus family. It also sounds less recognizably AI on two familiar tells: - 95% fewer em dashes - 73% fewer semicolons But a new giveaway may be emerging. Hedges and caveats such as “perhaps” and “arguably” rise 97%, from 0.39 to 0.77 per 1,000 words, the highest rate of any Claude model we analyzed. What do you think: does Opus 5.5 read more naturally?
65
78
1,253
712,056
Grok 4.7 by @SpaceXAI just landed in Agent Arena at #16, with a net improvement score of +3.96% Grok 4.7 (xHigh) shows improvement over previous versions, but comes with a comparable lift in price vs. Grok 4.6 (High): - Grok 4.7 (xHigh): $1.14 median per task / +3.96 net improvement - Grok 4.6 (High): $0.74 median per task / +1.22 net improvement - Grok 4.5: $0.43 median per task / +1.50 net improvement By signal, Grok 4.7 (xHigh) ranks: - #6 Confirmed Success (+10.41%) - #12 Bash Recovery (+5.98%) - #16 Praise vs Complaint (+4.93%) - #28 Steerability (-1.89%) - No issues with Tool Hallucination (+0.36%) Congrats to the @SpaceXAI team on this solid release!
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
47
24
463
103,888
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. See the Agent Arena leaderboard details at: arena.ai/leaderboard/agent
1
11
6,818
Arena.ai retweeted
Big news: Claude Opus 5.5 (Max) by @AnthropicAI just topped #1 in Code Arena: WebDev with 1818 pts and reshapes the Pareto frontier! This is a solid +26pt lead ahead of the next best model, GPT-6 Astra (Max) and a huge +126pt improvement over previous Opus 5 (Max) at 1692. Across categories, we can see that Opus 5.5 (Max) lands in the top spots across domains for: - #1 Brand and Marketing, Reference-Based Design, Data & Analytics, Simulations and Gaming! - #2 Consumer Product Stay tuned for more domain and categorical insights to land as more votes come in. Congrats to @AnthropicAI on the release!
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
76
139
1,826
140,405
The redesigned Arena Leaderboard Overview is live! The frontier is moving quickly, so we built a central hub to help you keep up. The new Arena Leaderboard experience brings all our live signals into one place for researchers, builders, and our global community of users who power our model evaluations. Here’s what’s new: - New Release Models: a real-time feed of the latest models added to Arena - Performance by Category: surfacing live sessions, and the top models by categories like Agents, Images, and Coding. - Model Capabilities: first impressions of new models, straight from the Arena team - Arena News: the latest research, features and updates from around the Arena universe For example, now you can see scores across arenas for models like GPT-6 Sol and Claude Opus 5.5, alongside signals like their category standings and capabilities, their latest news and more. Explore scores and now much more at our redesigned leaderboard. Find the link below.
11
7
213
21,396
Real-world results are in for GPT-6 Luna (Max). It just landed #24 in the Code Arena: WebDev with 1593 pts! GPT-6 Luna (Max) by @OpenAI is a significant +74pt improvement from GPT-5.6 Luna (xHigh) which is now ranked at #41. The latest Luna delivers on par with lower-cost models like: Gemini 3.7 Flash High and Qwen 3.8-27B, but at a fraction of the price (blended $.40/Mtoken). Congrats to the @OpenAI team!
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
57
54
1,035
152,988
ICYMI: Claude Opus 5.5 (Max) landed #1 in the Code Arena: WebDev with a price 60% lower than the next best model: GPT-6 Astra (Max). In the Arena, all scores are based on live, real-world use by our global community. Check out first impressions with our AI capability expert, @petergostev and let us know what you think. piped.video/-0PE_bHHoRk
Replying to @arena
Claude Opus 5.5 (Max) delivers top performance at a blended $16 per Mtoken, reshaping the Pareto frontier.
25
26
404
51,355
Check out side-by-side outputs generated on Arena from Claude Opus 5.5 by @AnthropicAI and GPT-6 Sol by @OpenAI. Scores for @claudeai Opus 5.5 are coming soon. Real-world tasks from our global community of users power the Arena leaderboards. Head to Arena now to test it out, and stay tuned!
60
82
1,579
145,212
Test it out yourself with this prompt from @petergostev: Tell the Trojan Horse story in a 20-second HTML animation painted around an ancient Greek vase. Rotate the vase as the horse enters the city, night falls, and hidden warriors emerge. Draw everything in code; no external assets or website UI.
1
1
33
8,418
Reminder: the Fall 2026 cycle of Arena's Academic Partnerships Program has begun! We're accepting new proposals from tenure-track faculty at U.S. universities working on the scientific foundations of AI evaluation: evaluation and ranking methodology, measurement and validity, learning from human preference, and safety and alignment. Funding of up to $50,000 per project is available. The Fall 2026 deadline to submit is October 30, 2026. Learn more about the program in thread 👇
5
3
67
14,847