Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring → arena.ai/jobs

US
Based in United States
Pinned Tweet
Introducing Agent Mode: Agentic AI is now measured in the Arena. Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more. It completes more complex tasks by using tools like web search, bash in a sandbox environment, image generation, file writing, and asking follow-up questions. Frontier models are waiting for you in Agent Mode to take on real-world tasks. GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and top open models. Test them yourself.
100
76
918
525,792
Arena.ai retweeted
GPT-6 Sol (Max) just landed in the Agent Arena with +7.7% net-improvement at #6 across 4K+ real-world agentic sessions from our global community of users! Compared with GPT 5.6 Sol (xHigh), this release is a +1.5-point lift in net improvement (+7.7% vs. +6.2%) and a point-rank move from #8 to #6, at half the per-token price. By signal GPT-6 Sol is strong in Confirmed Success at #4 with +11.4%, compared to 5.6 at #20 with 2.9%. Congrats to the @OpenAI team on GPT-6 Sol!
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
44
25
627
143,282
Arena.ai retweeted
Claude Opus 5.5 (Max) delivers top performance at a blended $16 per Mtoken, reshaping the Pareto frontier.
3
11
146
64,596
Arena.ai retweeted
Grok 4.7 by @SpaceXAI just landed in Agent Arena at #16, with a net improvement score of +3.96% Grok 4.7 (xHigh) shows improvement over previous versions, but comes with a comparable lift in price vs. Grok 4.6 (High): - Grok 4.7 (xHigh): $1.14 median per task / +3.96 net improvement - Grok 4.6 (High): $0.74 median per task / +1.22 net improvement - Grok 4.5: $0.43 median per task / +1.50 net improvement By signal, Grok 4.7 (xHigh) ranks: - #6 Confirmed Success (+10.41%) - #12 Bash Recovery (+5.98%) - #16 Praise vs Complaint (+4.93%) - #28 Steerability (-1.89%) - No issues with Tool Hallucination (+0.36%) Congrats to the @SpaceXAI team on this solid release!
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
50
27
492
119,742
Dive into all the recent releases placement in the Text Arena: arena.ai/leaderboard/text
1
6
10,019
Claude Opus 5.5 (High) reshaped the Text Arena Pareto frontier with 1509 pts at $16/MToken (input / output price per million tokens)!
6
2
49
13,653
Claude Opus 5.5 (High) debuts at #1 in Text Arena with 1509 pts! That’s an 18-pt improvement over Opus 5 (High), now at #11. Opus 4.6 (High) remains in the #2 spot, just 4 pts off the lead. This release gives @AnthropicAI all six top spots in Text Arena! Opus 5.5 (High) also landed on the Text Arena Pareto frontier at a blended $16/MToken, based on input/output pricing per million tokens! See its placement in the post below. Congrats to the @AnthropicAI team on the release!
107
97
1,520
245,296
Agent Arena’s Pareto frontier compares net improvement with median cost per task from real Agent Mode tasks. Dive into the Agent Arena Pareto frontier at: arena.ai/leaderboard/agent/p…
1
6
6,827
GPT-6 Sol (Max) by @OpenAI has reshaped the Agent Arena Pareto frontier with a +7.7% net improvement at $0.75 median cost/task! It sits just 0.60 percentage points below Claude Fable 5 (High) (+8.3%) while costing 56% less per task ($0.75 vs. $1.72). Dive into the Agent Arena Pareto frontier details at the link below. Congrats to the @OpenAI team on this release!
GPT-6 Sol (Max) just landed in the Agent Arena with +7.7% net-improvement at #6 across 4K+ real-world agentic sessions from our global community of users! Compared with GPT 5.6 Sol (xHigh), this release is a +1.5-point lift in net improvement (+7.7% vs. +6.2%) and a point-rank move from #8 to #6, at half the per-token price. By signal GPT-6 Sol is strong in Confirmed Success at #4 with +11.4%, compared to 5.6 at #20 with 2.9%. Congrats to the @OpenAI team on GPT-6 Sol!
29
29
342
49,707
Agent Arena measures net improvement relative to the average model across real-world, long-horizon tasks using web, filesystem, and terminal tools. Dive into the Agent Arena leaderboard at: arena.ai/leaderboard/agent
1
13
6,898
This Week in the Arena: the arrival of GPT-6 Sol, Claude Opus 5.5, and the latest AI news in less than 60 seconds. Watch the full update. piped.video/watch?v=JLiguQgf…
7
4
109
19,677
We analyzed how @claudeai’s Opus 5.5 writes compared with Opus 5 across high-reasoning Text Arena outputs. 10 of 12 writing measures moved in a better direction. Opus 5.5 should be easier to read: - Long content words fall from 41.7% to 38.6%, the lowest share of any Claude model we analyzed. - Sentences are 17% shorter on average, dropping from 12.14 to 10.03 words. The tradeoff is length. Answers get 6% wordier, rising from 453 to 481 words on average, making Opus 5.5 give the longest answers across the Opus family. It also sounds less recognizably AI on two familiar tells: - 95% fewer em dashes - 73% fewer semicolons But a new giveaway may be emerging. Hedges and caveats such as “perhaps” and “arguably” rise 97%, from 0.39 to 0.77 per 1,000 words, the highest rate of any Claude model we analyzed. What do you think: does Opus 5.5 read more naturally?
95
120
2,087
1,489,626
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. See the Agent Arena leaderboard details at: arena.ai/leaderboard/agent
1
12
7,926
Arena.ai retweeted
Big news: Claude Opus 5.5 (Max) by @AnthropicAI just topped #1 in Code Arena: WebDev with 1818 pts and reshapes the Pareto frontier! This is a solid +26pt lead ahead of the next best model, GPT-6 Astra (Max) and a huge +126pt improvement over previous Opus 5 (Max) at 1692. Across categories, we can see that Opus 5.5 (Max) lands in the top spots across domains for: - #1 Brand and Marketing, Reference-Based Design, Data & Analytics, Simulations and Gaming! - #2 Consumer Product Stay tuned for more domain and categorical insights to land as more votes come in. Congrats to @AnthropicAI on the release!
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
77
139
1,838
144,492
The redesigned Arena Leaderboard Overview is live! The frontier is moving quickly, so we built a central hub to help you keep up. The new Arena Leaderboard experience brings all our live signals into one place for researchers, builders, and our global community of users who power our model evaluations. Here’s what’s new: - New Release Models: a real-time feed of the latest models added to Arena - Performance by Category: surfacing live sessions, and the top models by categories like Agents, Images, and Coding. - Model Capabilities: first impressions of new models, straight from the Arena team - Arena News: the latest research, features and updates from around the Arena universe For example, now you can see scores across arenas for models like GPT-6 Sol and Claude Opus 5.5, alongside signals like their category standings and capabilities, their latest news and more. Explore scores and now much more at our redesigned leaderboard. Find the link below.
13
7
224
23,515
Real-world results are in for GPT-6 Luna (Max). It just landed #24 in the Code Arena: WebDev with 1593 pts! GPT-6 Luna (Max) by @OpenAI is a significant +74pt improvement from GPT-5.6 Luna (xHigh) which is now ranked at #41. The latest Luna delivers on par with lower-cost models like: Gemini 3.7 Flash High and Qwen 3.8-27B, but at a fraction of the price (blended $.40/Mtoken). Congrats to the @OpenAI team!
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
59
54
1,046
157,189
ICYMI: Claude Opus 5.5 (Max) landed #1 in the Code Arena: WebDev with a price 60% lower than the next best model: GPT-6 Astra (Max). In the Arena, all scores are based on live, real-world use by our global community. Check out first impressions with our AI capability expert, @petergostev and let us know what you think. piped.video/-0PE_bHHoRk
Replying to @arena
Claude Opus 5.5 (Max) delivers top performance at a blended $16 per Mtoken, reshaping the Pareto frontier.
26
26
405
52,922