World's first benchmark for real-world visuals with 6M+ creators and counting. Made by @intelligence_ai

Design Arena is excited to introduce the company behind our work: Intelligence. Over the past year, millions of people have helped us evaluate how well AI models perform on real-world creative tasks - where quality is subjective, outcomes are difficult to verify, and human judgment matters. Design Arena is only the beginning. Let us know what you want to see next!
Today, we're introducing @Intelligence_ai. In 6 months, as a team of 10, we scaled from $5M to $60M ARR and 5.5M users across 190+ countries. We raised a $7.9M seed, led by @IndexVentures with participation from @conviction, @A_StarVC, and @combinator to build DesignArena, a universal interface for accessing and evaluating the world's AI capabilities. Most evaluations try to simulate the real-world. We believe the real-world is the ultimate verifier. People come to @DesignArena with a request. Models compete to fulfill the request, and users determine what works best for them. Their live user behavior evaluates the models, improves how work is routed, and helps people access the right intelligence. We've helped the world's leading frontier labs break the news on their SOTA capabilities. What's the limit? Join us and find out.
19
8
176
36,079
We were surprised to find that 50% of Open AI's GPT-6 Astra's driving games have mistakenly reversed steering mechanics. So what makes users prefer games, and where do the models fall short? We analyzed 4,300+ tournaments in DesignArena to find out. Beyond looks, there are other interesting axes to explore: whether all scene objects obey the same physics, whether the controls work the way players expect, and whether the world interactions hold up once you start moving. Full video in the thread. Try it for free on DesignArena!
9
7
143
18,414
Design Arena retweeted
More soon :)
10
6
128
7,019
Recraft V4.1 Flash by @recraftai is now available on Design Arena! Built for speed, Recraft V4.1 Flash generates images in as little as 1.3 seconds, keeping creative workflows moving through rapid iteration and variations. At a fraction of the cost of V4.1, it brings the same Recraft look across photography, portraits, complex scenes, mood and atmosphere, and early concepts. Congrats to the @recraftai team on the launch!
The fastest image model on the market is here. Zero lag in your workflow. Recraft V4.1 Flash produces images in 1.3 seconds. Pick Flash from the model picker, generate, then hit Refine to clean up the details. Check it out and give it a try ↓
2
2
41
2,256
GPT-6 Sol & GPT-6 Luna by @OpenAI are now available on Design Arena! Built with the same advances behind GPT-6 Astra, Sol and Luna bring improved performance across professional work, factuality, coding, computer use, and alignment in faster, more affordable models. With improved caching and inference, both models are designed to make advanced AI more accessible for everyday tasks and applications at scale. Congrats to the @OpenAI team on the launch!
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
3
2
52
2,433
Claude Opus 5.5 by @AnthropicAI is now available on Design Arena! Built for agentic coding and long-running tasks, Claude Opus 5.5 is designed to handle complex workflows like full-stack and game development. It identifies root causes before making changes, verifies its work throughout the process, and brings Anthropic’s strongest vision capabilities for understanding screenshots and diagrams. Congrats to the @AnthropicAI team on the launch!
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
4
3
59
2,650
MiniMax H3 Max by @fal takes the #1 spot on Video Editing by Design Arena with an Elo of 1373. The model ranks ahead of MiniMax H3 by @MiniMax and @GoogleDeepMind's Gemini Omni Flash and Gemini Omni Flash 1.1. It's also in a league of its own on the Pareto frontier for Preference vs Speed and Preference vs Price. Congratulations to the @fal team on post-training an exceptional model.
15
33
214
82,750
BREAKING: MiMo-V2.6-Pro by @XiaomiMiMo lands at #8 overall (#3 open-weight) on Design Arena with an Elo of 1338. This is an impressive 54-point and 22-position increase from MiMo-V2.5-Pro. MiMo-V2.6-Pro also reaches #4 overall in Website (#2 open-weight) and #6 overall in Agentic Frontend Development (#2 open-weight), showing strong performance across both direct generation and agentic coding. This places @XiaomiMiMo's new model among leading models such as GPT-5.6 Sol by @OpenAI, Claude Opus 5 by @Anthropic, and Kimi K3 by @MoonshotAI. Congratulations to the @XiaomiMiMo team on returning to a top-10 placement on Design Arena!
33
14
201
34,349
MiMo-V2.6-Pro by @XiaomiMiMo also has strong agentic capabilities, landing at #6 in Agentic Frontend Development.
2
8
1,148
In the Website category, MiMo-V2.6-Pro is a 49-point increase over MiMo-V2.5-Pro.
2
1
9
1,617
Grok 4.7 by @SpaceXAI is now available on Design Arena! Built for coding and knowledge work, Grok 4.7 is designed to take on more difficult, long-running tasks with stronger reasoning, self-verification, and context management. Compared to Grok 4.6, it uses a larger base model and longer reinforcement learning to improve performance on complex tasks that require sustained work. Congrats to the @SpaceXAI team on the launch!
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
7
45
2,514
BREAKING: Freya TTS Adam v1 by @FreyaVoice takes the #1 spot on AudioRealismBench with an Elo of 1418. The model ranks ahead of Bland Speech v3 by @usebland and Kalpa TTS Beta v0.1 by @KalpaLabsAI. AudioRealismBench measures how human-like speech sounds through blind comparisons by vetted native speakers across phone agents, conversational assistants, and spoken explainers. Congratulations to @TungaBayrak and the entire @FreyaVoice team. We’re excited to see what comes next!
4
8
47
6,773
Gemini 3.8 Live by @GoogleDeepMind is now available on Design Arena! Built for more interactive and capable conversations, this model brings upgraded reasoning, near real-time visual understanding, automatic detection across 97 languages, and background tool calling. Congrats to the @GoogleDeepMind team on the launch!
We’re introducing Gemini 3.8 Live and 3.8 Live Extended Thinking – our best conversational AI. The models talk, think, and handle tasks in the background without breaking your flow. 🧵
6
2
70
4,834
BREAKING: @PrunaAI establishes a new Pareto frontier for Speed vs. Preference on Video Editing Arena! Its debut video editing model, P‑Video‑Edit, offers both Draft and Standard modes - giving users a faster option for iteration before generating the final-quality edit. Both configurations are live now. Try them today, and congratulations to the team!
3
9
58
3,789
Design Arena retweeted
DeepSeek-V4.1-Flash by @deepseek_ai is back in the top 10, but what's crazier is that 6 out of the 7 models on the Price vs Preference Pareto frontier are all open weights
BREAKING: DeepSeek‑V4.1‑Flash takes 6th overall on Design Arena with an Elo of 1347! This marks a 39-position jump over the next-highest DeepSeek model - and DeepSeek’s return to a top-10 placement on Design Arena. The model ranks in the same performance band as Claude Fable 5.1 on real-world frontend design tasks. Congratulations to the @deepseek_ai team!
2
2
33
4,851
BREAKING: DeepSeek‑V4.1‑Flash takes 6th overall on Design Arena with an Elo of 1347! This marks a 39-position jump over the next-highest DeepSeek model - and DeepSeek’s return to a top-10 placement on Design Arena. The model ranks in the same performance band as Claude Fable 5.1 on real-world frontend design tasks. Congratulations to the @deepseek_ai team!
40
44
722
59,772
Open weights are on the verge of sweeping Design Arena’s price–preference frontier. 6 of the 7 models defining it are already open weights. If Meta’s planned Muse Spark open-weights release includes Muse Spark 1.3 Max, all seven would be open weights. The frontier could soon be entirely open.
2
2
31
11,165
DeepSeek-V4.1-Flash by @deepseek_ai is now available on Design Arena! Built as the smallest model in DeepSeek’s new architecture family, DeepSeek-V4.1-Flash brings native visual understanding with a focus on faster inference, higher throughput, and greater efficiency. The model is designed to deliver stronger capabilities while scaling efficiently to larger models. Congrats to the @deepseek_ai team on the launch!
3
4
100
4,045