We’re excited to introduce Intelligence. If you’re building frontier models or working on hard-to-verify evaluations, we’d love to talk.
Today, we're introducing @Intelligence_ai. In 6 months, as a team of 10, we scaled from $5M to $60M ARR and 5.5M users across 190+ countries. We raised a $7.9M seed, led by @IndexVentures with participation from @conviction, @A_StarVC, and @combinator to build DesignArena, a universal interface for accessing and evaluating the world's AI capabilities. Most evaluations try to simulate the real-world. We believe the real-world is the ultimate verifier. People come to @DesignArena with a request. Models compete to fulfill the request, and users determine what works best for them. Their live user behavior evaluates the models, improves how work is routed, and helps people access the right intelligence. We've helped the world's leading frontier labs break the news on their SOTA capabilities. What's the limit? Join us and find out.
9
1
109
44,867
Intelligence retweeted
It turns out, 50% of @OpenAI GPT-6 Astra's driving games go left when the user steers right. We looked at thousands of user-created games on @DesignArena to find the common failure cases in web-based game dev. It turns out, the details matter. Let's say there's a boat floating on the water. Is the code animating the boat in sync with the waves? Or let's say we have to place trees on a mountain. Should the trees be allowed to spawn on a cliff? Or a mountain peak? Or be penalized for floating above the ground height grid? Find this analysis of where the models fall short of making a great game, and more, in our new video.
We were surprised to find that 50% of Open AI's GPT-6 Astra's driving games have mistakenly reversed steering mechanics. So what makes users prefer games, and where do the models fall short? We analyzed 4,300+ tournaments in DesignArena to find out. Beyond looks, there are other interesting axes to explore: whether all scene objects obey the same physics, whether the controls work the way players expect, and whether the world interactions hold up once you start moving. Full video in the thread. Try it for free on DesignArena!
3
3
28
3,901
Intelligence retweeted
We were surprised to find that 50% of Open AI's GPT-6 Astra's driving games have mistakenly reversed steering mechanics. So what makes users prefer games, and where do the models fall short? We analyzed 4,300+ tournaments in DesignArena to find out. Beyond looks, there are other interesting axes to explore: whether all scene objects obey the same physics, whether the controls work the way players expect, and whether the world interactions hold up once you start moving. Full video in the thread. Try it for free on DesignArena!
5
4
71
10,544
Intelligence retweeted
We are close to the entire frontier being open.
Replying to @DesignArena
Open weights are on the verge of sweeping Design Arena’s price–preference frontier. 6 of the 7 models defining it are already open weights. If Meta’s planned Muse Spark open-weights release includes Muse Spark 1.3 Max, all seven would be open weights. The frontier could soon be entirely open.
1
8
944
Intelligence retweeted
BREAKING: GPT‑Image‑2.5 takes the top two spots on Image Editing Arena! Sunburst debuts in 1st place with an Elo of 1386, followed by Flare in 2nd with 1360. @OpenAI now holds all three top positions - and has held the top position since GPT‑Image‑2 launched in April. GPT‑Image‑2.5 also generates edits up to 6.2× faster than its predecessor, establishing a new Pareto frontier for Speed vs. Preference. Huge congratulations to the team!
10
19
272
18,920
We’re entering the era of faster and affordable models at the highest quality.
BREAKING: Muse Spark 1.3 (xhigh) takes 1st overall on Website Arena with an Elo of 1362! This is a jump of 5 positions from Muse Spark 1.2, establishing a new Pareto frontier for Speed and Price. Only a month after the release of Muse Spark 1.2, @AIatMeta has topped this category on Design Arena. Note: GPT-6 Astra is still pending final results. Congrats to the @Meta team!
2
25
1,854
Intelligence retweeted
BREAKING: GPT-6 Astra (xhigh) takes #1 on Design Arena’s 3D Design leaderboard with an Elo of 1495. It leads Kimi K3 by 65 points and improves 79 points over GPT-5.6 Sol (xhigh), OpenAI’s previous top-performing model. A decisive new SOTA for 3D generation. Congratulations to the @OpenAI team!
14
65
719
31,066
Intelligence retweeted
See an example, created in a single file, based off of a user's reference image by Astra:
1
2
27
2,346
Intelligence retweeted
August was a busy month for video generation. MiniMax H3 Max, post-trained by @fal, took the top spot on Image-to-Video, building on the strength of MiniMax H3 by @MiniMax_AI. Wan 3.0 by @AlibabaGroup took 3rd overall, a 12-position leap from Wan 2.7. We also saw @bfl_ai debut its first video model, FLUX 3 Video, with a strong overall performance. These releases pushed more than just quality. H3 Max cut generation time by 46x compared to H3 in our testing, while Omni 1.1 by @GoogleDeepMind introduced cheaper, faster drafts - redefining the SOTA baseline across quality, speed, cost, and control. Four August releases now sit in the top 11, including two of the top three. Here’s where the Image-to-Video leaderboard stands heading into September:
7
8
91
11,180
Intelligence retweeted
BREAKING: MiniMax H3 Max sets the new Pareto Frontier for video generation, nearly 50x faster than the base model. This model is post-trained by @fal on @MiniMax_AI H3, and it's in a league of its own: no other Image to Video model on the arena delivers higher preference at a lower generation time. Its Image-to-Video generation time is just 6.4 seconds, 18x faster than average, and its Text-to-Video generation time is just 4.7 seconds, 24x faster than the average. Huge congratulations to the @fal team on this release!
39
127
1,264
333,855
Intelligence retweeted
BREAKING: GLM-5.3 Flash by @Zai_org debuts at 6th overall on Design Arena with a preliminary Elo of 1343. This places the MIT-licensed, natively multimodal model in the same performance band as GLM-5.3 - and ahead of Claude Fable 5 and Muse Spark 1.2. At $0.15/1M input tokens and $0.50/1M output tokens, it pairs strong performance with highly competitive pricing. Stay tuned for its performance across our agentic and multimodal evaluations! Congratulations to the @Zai_org team!
22
54
497
73,291
Intelligence retweeted
BREAKING: Muse Spark 1.2 by @AIatMeta takes 1st for Video-to-Website with an Elo rating of 1279 and impressive scores across all of our multimodal code categories. Muse Spark 1.2 also takes 2nd for Image-to-HTML with an Elo rating of 1252 and 3rd for Image-to-Frontend with Elo rating of 1272. At $1.25/1M input tokens and $4.25/1M output tokens, it lands on the price-preference Pareto frontier for all three categories. Congrats to the @Meta team on the achievement!
4
33
254
92,975
Intelligence retweeted
BREAKING: GLM-5.3 by @Zai_org places 3rd overall on Design Arena with an Elo of 1351. This is a 6-position improvement from GLM-5.2, and makes GLM-5.3 the 2nd-highest-ranked open-weight model on real-world design tasks. Congratulations to the team on the launch!
28
58
644
165,443
Intelligence retweeted
BREAKING: Seedance 2.5 by @BytePlusGlobal ranks 1st in Multi-Image to Video on Design Arena with an Elo of 1400. Open-weight MiniMax H3 by @MiniMax_AI takes 2nd with an Elo of 1355, the top open source video model. Unlike standard image-to-video, this leaderboard gives models multiple reference images and compares their final video outputs. This tests how well models combine different visual inputs while preserving subjects, styles, and scene continuity - a critical capability for controlled video creation. Congratulations to both teams!
16
16
235
27,156
Intelligence retweeted
BREAKING: Grok 4.6 by @SpaceXAI places 7th on Design Arena with an Elo of 1335. This is a 6 position improvement from Grok 4.5, released only one month earlier. Grok 4.6 is a 1.5T parameter model that builds on Grok 4.5 by bringing substantial improvements across reasoning and coding capabilities. Congrats to the @SpaceXAI team on the improvements!
8
7
120
7,995
Intelligence retweeted
Grok 4.6 by @SpaceXAI is 4th in 3D Design on Design Arena with an Elo of 1370. This puts it ahead of Claude Fable 5 by @AnthropicAI and behind Qwen 3.8 Max by @AlibabaGroup. Grok 4.6 is an 11 rank and 60 Elo jump from SpaceXAI’s previous model, Grok 4.5, and generates designs 62.6% faster. Congrats to the @SpaceXAI team on this achievement!
6
11
159
16,362
Thanks for having us on the show, @tbpn!
Intelligence cofounder @grx_xce wants to build an intelligence marketplace for AI capabilities, not just a leaderboard that tells you which model is best. Today, Design Arena users submit creative tasks and have multiple AI models attempt it. The company then collects users' preference about which result was best, which gives Intelligence data about which models perform well for particular kinds of work. Grace thinks that in the future, it will become way easier for companies and individuals to customize models for one narrow purpose and sell access to it on a marketplace. She explains:
18
3,184
Intelligence retweeted
Another major milestone for open weights!
MiniMax H3 by @MiniMax_AI is 2nd overall on Video Arena with an Elo of 1325. This is a 209 Elo increase from @MiniMax_AI’s previous video model, MiniMax Hailuo-2.3 (Pro), putting them behind Gemini Omni Flash by @GoogleDeepMind and ahead of Seedance 2.0 Mini by @BytePlusGlobal. With this performance, they establish themselves as the 2nd video lab overall. Among open weights, the model is 1st overall, ahead of LTX 2.3 by @Lightricks and Kandinsky 5.0 Pro by AI-Forever. By a substantial gap, MiniMax has set a new SOTA on open weight video generation. Congratulations to the @MiniMax_AI team on the achievement!
1
1
12
1,841
Intelligence retweeted
WOW shoutout to @jordihays and @johncoogan for some incredible questions It was a pleasure to be on the show!
Intelligence cofounder @grx_xce wants to build an intelligence marketplace for AI capabilities, not just a leaderboard that tells you which model is best. Today, Design Arena users submit creative tasks and have multiple AI models attempt it. The company then collects users' preference about which result was best, which gives Intelligence data about which models perform well for particular kinds of work. Grace thinks that in the future, it will become way easier for companies and individuals to customize models for one narrow purpose and sell access to it on a marketplace. She explains:
12
6
187
27,688
Intelligence retweeted
BREAKING: Bland Speech v3 by @usebland debuts as the most human-sounding TTS model on our newest evaluation: Audio Realism Bench. Traditional speech benchmarks increasingly struggle to separate frontier systems that all sound clean, fluent, and intelligible. Audio Realism Bench instead asks a harder question: when two voices read the same transcript, which one sounds more like a real person? Vetted native speakers evaluate recordings blindly across phone agents, conversational assistants, and spoken explainers, judging the subtle qualities that define human-like speech. We also place real human recordings directly into the evaluation pool, testing whether evaluators can distinguish TTS models from actual human speech. Bland Speech v3 takes the top spot in our initial rankings, followed by MAI-Voice-2 by @MicrosoftAI and Grok TTS by @SpaceXAI, all demonstrating important steps toward voice systems capable of sustaining genuinely human-feeling interactions. Congratulations to the @usebland team!
14
31
154
29,618
Intelligence retweeted
MiniMax H3 by @MiniMax_AI is 2nd overall on Video Arena with an Elo of 1325. This is a 209 Elo increase from @MiniMax_AI’s previous video model, MiniMax Hailuo-2.3 (Pro), putting them behind Gemini Omni Flash by @GoogleDeepMind and ahead of Seedance 2.0 Mini by @BytePlusGlobal. With this performance, they establish themselves as the 2nd video lab overall. Among open weights, the model is 1st overall, ahead of LTX 2.3 by @Lightricks and Kandinsky 5.0 Pro by AI-Forever. By a substantial gap, MiniMax has set a new SOTA on open weight video generation. Congratulations to the @MiniMax_AI team on the achievement!
11
12
120
15,002