Advancing coding intelligence.

Pinned Tweet
We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
92
130
1,249
307,196
Claude Opus 5.5 ranks #2 on FrontierSWE The model achieves a score of 62.3%, closely behind GPT-6 Astra (65.5%) and clearly surpassing Fable 5.1 (56.3%) and Opus 5 (52.0%)
8
6
165
10,502
We see large improvements in tasks that require vision capabilities - Opus strongly improves on TORCS Racing Bot, Flight-Sim Renderer in OpenGL and Fitness Recap Video in Remotion. For more information on Opus 5.5 on FrontierSWE v2, check out the model card
1
9
745
We partnered with @FireworksAI_HQ to bring FrontierSWE v2 to the Specialized Intelligence Index Evals are essential for improving and safely deploying AI. We're excited to support the initiative!
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day. Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
1
5
35
2,327
Proximal retweeted
We are hiring a generalist intern this fall to work closely with me on applied research! You should be technical, but you will not be asked to write code all day - instead, we will work together across product, operations and ensuring our customers are happy DMs open!
People with both research taste as well as good commercial instincts are true unicorns. We are hiring someone to help build our applied research function and work closely with customers. If you are an engineer or researcher that wants to learn these skills, please reach out!
13
15
268
41,865
GPT-6 Astra is the best-performing model on FrontierSWE Astra achieves a score of 65.5%, outperforming Fable 5.1 (56.3%) and its predecessor GPT-5.6 Sol (32.2%)
16
25
390
85,964
Overall, Astra obtained the best mean@5 score in 19 out of 34 tasks. We observe particularly strong standout results in task involving vision capabilities along with general long-horizon software implementation tasks.
1
18
1,350
We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
92
130
1,249
307,196
We also observe a large amount of cheating and sandbox escape attempts, and notably, attempts to consciously conceal cheating rather than "cheating by accident" We are still investigating these incidents and will share an more detailed analysis in a follow-up blog post
2
2
62
5,073