Cofounder @ProximalHQ | prev founding eng / research @Cursor_AI

San Francisco
Excited to release FrontierSWE v2! The team put a lot of good work into this and it's one of the coolest benchmarks I've seen
We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
1
1
20
1,063
Grok 4.5 is a really impressive model
Grok 4.5 ranks #2 on FrontierSWE It is only outperformed by Claude Fable 5 and ranks higher than Opus 4.8, GPT-5.5 and GLM-5.2
9
1,538
It's really beautiful how much of writing good software is understanding systems deeply and simplifying them into their core abstractions
6
522
Mostly red PRs > Mostly green PRs
4
535
Fable 5 is now the best model on FrontierSWE (by a large margin). In the training task, the solve rate of Fable is 67.8% compared to the next highest score of 3.8% for Opus 4.8.
Claude Fable 5 ranks #1 on FrontierSWE. This represents the biggest capability jump we have observed since releasing the benchmark On many tasks, Fable 5 works productively for close to 20 hours and fully saturates tasks that were effectively out of reach for earlier models
4
741
Navid Pour retweeted
We believe that better training data will come from creative research and engineering ideas, not from hiring annotators. Here are some of the open problems we are working on:
2
17
86
20,826
Navid Pour retweeted
how do i invest in Wang Ji Rice Dumplings
70
515
9,524
641,345
This is pretty cool! TLDR: Robot foundation models scale better by increasing data diversity than increasing model size
Do we finally have a "Chinchilla Moment" for robotics? 🤖📉 Turns out, you can't just treat the physical world like next-token prediction. New scaling laws show a massive twist: data diversity completely crushes raw model size for physical AI. 👇 open.substack.com/pub/paoloa…
2
20
6,206
Navid Pour retweeted
Opus 4.8 is indeed #1 on FrontierSWE
Anthropic says Opus 4.8 ranks 1st on FrontierSWE
7
18
474
31,732
We evaluated Opus 4.8 leading up to the launch. I personally prefer Opus models to GPT because they feel higher "taste" to me. My only issues with Opus models were token efficiency and covering edge cases better, which has been fixed in Opus 4.8. It's a seriously good model.
We evaluated Claude Opus 4.8 on FrontierSWE ahead of today's release. It is now the best-performing model on FrontierSWE.
1
23
1,815
Composer 2.5 used far less token than Kimi 2.6 on Frontier SWE (roughly 1/3 of the token usage). It's very token efficient and fast, but not on par in performance with OpenAI and Claude models on long horizon tasks. It gives up too early.
Composer 2.5 is ranked #5 on FrontierSWE The model is broadly on par with Gemini 3.1 Pro, with a slight edge in our evaluation, and it beats all open source models. We still observe a significant performance gap between Composer and models from Anthropic and OpenAI
5
769
GPT 5.5 feels like a very detail-oriented (and somehow fast) mid-level engineer. Doesn’t make mistakes, but lacks good engineering taste. Opus 4.7 feels like a senior engineer that is a little detached from the details of the code. Has decent engineering taste, but makes small mistakes while writing the code. It feels like maybe 5.5 is optimized to write code that is LLM exensible, while 4.7 is optimized for human readability and engineering taste.
3
21
3,635
Curious how model providers will make decisions as "human-taste" diverges from "model-taste" for code
1
478
Come hang out with us in North Beach!
Hosting a research meetup in our North Beach office on Thursday! Come by for food, drinks and talks: @jyangballin (MSL) will present ProgramBench @rawsh0 & @rishiiyer01 (Zyphra) will talk about ZAYA-8B @evan_j_chu and I will speak FrontierSWE and our research bets!
11
1,269
Navid Pour retweeted
We are hiring research fellows to help us improve FrontierSWE! If you want to help build the hardest real-world coding benchmark, reach out! Fellows can work with us for a few weeks up to months and will be supported with compute and a generous stipend
Introducing FrontierSWE, an ultra-long horizon coding benchmark. We test agents on some of the hardest technical tasks like optimizing a video rendering library or training a model to predict the quantum properties of molecules. Despite having 20 hours, they rarely succeed
5
18
309
41,831
GPT 5.5 goes brrrrrrr
GPT-5.5 is the best-performing model on FrontierSWE. The model substantially outperforms Opus 4.7 in both mean@5 and best@5 rankings while working faster.
1
8
504
DeepSeek V4 & Kimi K2.6 on Frontier SWE. It's interesting that this tracks very well with my personal experience. Anthropic + OpenAI are meaningfully different to use than Gemini and OS models (though their performance is basically on par on other benchmarks).
DeepSeek V4 Pro is the best open source model on FrontierSWE, closely followed by Kimi K2.6. V4 exhibits noticeably fewer reward hacking attempts than most other models. In the best@5 ranking it performs as well as Gemini 3.1 Pro
2
1
17
987