M3.1-Flash Max is running at around 116 t/s, and the agents now feel a lot more like the ones in Grok Bot, It’s been running my new benchmark for more than 50 minutes, and so far it genuinely feels like it’s performing at a frontier model level, My first impressions: it’s extremely fast, you can message the agents while they’re working without any issues, which is basically impossible in some other harnesses, and it also has a cloud mode that runs everything remotely without using your local memory, Really impressed so far, I’ll post the full results soon, including a complete comparison with other models
Sep 27, 2026 · 7:29 AM UTC
5
1
39
2,734





