We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash.
Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task.
Here’s how all 6 models compared 🧵🧵🧵