Sonnet 5.5 vs Opus 5.5 High on my-mini-bench and comparison with Sonnet 5
I tested them on tasks from my own work: data conversion, frontend, refactoring and running task evaluations.
Setup: 10 tasks × 3 runs, High reasoning, in pi-agent and Claude Code through harbor.
Sonnet 5.5 vs Opus 5.5 in pi:
> Both passed 30/30.
> Time: 46s vs 103s: ~2.2× faster.
> Cost: $0.10 vs $0.30 per task: ~3× cheaper.
Compared with Sonnet 5 in pi:
> Passes: 25/30 → 30/30.
> Turns: 12.7 → 4.3
> Output tokens: 17.7k → 6.5k
The trajectories show a different approach to reading files, writing code, and testing. Examples below 👇
This benchmark is already saturated, but I use it to find the best setup for my tasks: cheaper, faster, and still gets the job done. It also lets me quickly compare models and see how their behavior differs.