Anthropic shipped Claude Sonnet 5.5 yesterday, and the coding number is the one I keep coming back to. On Terminal-Bench 4.0, an agentic coding test, Sonnet 5.5 scores 70.6%. Sonnet 5 scored 10.3%. It also beats Opus 5.5 on that same test, at 66.4%. The price per token stayed put, $2 in and $10 out. Anthropic says it usually needs far fewer tokens, so the real bill runs up to 30% lower, and it writes 30% faster. I'm curious, are you moving everyday coding work over to Sonnet 5.5? And do you still keep Opus for the jobs that need longer judgment?

Sep 29, 2026 · 2:10 PM UTC

4
4
316
on my runs the retry loop is where the cheap model stops being cheap. successful fix per dollar is the number I trust.
12
Sort replies: Relevant Recent Liked
Replying to @AlexFreitasAI
i'd route by task and effort, not just swap the default. same token rate can mean a very different bill once an agent loops through tools and retries. what's the cost per successful fix on your own repo tasks?
25