Claude Fable 5.1: 55.8% on Terminal-Bench 4.0.
Opus 5: 52.3%. Old Fable 5: 42.0%.
A ".1" release just jumped past Opus 5 on agentic coding. And it runs cheaper.
The quieter number is the bigger one. On Terminal-Bench-Science 0.1, Fable 5.1 scores 52.6%. Fable 5 scored 24.7%. It more than doubled its own science score in a single point release.
Same token price as Fable 5. $10 in, $50 out. Nothing moved on the sticker.
What moved is effort. It reaches the same answer at a lower effort level, so typical jobs cost about 25% less to run, and heavy agent jobs up to ~45% less. Anthropic's own line: "similar or better results at a much lower cost."
Proof it is not a benchmark mirage: a builder had Fable 5.1 one-shot a playable Mario Kart game. Not a snippet. The whole thing, first prompt.
If you run agents that ship real work, this is a free upgrade you have to opt into. Repoint your most expensive task at Fable 5.1 and measure cost per finished task, not per token.
The frontier didn't post a shiny new headline number this week. It got cheaper to reach.
Bookmark this before your next model pick.
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1.
They're the world’s most advanced models for coding and knowledge work.