This is a perfect example of why production use of LLMs (local+frontier) should include automated benchmarks for the specific tasks relevant to an organization/team.
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why. Measured five different ways, August delivered dramatically fewer thinking tokens than July.

Sep 21, 2026 · 2:17 PM UTC

4
16
2,198
Sort replies: Relevant Recent Liked