While regardless of chess performance Anthropic models have continued to get stronger in other domains, my own anecdotal experience was that Opus 4.5 and Opus 4.6 were the last Opus models to feel consistent.
Ever since Opus 4.7 I feel like I am constantly having to correct and double-check the model's work.
One moment the model feels smart, the next it is apologizing to me for the umpteenth time for having made a mistake.
These days anytime I use an Anthropic model for any serious work I am constantly farming out adversarial reviews to codex/OpenAI models.
This change in anecdotal vibes for me coincided with Anthropic's shift to "adaptive thinking".
The older model's used the now deprecated "extended thinking" feature. If we add in data on the older model's we see a profile far closer to the OpenAI models as far as thinking token allocation.
Adaptive thinking definitely saves money on a per API call basis, but I do wonder to what extent that comes at the cost of quality.
8/