Comparing thinking tokens across models/providers is fraught, and of limited value.
Different models, different backend implementations, different tokenizers, etc.
Stronger models should be able to solve more complex problems using less thinking tokens/test time compute.
The older pre-adaptive thinking Opus models are weaker so even with higher thinking token allocation their performance on chess doesn't improve enough to catchup to latest OpenAI models.
With that said, I do think there is value in tracking reasoning/thinking token allocation to see how often low/zero thinking token allocation corresponds to poor performance.