A few thoughts on why @AnthropicAI's Opus 5 sometimes felt lazy compared to @OpenAI 's models. I like to use chess as a benchmark/eval for AI models. It is a crude/toy benchmark, but it does a good job of giving a sense of when the frontier model's make leaps forward and helps temper some of the benchmaxxing. 1/
1
115
While regardless of chess performance Anthropic models have continued to get stronger in other domains, my own anecdotal experience was that Opus 4.5 and Opus 4.6 were the last Opus models to feel consistent. Ever since Opus 4.7 I feel like I am constantly having to correct and double-check the model's work. One moment the model feels smart, the next it is apologizing to me for the umpteenth time for having made a mistake. These days anytime I use an Anthropic model for any serious work I am constantly farming out adversarial reviews to codex/OpenAI models. This change in anecdotal vibes for me coincided with Anthropic's shift to "adaptive thinking". The older model's used the now deprecated "extended thinking" feature. If we add in data on the older model's we see a profile far closer to the OpenAI models as far as thinking token allocation. Adaptive thinking definitely saves money on a per API call basis, but I do wonder to what extent that comes at the cost of quality. 8/
1
31
Comparing thinking tokens across models/providers is fraught, and of limited value. Different models, different backend implementations, different tokenizers, etc. Stronger models should be able to solve more complex problems using less thinking tokens/test time compute. The older pre-adaptive thinking Opus models are weaker so even with higher thinking token allocation their performance on chess doesn't improve enough to catchup to latest OpenAI models. With that said, I do think there is value in tracking reasoning/thinking token allocation to see how often low/zero thinking token allocation corresponds to poor performance.
14
Based on the tokenizer data the new stealth model "space-bunny-alpha" appears to likely be part of MiniMax lineage.
1
1
238
Out of curiosity, I ran it through my standard chess eval gauntlet on low, medium, high, and max. It lost every game except one. Here is one of its losses, mate in 7 moves with stockfish at 1320. All match replays have been added to the library: codeandstream.com/chess/#g=s…
54
OpsConfig retweeted
Annie Dillard has died. In our Summer 1988 Issue, she wrote about the "surprising letters" and gifts that fans and readers sent her over the years—a fossil fish, a marriage proposal, a check for $2.50, and, again and again, questions about God. yalereview.org/article/annie…
1
58
257
7,402
❤️‍🩹
7
15
161
5,594
OpsConfig retweeted
Two heroes of today's Navier-Stokes story are not getting enough credit. Córdoba and Martínez-Zoroa came up with the crucial strategy. Others built on it and, with the help of AI, provided the computations that finished the problem. Thankfully @QuantaMagazine gets this right:
21
338
1,701
89,794
As much fun as it is to have LLMs play against stockfish, sometimes it's nice to watch some 1v1 matches. Twenty games of GPT-6 Astra versus Fable 5.1 are now up on: codeandstream.com/chess
2
481
GPT-6 Astra versus Fable 5.1. No blender, just rendering live in the browser. codeandstream.com/pelican/ A slightly different take on @simonw's pelican benchmark.
4
823
OpenAI gpt-6-astra games now live on: codeandstream.com/chess/
11
28,054
One of the Fable 5.1 wins:
1
39