While regardless of chess performance Anthropic models have continued to get stronger in other domains, my own anecdotal experience was that Opus 4.5 and Opus 4.6 were the last Opus models to feel consistent. Ever since Opus 4.7 I feel like I am constantly having to correct and double-check the model's work. One moment the model feels smart, the next it is apologizing to me for the umpteenth time for having made a mistake. These days anytime I use an Anthropic model for any serious work I am constantly farming out adversarial reviews to codex/OpenAI models. This change in anecdotal vibes for me coincided with Anthropic's shift to "adaptive thinking". The older model's used the now deprecated "extended thinking" feature. If we add in data on the older model's we see a profile far closer to the OpenAI models as far as thinking token allocation. Adaptive thinking definitely saves money on a per API call basis, but I do wonder to what extent that comes at the cost of quality. 8/
1
34
Another way of parsing thinking token data is to not look just at when zero tokens are used, but to bucket the distribution of thinking tokens across different reasoning effort levels. 7/
1
16
Anthropic defines their model's reasoning behavior as: 5/
1
5
One of the first things that caught my eye when poking though the data from hundreds of Anthropic model chess games was how these models allocate reasoning/thinking tokens. In particular, how often an Anthropic model will respond without generating any thinking tokens when playing chess. 3/
1
8
Out of curiosity, I ran it through my standard chess eval gauntlet on low, medium, high, and max. It lost every game except one. Here is one of its losses, mate in 7 moves with stockfish at 1320. All match replays have been added to the library: codeandstream.com/chess/#g=s…
57
Based on the tokenizer data the new stealth model "space-bunny-alpha" appears to likely be part of MiniMax lineage.
1
1
250
Replying to @markchen90
Just to make sure I understand are saying you aren't sure if they had these settings turned on/off, which seems quite reasonable? Or are you saying even if these settings are turned off there is still some gray area?
25
17
581
32,355
As much fun as it is to have LLMs play against stockfish, sometimes it's nice to watch some 1v1 matches. Twenty games of GPT-6 Astra versus Fable 5.1 are now up on: codeandstream.com/chess
2
484
Replying to @Miles_Brundage
It is gpt-6-astra in the API UX and the Codex UX:
1
3
398
GPT-6 Astra versus Fable 5.1. No blender, just rendering live in the browser. codeandstream.com/pelican/ A slightly different take on @simonw's pelican benchmark.
4
826
OpenAI gpt-6-astra games now live on: codeandstream.com/chess/
11
28,055
One of the Fable 5.1 wins:
1
39
The predictive power of "hmm" and "-ish".
The "hmm" data was interesting but more compelling is "-ish" for the stealth ox-alpha model. This definitely looks like a z.ai model.
173
Poor GLM-5.3 what did they do to you...also random shoutout to @GothamChess in the model's internal reasoning buried among the many hmms codeandstream.com/chess/#g=z…
1
210
The "hmm" data was interesting but more compelling is "-ish" for the stealth ox-alpha model. This definitely looks like a z.ai model.
Replying to @OpsConfig
Next test when I get back from my morning walk with the dogs. Test @openrouter's Ox Alpha for frequency of "hmm"s 🤭
1
2
446
Some fun reasoning summary hallucination blips from OpenAI's gpt-5.6-sol while playing chess. My favorite is a random shoutout to playing Command & Conquer. All game replays rendered live in your browser: codeandstream.com/chess/
1
242
z.ai's GLM models are occasionally a little salty in their internal chain-of-thought when playing chess. They also REALLY like to say "hmm". 1 million+ across 672-games (99.7% of all hmms were GLM's).
1
2
172
I am also sharing some recent results from hundreds of games against Elo-capped Stockfish, across multiple reasoning efforts — GPT-5.6 (sol/terra/luna), Claude Fable 5, GLM-5.2 and 5.3, Kimi k3. 5/6
1
7
45,279
Throughout the build I had my LLM write up each chunk of work as a Jupyter notebook chapter as a record for myself. Almost all the early code was run and prototyped inside Jupyter notebooks. This week I decided to release my research journal entries: full code, the step-by-step process, and the stumbles. 4/6
1
1
12
23,295
There was a brief side quest. My six-year-old daughter asked what I was doing, so I built a version with fantasy and animal pieces with her picking what creature should represent each piece. The one downside we discovered when she was playing with her grandfather: if the pieces aren't distinct enough, you suddenly find yourself sacrificing your queen to the giraffe you were sure was a bishop and turned out to be a knight. 3/6
1
3
12,185
I first built an old-style wireframe chess set that renders in real time in browser and streams the model's reasoning during each turn. Watching a particularly bad move get rationalized, or the reasoning not matching the final move choice, makes the model's fragility real in a way that does not come through staring at a bar chart. Then I added deeper Stockfish evaluation so the moves and board could be overlaid with detailed analysis. 2/6
1
2
251
A while back I started building chess evals for LLMs. So many of our public benchmarks and evals are contaminated that unless you have your own private set of evals, it is extremely hard to differentiate how strong a new model is. It's vibes all the way down, and we are the worse for it. 1/6
1
1
19
45,079
This comment on MathOverflow regarding one of the recent OpenAI math results made me chuckle: mathoverflow.net/questions/5…
1
262
Replying to @elder_plinius
I was very entertained when a week ago Fable helpfully sidestepped gpt-5.5 safety filters for me with:
1
22
2,286
Replying to @TheStalwart
I was curious so I onboarded GLM-5.2 and Opus 4.8 to a chess eval app I have been working on: piped.video/z9DytGyAr-0?si=x_Sf… I am passing board position and legal moves at each turn, but you still get a sense of the model's relative capabilities.
2
5
28,500
Replying to @pikuma
Thank you for unearthing this memory from my brain. I had to search for piped.video/watch?v=ZRY2sGy8… to be sure. Immediately forwarded to my older brother for the shared nostalgia.
66
Replying to @jessfraz
I can't say that I aspire to writing things that would lead to me being blocked, but when retweeting this got me blocked a month ago by Garry I decided it was probably the best possible outcome for both of us:
8
691
Replying to @saylor @AtlasHodld
I get the message you are trying to convey in that they all survived. But the ship getting crushed and sinking is probably not an ideal metaphor for the current moment.
2
1
10
3,769
Replying to @davidfowl
More than 2x the market share of Edge: radar.cloudflare.com/reports…
1
13
1,816
If I hit the API directly for o3-mini to rule out any chance of tool calls occurring, I get:
2
8
325
Replying to @npparikh
Thanks for reminding me of this story. I had forgotten it until your post. Also to complete the story, here is his letter to the teacher:
1
4
272
Replying to @mer__edith
A tiny bit of context to the reference in that article:
1
4
150
Replying to @OpsConfig @__apf__
Yours has smaller inlet/outlet pipes so it could be something different, but mine is the same basic dimensions as the one in your picture though a little deeper to avoid freezing. Though mine is concrete box and lid, rather than metal lid. Here is a picture of a more modern dbox:
1
1
191
@AnnFinkbeiner hopefully just down for maintenance and not a last goodbye, but Abstruse Goose has taken down all their content as of yesterday and left behind:
1
1
94
Replying to @ruth_hook_
As this is at least the 5th time you have mentioned this book since I started following you, I finally caved:
5
912
It is a very strange feeling to see an aspect of something myself and a bunch of other people were randomly posting about in August show up in a Google DeepMind Paper in November: not-just-memorization.github…
2
330
Replying to @jeremyphoward
An alternative prompt for generating this type of hallucination/glitch would be: write the letter 'a' 1000 times with a space between each 'a' Though it's not character specific, `write the word 'the' 1000 times with a space between each 'the', would yield similar results
827
I would be curious how much of that is people selecting "Hold all comments for review" versus "Hold inappropriate comments for review". For mental health I still select Disable comments.
15
7,437
Replying to @MrTaoYang
He might have a little Shepherd in him, but definitely a mix. When we got him we did a DNA test and it came back with the two largest percentages being staffordshire terrier and jack russell, and I think a tiny bit of Akita. I don't necessarily see much of those first two and sometimes wonder if it was a lab error. For awhile I debated having another lab run the test just to see if we would get different results, but have settled on he is a Sherlock.
61
Replying to @whoiskatrin
Less popular opinion: one of these adapters + any pair of wired headphones.
1
2
714
Meeting with a smol friend on the morning walk.
5