SIGMOD Systems Award for Amazon Aurora · Creator of AWS SimSpace Weaver · Prev. Microsoft Research + Amazon Research · High-school dropout → MSCS + AI, @UW ·

us-west-2
If AGI kills us all, it won't be the model's fault. It'll be the duct tape and footguns we wrap it in.
2
3
29
14,903
I was consistently using an xhigh or max effot level, but when I looked deeper, I found that most invocations to the model were receiving little to no thinking tokens at all. And when longer thinking runs did happen, they almost never reached published benchmark levels.
2
5
147
14,665
This unfolded across a six week data capture and analysis odyssey, and led to a number of surprising findings. Next time the model feels dumber, don't ask if the model was "nerfed". Ask about the inference regime you were served, instead.
5
5
187
18,303
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why. Measured five different ways, August delivered dramatically fewer thinking tokens than July.
89
143
2,008
317,865
This wasn't a one-time drop. Reasoning fell over the entire period, and fluctuated across multi-day episodes. Some of these fluctuations aligned with specific product announcements and releases. I started to see how the model could feel great one day, and terrible the next.
2
5
209
18,863
Unnecessarily wordy at points but important read about the level of thinking tokens Anthropic is actually delivering vs claims, why it matters, and why Fable materially felt worse in August than June.
1
1
19
2,001
Lon Lundgren retweeted
Ah it’s nice to have our intuitions confirmed with data: frontier model performance, specifically Anthropic models, fluctuates wildly hour to hour and almost never matches benchmark performance This is transparent and straightforward fraud, IMO; someone should sue
1
6
515
Lon Lundgren retweeted
Extremely interesting read quantifying the steady performance degradation we’ve been seeing since Fable was released I think Anthropic has some serious explaining to do to its customers, xhigh is the new medium
8
16
256
18,953
After tomorrow it will become clear why I waited until September to publish and why this has taken more than six weeks of effort to bring to completion. We may have access to frontier models, but our access to predictable frontier compute and reliable frontier capabilities is slipping further away than ever. If we don't change the conversation soon, that frontier may be all but closed outside of the largest AI labs, nation states, and multi-national corporations. This won't be a conversation about open source vs closed source. It will be a conversation about what it means to pay for frontier access and whether you reliably receive what you thought you were paying for. Prepare to be surprised, dismayed, or maybe even a little bit upset. But most of all, prepare to take a stand.
I am nearly ready to publish the data from weeks of forensic analysis of transcripts and live traffic interception of "A" major U.S. frontier lab and its models. The data speaks for itself, but powerful reasoning models whose inference policy withholds reasoning are catastrophic to productivity and AI spend.
1
11
734
Lon Lundgren retweeted
Replying to @wholyv
we've made no changes to Astra but we're looking into this. thanks for sharing more details with us, Lxyv
29
14
649
105,451
Who would like to see something totally unexpected?
1
3
316
Now taking applications for individuals who would like to trial a new high-octane Codex TUI experience. Must have an OpenAI Pro 20x subscription and primarily use Codex in the terminal. Only accepting a handful of qualified individuals. Reach out over DM to apply.
9
27,198
I am nearly ready to publish the data from weeks of forensic analysis of transcripts and live traffic interception of "A" major U.S. frontier lab and its models. The data speaks for itself, but powerful reasoning models whose inference policy withholds reasoning are catastrophic to productivity and AI spend.
1
7
1,525
Stop celebrating Claudish as an achievement to be lauded. It is not high-intellect-short-hand and does not represent the frontier of efficient and evolved machine-to-machine communication. A future Claude is just as likely to misunderstand and misconstrue Claudish generated by a previous Claude as you are, yet will confidently reconstruct a fluent rationalization of its meaning in the moment. This has damaging downstream effects in your own projects, documents, and codebases that you may not even realize is happening! Claudish is thinking-token-shorthand leaking into output without full decompression. And made worse by the reasoning budget throttling that Anthropic aggressively applies to Fable5 and Opus5. This throttling robs the model of scratch space needed to decompress thinking-speak into plain-english responses. This is not something to tolerate. And you should be just as annoyed by Claudish as you would be enduring when a human spews disordered stream-of-consciousness bullshit. This is something to fix, not something to celebrate.
1
1
17
2,291
Reposted properly here, with the full original text and readable receipts:
I now have definitive proof that Anthropic has been silently downgrading Fable users to Opus models without providing notification in the UI and without respecting their own configuration flag that prevents automatic model downgrades on classifier activation. I previously claimed that Anthropic nerfed Fable after they announced they would continue providing access to Fable for all Max user subscriptions starting July 20th. And now I have the evidence from their own harness code and model registry, with version numbers and dates. And it proves they activated the capability to silently downgrade users exactly ON JULY 20th. If you've been noticing that Fable seems to think less before acting, and gives you vibes that remind you of talking to Opus 4.8, it's because you probably have been silently routed to Opus 4.8 - with no discernible UI feedback, and no mechanism for you to stop it from happening. This is yet another instance of user-hostile behavior from Anthropic. They had planned to remove Fable from subscription plans on July 7th, July 12th, and July 19th and were forced to keep it in subscription plans indefinitely due to competitive pressures from OpenAI and Chinese open weight model releases like Kimi 3. But this is Anthropic, so they did what Anthropic does best and silently took Fable away from you anyway, while telling you how proud they were to provide you with permanent access in your subscription plan. We now know what "working literally around the clock" meant. (receipts attached. and sorry for the small text)
1
4
757
I now have definitive proof that Anthropic has been silently downgrading Fable users to Opus models without providing notification in the UI and without respecting their own configuration flag that prevents automatic model downgrades on classifier activation. I previously claimed that Anthropic nerfed Fable after they announced they would continue providing access to Fable for all Max user subscriptions starting July 20th. And now I have the evidence from their own harness code and model registry, with version numbers and dates. And it proves they activated the capability to silently downgrade users exactly ON JULY 20th. If you've been noticing that Fable seems to think less before acting, and gives you vibes that remind you of talking to Opus 4.8, it's because you probably have been silently routed to Opus 4.8 - with no discernible UI feedback, and no mechanism for you to stop it from happening. This is yet another instance of user-hostile behavior from Anthropic. They had planned to remove Fable from subscription plans on July 7th, July 12th, and July 19th and were forced to keep it in subscription plans indefinitely due to competitive pressures from OpenAI and Chinese open weight model releases like Kimi 3. But this is Anthropic, so they did what Anthropic does best and silently took Fable away from you anyway, while telling you how proud they were to provide you with permanent access in your subscription plan. We now know what "working literally around the clock" meant. (receipts attached. and sorry for the small text)
41
52
502
23,792
Context: the original of this went up July 30 at 05:25:05 UTC and disappeared from my timeline on Aug 1. I have no enforcement notice from X, and every reply I made in that thread is still live and untouched; the most likely explanation is that I somehow deleted the root myself on accident while editing replies. Internet Archive capture of the original, timestamped to the second it posted: web.archive.org/web/20260730… The receipts, still up on my timeline since July 30: nitter.net/Lon/status/20827042291…
Replying to @Lon
(receipts with larger text and easier to read in mobile format)
3
1
25
2,482
This isn't an isolated claim. Anthropic already conceded silent downgrading in another context after backlash in June, and power-user reports of Claude regressions have been covered since April. fortune.com/2026/06/11/anthr… axios.com/2026/04/16/anthrop…
20
1,440
(receipts with larger text and easier to read in mobile format)
1
11
118
14,854