Last year, Anthropic was optimizing inference across 3 different types of compute (Nvidia, AWS trainium, Google tpus). There were mistakes made during these optimizations that resulted in actual degradation. Initially Anthropic denied it, but eventually they found, fixed and explained what happened. There have been no notable instances of degradation since.
Sadly, this one instance has made half of tech twitter’s brains fall out.
Separately, there’s the statistics side. You’re more likely to see the stupid spikes over longer windows.
If a model has a 1/50 chance of doing weird shit, and you do 20 prompts a day, there’s a ~30% chance you’ll have encountered weird shit on day 1.
By day 5, it’s closer to a 90% chance 🙃
Tl;dr, people are stupid, don’t understand non-determinism, and it happened once so they feel righteous