Everyone! This is a straw-person argument. The Stochastic Parrot paper was about LLMs of 2021, not the AI of today, which are not LLMs but complex software systems with vast post training and many external software components.
So the point of the stochastic parrot argument is that LLMs have zero understanding of language or anything else, they are just regurgitating training data. This hypothesis has been thoroughly disproven by LLMs solving millennium problems our smartest mathematicians have failed.
149
62
564
287,949
Don't some of the authors still apply it to models today? I think it's that continued application that people are objecting to.
2
197
12,311
I've seen them talking about it in the context of *LLMs*. What we have now are not LLMs. But the fact that many people still use these term to describe today's systems makes things confusing, I agree.
20
3
67
62,464
Do you really think they'd sign onto the statement "Current models are not pure LLMs and so the criticisms we made in this paper no longer apply to them"? I really don't think so based on what I've seen but open to being wrong.
1
1
159
7,115
See
In case helpful, I think Emily Bender and I have done the opposite — explain how it refers to LLMs. I wrote this: medium.com/@margarmitchell/n… And she wrote this: medium.com/@emilymenonbender…
2
6
10,848
Replying to @MelMitchell1
Bender's FAQ here mostly leaves me thinking that current models are still very limited for the basic reasons she'd outlined in the paper. She specifically says the framing is "extremely relevant" to current models.

Sep 26, 2026 · 5:16 PM UTC

8
1
143
23,198
Sort replies: Relevant Recent Liked
Let's be honest. This is the only definition ever offered for "stochastic parrots" in the paper. In other words, no definition at all. I've seen no disavowal by any of the authors of this awful rhetorical device that has done more harm to the understanding of LLMs by academics. I still see this cited by people as "evidence" that LLMs "don't really think" and therefore could not be used for "research". This was a bad paper written full of bad claims even for its time. I must admit I fell for it then but Google were right to not want to be associated with it.
1
21
1,115
Let's see if these questions get an answer
Say I take a transformer and… 1. Pretrain it on a bunch of text 2. Teach it special tokens for calling tools 3. Teach it how to use the command line 4. Teach it to use CoT to do longer reasoning and planning 5. Teach it how to engage in multi-turn conversation
7
822
My impression from the two articles is that Margaret Mitchell and Emily Bender have fairly different overall views, but will still generally state agreement with each other. I think you’d have an easier time getting Mitchell to agree to some version of your statement than Bender.
1
3
354
what is the claim supposed to be here? does anyone dispute the fact that current models are stochastic generative models trained to broadly reproduce their training data or that this is a significant factor in determining their capabilities and hence shortcomings?
55
Doesn't seem they think they are limited. SP is a description, not prescriptive, in order to make people understand what they are. The principles of cybernetics (the superior 80s version of algo and ai safety) still applies.
2
791
Reading the FAQ, I think you're right. Bender calls the framing 'extremely relevant' to current models and to systems built with them, so the 'these aren't LLMs anymore' defense doesn't really hold.
1
441
It's funny because their best OG point -- that the datasets used to train 2021-era LLMs are too vast and inscrutable -- is actually very relevant for those worried about AI safety! We should be aggressively auditing training sets to promote alignment and reduce biorisk!
1
494