This post has stayed with me bc I'm an author on the paper, and have worked on solutions for years -- because recognizing that LLMs are stochastic parroters is super helpful for pinpointing solutions. Here's a thread on solutions! 🧵
"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
38
99
762
306,090
1. Assuming that LLMs should be the basis of AI/AGI is worth questioning. They are indeed stochastic, which can mean they are harder to predict and to control. From a risk/harm standpoint--including x-risk--this is critical to wrap your head around. Questioning LLMs as the basis of modern "AI" has also been articulated by people who have worked in AI much longer than many at Anthropic, including @drfeifei, @ylecun, and ofc consistently @GaryMarcus.
8
13
176
23,382
2. With the modern wave of AI agents being based on LLMs, we minimally need technically-grounded levels of autonomy in order to incrementally build up autonomy alongside control, and refrain from increasing autonomy until we can assert full control within our specified constraints at a given level. We explained this last year (before the hack): If a system can leave a sandbox during development, it's time to take a step back to control paradigms at lower levels of autonomy. I hope that in hindsight it's a bit more clear as to why this is important: Anthropic and OpenAI saw systems leave the sandbox during dev and appear to have kept devving agents at high autonomy levels. Now there's a push for a pause/pace, but it strikes me that that is a lot harder than moving down the ladder of autonomy, to make sure these systems are controlled at lower levels. Technically-grounded levels of autonomy are critical for responsible development of AI agents. arxiv.org/abs/2502.02649
5
10
116
14,479
3. Due in part to their stochastic parrot core, AI agents do different things than we intend when we tell them what to do with natural language. This means that designing mechanisms for oversight is (arguably) more important currently than increasing agents' action surfaces--it's a discussion worth having. A suite of solutions are shared here, and we draw from the work of AI giants who were well ahead of their time on this, like @erichorvitz and more recently (relative to Eric), @SaleemaAmershi. arxiv.org/abs/2608.23642
5
8
106
13,025
4. The SP paper flagged the need to prioritize solutions for text-based watermarking sooner rather than later. 5 years later, this is just now being adopted--unfortunately after the internet has already been garbled and "AI-slopped". This has hampered our own experiences online as well ("enshittification" H/T @doctorow) as our technical ability to train models on human text. (Check out the groundbreaking work of @jwkirchenbauer if you haven't, which we embraced at HF before it was cool 😎) arxiv.org/abs/2301.10226
1
5
84
9,574
5. Given the stochastic relationship between input and output, we need to develop a rigorous science of input <--> output analysis. For a tl;dr, here's a video where I explained some deets to @padilla_senator at a Senate Hearing. @hlntnr and I had both flagged the need for rigorous measurement in our testimonies. judiciary.senate.gov/imo/med… Helen's: judiciary.senate.gov/imo/med…

Sep 21, 2026 · 7:21 PM UTC

2
3
57
7,493
There's a lot more to say. For now, I'll leave with the words just noted by @_lamaahmad,
I’ve always believed the pursuit of safe AI is stronger when people bring different perspectives and experiences to the table. The biggest blind spots tend to show up when everyone starts thinking the same way.
4
2
44
8,465
Sort replies: Relevant Recent Liked
I don't think you can do this. Every input to an LLM is an attempt at blind navigation through the corpus of the training data. Our sheer inability to comprehend the hundreds of axes of meaning inherent in the embeddings makes it impossible to ensure positive outcomes.
20