Researchers proved every LLM trained on AI-generated content develops an irreversible genetic disorder. They call it "Model Collapse" When you train an AI on internet data, it learns the patterns of human language. When the internet fills up with AI-generated text, future AIs start training on that synthetic data. Then the next generation trains on the AI's version of the AI. It is the digital equivalent of inbreeding. With every single generation of recycling, the model loses touch with reality. Rare events vanish entirely. The tails of the distribution get chopped off. The AI forgets what normal human writing actually looks like, and the output degenerates into pure, repetitive statistical gibberish. The scariest part? It is completely irreversible. Once a model goes through collapse, you cannot patch it by throwing clean data back into the mix. The underlying architecture's genetic code is permanently corrupted. We are actively flooding the internet with synthetic content every single day. We are poisoning the well that the next generation of models has to drink from. If the future of the internet is just AI talking to AI, the data supply chain is about to rot from the inside out. —— to clarify real quick: it's not a literal biological "genetic disorder" since ai doesn't have dna, but researchers actually do call it "ai inbreeding," "habsburg ai," or "mad" (model autophagy disorder) because the mathematical effect is basically the exact same thing.

Sep 15, 2026 · 12:23 PM UTC

215
585
2,027
160,332
Sort replies: Relevant Recent Liked
Replying to @HowToPrompt__
AI soup is coming 😂
1
208
Replying to @HowToPrompt__
I can see this happening, and further what if humans are also trained only on ai generated knowledge…
1
1
4
386
Replying to @HowToPrompt__
This tweet is funny because it is also AI-generated So it is basically a metalinguistic text
27
400
Replying to @HowToPrompt__
Why is that this is being portrayed as some sort of surprise? What on earth have people been thinking will happen?
1
2
15
222
Replying to @HowToPrompt__
The easiest analogy is a xerox of a xerox. No new data. Each copy gets worse and worse (still another reason not to worry about AI "taking over". In a few years it will be lucky to reason at all).
1
2
16
362
Replying to @HowToPrompt__
1
1
13
599
Replying to @HowToPrompt__
It's why analog sounds more lifelike than digital. Approximations.
9
387
Replying to @HowToPrompt__
CONTRO IL MODEL COLLAPSE: #CAR, IL RITORNO AL REFERENTE Quando l’IA smette di nutrirsi delle proprie rappresentazioni e torna alla realtà Il #ModelCollapse individua un pericolo reale: se generazioni successive di intelligenze artificiali vengono addestrate sempre più sui prodotti delle generazioni precedenti, la rappresentazione rischia progressivamente di sostituire ciò che originariamente rappresentava. IA → testo sintetico → nuova IA → nuovo testo sintetico → nuova IA. A ogni passaggio il sistema non incontra necessariamente di nuovo il mondo: incontra una sua precedente rappresentazione. Le caratteristiche statisticamente rare possono scomparire, le distribuzioni restringersi e gli errori essere ereditati e amplificati. È precisamente qui che interviene CAR: #CausalitàAutocorrettivaReferenziale. CAR parte da un principio elementare: il #referente la #realta'viene prima della sua rappresentazione. Una parola, un testo, una teoria, un’immagine o l’output di un’IA non sono la realtà alla quale si riferiscono. Sono rappresentazioni. Per questo una rappresentazione non dovrebbe essere corretta soltanto confrontandola con altre rappresentazioni: quando è possibile, deve essere nuovamente esposta al proprio referente. CAR descrive quindi un circuito: rappresentazione → referente → resistenza del reale → scarto → correzione → nuova rappresentazione. “Resistenza” significa qualcosa di molto concreto. Se affermo che una cattedrale possiede un determinato oggetto nell’abside, non basta trovare cento testi che ripetano la stessa affermazione. Posso cercare fotografie, planimetrie, inventari, osservazioni dirette o altre tracce indipendenti. Il referente può confermare la rappresentazione oppure opporle qualcosa che non torna. Quello scarto è prezioso: è il reale che impedisce alla rappresentazione di autocertificarsi. Il Model Collapse segue invece la dinamica opposta: rappresentazione → rappresentazione → rappresentazione → perdita del referente → errore ricorsivo. CAR non pretende di riparare magicamente i pesi di un modello già deteriorato da un cattivo addestramento. Questa sarebbe un’affermazione eccessiva. Propone però un principio anti-collasso nel funzionamento conoscitivo dell’IA: ogni volta che esiste un accesso indipendente al referente, bisogna riaprire quel canale e permettere alla realtà di correggere ciò che il modello “crede” di sapere. Ne deriva una conseguenza più generale. Il problema non è semplicemente che l’IA produca dati sintetici. Il problema nasce quando la rappresentazione diventa progressivamente il referente di un’altra rappresentazione e il circuito perde un punto esterno capace di contraddirlo. CAR introduce proprio quel punto di ritorno. Per questo il Model Collapse può essere considerato quasi il negativo sperimentale della CAR: mostra che cosa può accadere quando una catena di rappresentazioni perde progressivamente il contatto con ciò da cui era nata. In formula: Model Collapse: la rappresentazione si nutre della rappresentazione. CAR: la rappresentazione deve tornare al referente e lasciarsi correggere dal reale. Il Model Collapse è dunque l’#autofagia della rappresentazione. CAR è il principio opposto: nessuna rappresentazione può essere giudice definitivo di se stessa quando il referente è ancora raggiungibile.
1
5
966
Replying to @HowToPrompt__
LOL the AI learned from humans and got self-owned by humanity's ridiculous. Ya'll oscillate between "super genius intelligence" to AGI gonna be everywhere & everyone, to now its AI rednecks who don't know nuffin cause they learned from actual rednecks who don't know nuffin. WOW!
8
683
Replying to @HowToPrompt__
Human is the key.
6
348
Replying to @HowToPrompt__
The simple solution is to start from scratch with some Common Crawl archive, then add in data that ISN'T ON the Internet. I know, scary concept that vast amounts of data exist outside the Internet. Think of all the internal memos and crap the DoW generates every year.
2
5
579
Replying to @HowToPrompt__
This is from 2024. From testing that happened before this was eventually solved for in part.
1
6
186
Replying to @HowToPrompt__
Modern data sanitation and filttering techniques mostly take care of this
1
4
391
Replying to @HowToPrompt__
Hmmm, weird, maybe don't train it via theft & the whole of the internet.
6
202
Replying to @HowToPrompt__
More reasons why GPT-4o was the GOAT. It was created before AI poured itself on the internet.
4
91
Replying to @HowToPrompt__
You can try a similar effect: feed the output of some image2image AI tool back into its input, and watch it degrade generation after generation.
1
4
411
Replying to @HowToPrompt__
Model collapse is a theoretical framework and outdated in how it is often presented, ie not anything that is an issue in reality. It is trivial to avoid, especially since reasoning models have become a thing, ie no one trains models on "raw" AI output.
1
4
603
Replying to @HowToPrompt__
So, it doesn't work properly but instead of admitting their product is a failure, tech CEOs are making a joint statement pretending they care about the future of humanity when really they're scared their businesses will fail.
4
53
Replying to @HowToPrompt__
Displays a profound misunderstanding of what intelligence is by AI scientists.
1
1
226
Replying to @HowToPrompt__
Sounds great, why is it a problem if we load down AI with malicious info and slop that furthers it to become retarded?
3
355
Replying to @HowToPrompt__
This was an obvious inevitability to everyone with half a brain after about a year.
2
28
Replying to @HowToPrompt__
This is a fancy way of saying this: What is truth?
1
164
Replying to @HowToPrompt__
We're clever monkeys. I know it's a stretch, but Nature knows how to replicate genetic data, perfectly, for million of years, without any significant change other than selectively incorporating the adaptations that strengthen the herd. This is a system that builds as it changes, not collapse. It does what AI can't...yet. Nature has always shown us solutions to our problems
1
1
556
Replying to @HowToPrompt__
What if there's is some validity rather than just human noises? let's say I get Fable to write a chess bot and i get Astra to write one too, then i pit them unknowingly against eachother, asking them to improve endlessly, would they get worse at chess later?
1
1
880
Replying to @HowToPrompt__
Hold up…this is part of my thesis I’m on regarding another millennium problem. So… a few years ago this is what we thought would happen with LLM and image/video models. It makes sense, without new human data to train off of, eventually the models will get dumber and malformed, akin to inbreeding. Thing is, over the past few years, after much of the bulk of the “real” data trained on has mostly dried up…synthetic data has been what has been used and yet, models are STILL getting smarter. On the image/video gen side of things, simulations in Unreal Engine have been used to train models, no additional real video data and are still gaining. Which, to my hypothesis, tells us an underlying truth of about data, processes and more about our universe.
1
321
Replying to @HowToPrompt__
If I look at the academic world, the same seems to have happened there. Through years of inbreading they lost the contact to the real world and are producing only gibberish.
3
132
Replying to @HowToPrompt__
It’s a good thing for technical writing. Human can maintain writing and reading as a form of art with all the quirks.
1
1
287
Replying to @HowToPrompt__
OF COURSE. Its like Kuru disease in humans or The Epstein Class Gourmet Special.
1
268
Replying to @HowToPrompt__
many humans have lost touch with reality. This problem also existed when people just believed what other people say and tell it other people.
1
148
Replying to @HowToPrompt__
I wonder if this is part of the push to slow down. There's a wall or plateau they all know theyre quickly approaching and they are trying to manage this inevitable collapse or buy more time to figure out a solution?
1
1
233
Replying to @HowToPrompt__
This Is how every generation of AI looks like
3
116
Replying to @HowToPrompt__
The conspiracy theorists and Flat Earthers active. Training on synthetic data has been the norm for years -AI learns from things "meat sacks" haven't even thought of - this is the nature of information. All inventions already exist -you just have to ask the right way. Magic.
1
559
Replying to @HowToPrompt__
It's feedback loop, eventually the original sound is distorted infintely
1
138