Bon, une autre affaire? 🤷🏻😅
ChatGPT has now a big problem.
Researchers at Oxford and Cambridge exposed a massive threat to large language models.”
They call it “model collapse."
Internet ecosystem is rapidly changing, and generative AI will soon contribute much of the text found online. This forces us to consider what happens to future iterations like gpt-n when they are trained on data scraped from the web that was already generated by an llm.
According to the research, indiscriminately using model-generated content in training causes "irreversible defects" in the resulting ai. the model loses the "tails of the original content distribution." in other words, it begins to forget the creative, fringe, and unique nuances of actual human writing, collapsing into a repetitive echo chamber.
This isn't just a chatgpt issue.. the researchers built theoretical intuition showing this collapse is ubiquitous across learned generative models, occurring in large language models as well as in variational autoencoders and gaussian mixture models.
Tech companies rely on scraping the internet for large-scale data to build smarter models. However, the paper warns that if we want to sustain the benefits of training on web data, model collapse must be taken seriously.
Ultimate takeaway? data collected from genuine human interactions is going to become increasingly valuable in a web filled with ai content.