Three years into the AI copyright fight, the headline question is still open: no court has settled whether training on copyrighted work is legal. But the rulings so far point somewhere more specific. Courts are not really deciding whether machines may learn from text. They are asking two narrower questions. Where did the material come from, and can anyone prove the output harmed the market for the original? Almost everything decided so far turns on one of those two.
Provenance first. On 20 July 2026 a federal court gave final approval to the $1.5 billion settlement in Bartz v. Anthropic, roughly $3,000 for each of about 500,000 books, plus an obligation to destroy the pirated files. The underlying ruling had split the question cleanly: training on lawfully acquired books was fair use, keeping a permanent library of pirated copies was not. Anthropic was penalised for how it got the material, not for training on it. The release covers acquisition and copying up to August 2025 and nothing else, so output claims survive.
That sequence has been said out loud. Speaking to an AI class at Stanford in 2024, the former Google chief executive Eric Schmidt described the play for an AI startup: tell the model to build your competitor and "steal all the music", get it in front of people, and "hire a whole bunch of lawyers to go clean the mess up" if it works. He asked for the video to be taken down and said later that he had not meant it literally. It remains a fair description of the incentive. Clearing rights up front costs money now, while a ruling years away is uncertain and discounted. Bartz is the first time that bill has been presented, and whether $1.5 billion is large enough to change the calculation is an open question.
Then harm, which is where most cases have actually turned. In Kadrey v. Meta the authors lost, and the judge said plainly that they lost because they had not proved market harm, not because training is lawful. He also flagged a dilution theory a better-argued case could win. Thomson Reuters v. Ross failed for the mirror-image reason: the copying produced a product that competed directly with the original. Harm is the hinge, and it is hard to evidence.
Neither question is settled at the level that counts: no appeal court has ruled on any of it. The first to hear argument was the Third Circuit on 11 June 2026, in Ross, and the panel spent most of its time on exactly these points, whether the use was transformative and whether it harmed the market, including a licensing market for training data. No decision has issued. NYT v. OpenAI has an order to hand over 20 million anonymised ChatGPT logs, summary judgment pending, and no trial date. Getty's UK case failed on its central copyright theory and is under appeal.
One filing from last month matters more to this audience than any of the above. On 14 August a group of textbook authors sued OpenAI, following a parallel suit against Meta in July, and their harm argument is built around how academic work is actually bought. Textbook adoption is an institutional decision rather than a consumer one, so a substitute does not have to be as good as the book. It only has to be good enough for the committee that chooses it. That is a far easier harm to demonstrate than a lost novel sale.
To summarise. Whether AI training is lawful in the abstract remains undecided, and will stay that way until an appeal court speaks. What the cases turn on is provenance and provable harm. For researchers, NGOs and archives that has a practical edge: clean records of what you hold and where it came from are becoming legally load-bearing. And every plaintiff named here is a large publisher or a well-organised class, not a small archive.
Next: the incentive trap.
#AIcopyright #ScholComm