Competing summary-judgment motions in the New York Times’ case against OpenAI and Microsoft ask Judge Sidney Stein to decide whether training on copyrighted work is fair use. Almost every AI copyright ruling so far has been at the pleading stage, deciding only whether a claim may proceed; this would decide the question. Across 105 classifiable US filings — 89 litigation families — 77 per cent sit in two courts, and the count has gone 8, 15, 32, 47 across four years. What is inside them is moving: of 47 filings dated 2026, 33 allege training, 28 raise DMCA theories, and 14 allege the material was unlawfully acquired, against six of 32 a year earlier. Courts have started separating three questions — whether training is fair use, whether the material was lawfully obtained, and whether outputs infringe. Two pleading-stage rulings show the other pattern: core copyright claims surviving while DMCA theories fail. Cohere lost its motion to dismiss entirely, on 75 cited examples of alleged copying.
The Whole Story
Every generative model is built out of work somebody else made. Since January 2023, when visual artists sued Stability AI, Midjourney and DeviantArt in California, the people who made it have been arguing in court that assembling a training corpus is copying at industrial scale. Getty Images followed weeks later over more than twelve million photographs, and The New York Times sued OpenAI and Microsoft that December, asking not only for damages but for the destruction of models built from its articles. The defence has been the same everywhere and rests on an old idea: that learning from a work is not copying it, and that every major copyright system already carries an exception broad enough to cover the reading. What the courts have found is that the exceptions are not the same in any two countries, and that the argument turns less on what a model does than on where the case is brought.
American law has moved, but not where either side expected. The first merits ruling on fair use in an AI case, in February 2025, went against the AI company: Judge Stephanos Bibas held that ROSS Intelligence infringed by copying 2,243 Westlaw headnotes, and that a market for AI training data was itself something copyright protects — though ROSS's system was not generative, which limits how far the holding reaches. Then in June 2025 Judge William Alsup drew the line that still governs practice: training a model on lawfully acquired books is "exceedingly transformative" fair use, and so is digitising books you bought, but keeping a permanent library of pirated copies to do it is not. Two days later Judge Vince Chhabria granted Meta summary judgment on almost identical facts — reluctantly, "on this record" only, and faulting the thirteen authors who sued for never building a market-harm case. He wrote that he expected the opposite result in most cases properly argued. So the operative American rule is about provenance rather than training: no court has held that unlicensed training is generally unlawful, and the clearest thing anyone has won is that you may not torrent the corpus. That rule is now being tested head-on. In September 2026 OpenAI told the Manhattan court hearing the consolidated New York Times and Authors Guild cases that compilations from Library Genesis had gone into GPT-3 and GPT-3.5 — and asked it to hold that this makes no difference, because where the ultimate purpose is transformative the downloading in service of it is transformative too. That is the position Chhabria took in Meta's case and Alsup rejected in Anthropic's, and the two answers cannot both survive.
That is why the largest number in the field attaches to piracy rather than to training. Anthropic settled the authors' class action in September 2025 for about $1.5 billion — the largest copyright recovery on record — covering an estimated 500,000 works at roughly $3,000 each and requiring the downloaded files be destroyed, after filings showed it had taken more than seven million books it knew to be pirated from Books3, LibGen and the Pirate Library Mirror before switching to buying and scanning physical copies. A successor judge approved the deal in July 2026, noting the per-work figure was four times the minimum statutory damages for willful infringement and that 92% of eligible authors had opted in. It bought no peace: music publishers pressing a separate fight over song lyrics filed an amended complaint the following day built on what discovery had turned up, and are seeking up to $150,000 a song. Where the law is unsettled, private ordering fills in — Microsoft began indemnifying its Copilot customers against copyright claims in September 2023, promising to defend them and pay any judgment.
Outside the United States the same conduct meets different exceptions and gets different answers. In November 2025 the High Court in London largely cleared Stability AI, holding that a model which does not store the works is not an "infringing copy" — but Getty had already withdrawn its central training claim mid-trial because the copying happened outside the UK, and the court found for Getty on some watermark claims. A week later the Regional Court of Munich I went the other way in GEMA's suit against OpenAI, holding that Germany's text-and-data-mining exception does not apply because models permanently memorise rather than transiently analyse, that storage inside a model is reproduction, and that a model reciting lyrics on request is communicating them to the public; OpenAI is appealing. Then in July 2026 a Delhi judge refused to stop OpenAI at all, reading India's exception for "private or personal use, including research" to cover a closed corporate training process, declining to import the American four-factor test, and taking jurisdiction even though the servers sit abroad — an interim finding on a prima facie standard, not a final one.
Behind the litigation, the institution that would ordinarily settle American copyright policy has been fighting over its own leadership. The US Copyright Office ran a three-part study across 2024 and 2025 — on digital replicas, on whether AI output can be copyrighted, and finally on training itself, which floated a theory that AI output can dilute the market for originals even where nothing is copied. That third part was released in an unusual pre-publication form on 9 May 2025. The Register of Copyrights was fired the next day and sued the president a month later over it; a district judge first refused her emergency relief, but an appeals court reinstated her, and in June 2026 the Supreme Court declined to remove her while the case proceeds — even though days earlier, in a separate case, it had expanded the president's power to fire officials of that kind, so the reasoning likeliest to decide her fate now runs against her. In September 2026 the executive that removed her went further: the United States filed a Statement of Interest in the consolidated New York Times litigation telling the court her Office's market-dilution theory was "deeply flawed" and that her understanding "does not warrant deference" — the first time the federal government has argued the merits of the training question in any court, and squarely on the side of the AI companies. So the questions stay open in every direction — but one of them is close to an answer. The image cases are heading to trial and an appeal is pending in Munich, while the New York Times and Authors Guild cases are now fully briefed for a fair-use ruling, with the United States arguing inside them for the AI companies and the two American judges who have looked hardest at training having reached opposite instincts about how it ends.