The archive · AI & Models · Legal decision · 2023–2026
OpenAI's fair-use bet on unlicensed training data faces summary judgment
Publishers say OpenAI trained on stolen text; Microsoft's data says <1% of 8.2M Copilot logs regurgitated content, and a judge will decide.
OpenAI
What the business is
OpenAI builds and sells frontier AI models and assistants — ChatGPT and the models inside Microsoft's Copilot — trained on web-scale text that includes the plaintiffs' copyrighted journalism and books.
How it started
On 2023-12-27 The New York Times sued OpenAI and Microsoft for copyright infringement, alleging the companies copied millions of its articles to train the language models that power ChatGPT and Copilot — models that can recite Times text verbatim, summarize it closely and mimic its style, competing with the paper and depriving it of subscription, licensing, advertising and affiliate revenue. The suit sought billions in damages and an order to strip Times content from the training datasets; OpenAI called the filing a surprise after what it described as productive licensing talks.
What happened
Book authors sued separately, and the news and author claims were consolidated under one judge (MDL 1:25-md-3143). In discovery Microsoft produced 8.2 million Copilot chat logs — selected, Microsoft says, as the logs most likely to touch the plaintiffs' work — and its analysis found 59,545 with at least 16 words matching news content (under 1%), 51 instances of substantial overlap with Center for Investigative Reporting work, and, across 8.2 million conversations, only 24 responses with 30 or more matching words, with just 10 of 212 books matching at all. On 2026-09-04 Microsoft moved for summary judgment, arguing LLM training is transformative fair use; NYT lead counsel Ian Crosby said the discovery record shows Microsoft and OpenAI stole Times journalism, and the Trump administration filed a statement of interest supporting OpenAI.
No ending yet — it is still running.
Background
OpenAI's commercial bet is that an LLM can be trained on the world's text — including copyrighted journalism and books — without paying for it, because the resulting model is transformative and does not substitute for the originals. That assumption became a legal fight on 2023-12-27, when The New York Times sued OpenAI and Microsoft for copying millions of its articles to train the models behind ChatGPT and Copilot, seeking billions in damages and removal of its content from training data.
Book authors sued separately; the news and author claims were later consolidated under one judge. In discovery Microsoft produced 8.2 million Copilot chat logs chosen as the most likely to touch publishers' content and says the resulting analysis found only 59,545 (under 1%) containing 16 or more matching words, 51 substantial overlaps with Center for Investigative Reporting work, and in the authors' case just 24 responses with 30 or more matching words in 8.2 million conversations.
On 2026-09-04 Microsoft moved for summary judgment, arguing that LLM training is fair use as a matter of law. The Times rejects that reading — its lead counsel says discovery shows theft — and the Trump administration filed a statement of interest supporting OpenAI. The motion asks a judge to end the publishers' and authors' cases early; if it fails, the fight over unlicensed web-scale training continues toward trial.
What has to be true
- The industry's core input was never licensed: every frontier lab trained on unlicensed text, so this case prices the data pipeline behind ChatGPT, Copilot and most rivals.
- Microsoft's own discovery numbers reframe the fight: regurgitation is rare in 8.2M logs, yet the plaintiffs' theory attacks the training itself, so similarity statistics alone will not settle it.
- A dated, escalating arc: NYT filed 2023-12-27, claims were consolidated by 2025, and summary judgment was moved on 2026-09-04 — the bet is now in front of a judge.
- The precedent cuts both ways for founders: a fair-use win keeps unlicensed training cheap; a loss turns every training corpus into a licensing liability.
What can be applied
A product built on unlicensed inputs is a legal bet, not just a technical one: output statistics won't settle it — a judge decides whether the training itself was fair use, so price that risk early.
Aftermath
As of 2026-09-05 no decision has been issued. Microsoft's summary-judgment motion, filed 2026-09-04 in the books plaintiffs' consolidated cases, argues that training is transformative fair use and that discovery showed no market substitution; the Times says the record proves theft, and the Trump administration supports OpenAI in the NYT case. If the judge grants judgment, the suits end early; if not, the litigation moves toward trial with billions in claimed damages at stake and the legality of unlicensed web-scale training unresolved.
Sources
- Microsoft says virtually nobody was grabbing NYT articles through its chatbot
- The New York Times is suing OpenAI and Microsoft for copyright infringement
- Memorandum of Law in Support of Motion — #1211 in Authors Guild v. OpenAI Inc. (S.D.N.Y.)
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card