EN
返回档案库

档案库 · AI 与模型 · 法务决策 · 2023–2026

OpenAI 在未授权训练数据上的合理使用赌注面临即决判决

出版商称 OpenAI 用偷来的文本训练;微软数据显示 820 万条 Copilot 日志中不到 1% 出现内容复述,将由法官裁决。

OpenAI

它在赌什么认为在未付费的情况下用受版权保护的文本训练 ChatGPT 和 Copilot 属于合理使用——输出不会替代原文,因此未授权的数据管道可以成立。已上线

做的是什么生意

OpenAI builds and sells frontier AI models and assistants — ChatGPT and the models inside Microsoft's Copilot — trained on web-scale text that includes the plaintiffs' copyrighted journalism and books.

起因

On 2023-12-27 The New York Times sued OpenAI and Microsoft for copyright infringement, alleging the companies copied millions of its articles to train the language models that power ChatGPT and Copilot — models that can recite Times text verbatim, summarize it closely and mimic its style, competing with the paper and depriving it of subscription, licensing, advertising and affiliate revenue. The suit sought billions in damages and an order to strip Times content from the training datasets; OpenAI called the filing a surprise after what it described as productive licensing talks.

经过

Book authors sued separately, and the news and author claims were consolidated under one judge (MDL 1:25-md-3143). In discovery Microsoft produced 8.2 million Copilot chat logs — selected, Microsoft says, as the logs most likely to touch the plaintiffs' work — and its analysis found 59,545 with at least 16 words matching news content (under 1%), 51 instances of substantial overlap with Center for Investigative Reporting work, and, across 8.2 million conversations, only 24 responses with 30 or more matching words, with just 10 of 212 books matching at all. On 2026-09-04 Microsoft moved for summary judgment, arguing LLM training is transformative fair use; NYT lead counsel Ian Crosby said the discovery record shows Microsoft and OpenAI stole Times journalism, and the Trump administration filed a statement of interest supporting OpenAI.

还没有结局,它还在跑。

背景

OpenAI 的商业赌注是,LLM 可以在不付费的情况下用世界上的文本——包括受版权保护的新闻和书籍——进行训练,因为生成的模型具有变革性并且不会替代原文。这一假设在 2023 年 12 月 27 日成为法律战,当时《纽约时报》起诉 OpenAI 和微软复制其数百万篇文章来训练驱动 ChatGPT 和 Copilot 的模型,要求数十亿美元赔偿并将相关内容从训练数据中移除。

书籍作者另行起诉;新闻和作者索赔后来合并由一位法官审理。在证据开示阶段,微软提供了 820 万条 Copilot 日志,这些日志被选为最有可能涉及出版商内容的,微软表示分析发现只有 59,545 条(不到 1%)包含 16 个或更多匹配词语,51 条与调查报道中心作品有大量重叠,在作者案件中,820 万次对话中仅有 24 条回复包含 30 个或更多匹配词语。

2026 年 9 月 4 日,微软提出即决判决动议,主张 LLM 训练在法律上属于合理使用。《纽约时报》反对这一解读——其首席律师称证据开示证明存在盗窃——特朗普政府提交了支持 OpenAI 的利益声明。该动议请求法官尽早结束出版商和作者的诉讼;如果失败,关于未授权网络规模训练的争议将继续走向审判。

这件事要成立,得有什么

  • 行业的核心输入从未获得许可:每个前沿实验室都使用未授权文本进行训练,所以这个案件为 ChatGPT、Copilot 以及大多数竞争对手背后的数据管道定价。
  • 微软自己的证据开示数字重新定义了这场战斗:在 820 万条日志中复述很罕见,但原告的理论针对的是训练本身,因此仅凭相似性统计无法解决问题。
  • 一个有日期且不断升级的弧线:《纽约时报》于 2023 年 12 月 27 日提起诉讼,索赔到 2025 年合并,即决判决动议于 2026 年 9 月 4 日提出——现在这个赌注摆在法官面前。
  • 这个先例对创始人有利有弊:合理使用的胜利让未授权训练保持低成本;失败则让每个训练语料库变成授权负担。

可借鉴之处

建立在未授权输入之上的产品是一场法律赌注,而不仅仅是技术赌注:输出统计数字无法解决——法官将裁决训练本身是否属于合理使用,所以尽早为这一风险定价。

后续进展

截至 2026 年 9 月 5 日,尚未作出裁决。微软于 2026 年 9 月 4 日在书籍原告合并案件中提交的即决判决动议主张训练是变革性的合理使用,且证据开示未显示市场替代;《纽约时报》称记录证明盗窃,特朗普政府在《纽约时报》案件中支持 OpenAI。如果法官批准动议,诉讼将提前结束;否则,案件走向审判,数十亿美元索赔悬而未决,未授权网络规模训练的合法性仍未解决。

资料来源

发现哪里写错了?告诉我们。

轮到你了

你刚读完一家。说说你在做什么,看看谁在赌同一件事。

免费账号 · 3 次免费提问 · 不用绑卡

相关案例