Ground Truth.
AI, checked against the source.

News · 2026-09-19

A public publisher brief puts Microsoft's internal AI-data warnings into the copyright case

A September 17 public filing in the consolidated publisher copyright case against OpenAI and Microsoft alleges industrial-scale copying and quotes a Microsoft executive warning that model data collection could be viewed as the largest theft of labor in human history. The legal news is the release of a party's evidence and argument, not a ruling that the companies infringed.

Key facts

The document's most memorable phrase is not the most important evidence. The underlying claim is a proposed feedback loop: AI products ingest journalism, answer users without sending them to publishers, reduce traffic and revenue, and then starve the reporting ecosystem that supplies future information. The filing calls this a doom loop. Like a restaurant that uses a farm's produce to sell meals while depriving the farm of its customers, the alleged harm is not merely a copy being made; it is a product substituting for the source relationship.

The plaintiffs point to discovery involving Common Crawl-derived data, Bing-indexed content and projects called Taxi and Mango. They allege Project Mango contained at least 160,903 unique publisher works and cite internal discussions about paywall circumvention. These are litigants' allegations and selected discovery excerpts, not neutral fact findings. That distinction also applies to a cited comparison alleging Copilot could reduce click-through to The New York Times by as much as 93 percent relative to conventional Bing.

OpenAI's case page and its summary-judgment memorandum provide the essential countercase. OpenAI argues training learns statistical relationships rather than distributing expressive copies, that facts are not copyrightable, and that outputs are rarely verbatim. It also disputes market substitution with economist evidence. Those arguments are not erased by an internal employee's rhetorical warning.

Still, the filing changes the public texture of the debate by putting the labor-market mechanism beside copyright doctrine. The sophisticated question is whether a legally transformative training process can nevertheless create a commercially substitutive interface that undermines the market the law is meant to protect. That will be decided through contested evidence and law, not a viral quotation. Readers tracking the issue should distinguish training data attribution from legal liability: being able to trace a model's inputs does not itself answer fair use, market harm or consent.


Primary source, verified: read the paper →

Key questions

Did a judge rule that Microsoft stole publishers' work?

No: the document is the News Plaintiffs' summary-judgment brief and its allegations have not been adjudicated.

What did the Microsoft memo say?

The filing attributes to Microsoft applied-science leader Brent Hecht a January 2023 description of models hoovering up work as an unprecedented theft of labor.

What is the claimed doom loop?

The plaintiffs say internal material described AI answers reducing publisher traffic and revenue, which then weakens the source ecosystem the systems rely on.
Cite this

APA

Ground Truth. (2026, September 19). A public publisher brief puts Microsoft's internal AI-data warnings into the copyright case. Ground Truth. https://groundtruth.day/news/microsoft-openai-news-plaintiffs-summary-judgment-brief.html

BibTeX

@misc{groundtruth:microsoft-openai-news-plaintiffs-summary-judgment-brief,
  title  = {A public publisher brief puts Microsoft's internal AI-data warnings into the copyright case},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/microsoft-openai-news-plaintiffs-summary-judgment-brief.html}
}

Topics: copyright · policy · publishing · microsoft · openai

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.