Unsealed filings: a Microsoft exec called AI scraping "the largest theft of labor in human history" -- in a memo to his own company
Newly unredacted documents in the New York Times' copyright suit against OpenAI and Microsoft show internal alarm, in January 2023, over the same scraping practices both companies kept using -- plus evidence Times click-through traffic fell as much as 93% after Copilot's launch.
Newly unredacted filings in In Re: OpenAI Inc. Copyright Infringement Litigation (1:25-md-03143, S.D.N.Y.) -- the consolidated case bundling the New York Times, Daily News, Center for Investigative Reporting and other publishers' claims against OpenAI and Microsoft, and the same litigation whose cross-motions for summary judgment Merit AC covered on September 13 -- show a Microsoft executive privately calling the industry's scraping practices "the largest theft of labor in human history," TechCrunch reported September 17. The line comes from a January 2023 internal memo by Brent Hecht, Microsoft's director of applied science, who called it "an astonishing theft of unprecedented proportions" -- more than a year before the Times filed suit, and years before either company changed the underlying practice.
What the documents show
Per the filings, OpenAI employees developed methods to pull New York Times content past its paywall undetected, and training datasets built by both companies stripped copyright notices before use -- one Common Crawl-derived set alone reportedly containing more than 2 million documents from nytimes.com, part of a broader collection of over 91,692 copies of Times, Daily News and other publisher works. Separately, the filings cite evidence that Microsoft's Copilot cut the Times' click-through rate by as much as 93% compared with traditional search referrals, and quote OpenAI's Nick Turley describing publishers as facing an "existential threat" from chatbots -- while Microsoft CEO Satya Nadella has testified that paywalled content should require licensing.
Why this matters beyond the courtroom
An internal admission that a company's own AI training methods amounted to "theft," made a year before litigation and never acted on, is a different kind of evidence than the fair-use doctrine arguments both sides are otherwise making to Judge Sidney Stein. For any buyer weighing AI vendor risk, it's a reminder that training-data provenance isn't just a compliance checkbox -- it's exactly the kind of internal paper trail that turns into deposition material, and eventually into licensing costs (or judgments) that get passed down the vendor chain to enterprise customers relying on these models today.