Unsealed filings in the three-year-old New York Times copyright case put the frontier’s own words on the record for how training data was gathered. Microsoft’s internal January 2023 memo, from applied-science director Brent Hecht, called the scrape “an astonishing theft” and “the largest theft of labor in human history”; Microsoft’s own presentation showed its Copilot “answer engine” cost the Times up to 93% of its click-through versus ordinary Bing; and a deposition has Satya Nadella saying paywalled content “should be licensed by anyone who wants to use it” — adding, per the filings, that he would want his own teams told how that data was sourced. The quotes come out of the Times’ brief above still-sealed exhibits, so this is filings, not verdicts — but it is the strongest evidence to land so far in the case that decides who actually owns the frontier’s input.
The fair-use defense just met the defendant’s own spreadsheets
What happened. Per filings reported by TechCrunch and Ars Technica, an internal January 2023 memo by Brent Hecht, Microsoft’s director of applied science, described the training scrape as “the largest theft of labor in human history.” Microsoft’s own deck showed Copilot cutting click-through to the Times’ domain by up to 93% versus Bing; OpenAI’s head of ChatGPT, Nick Turley, called the threat to publishers “existential” and the chatbot “largely substitutive, period”; and the filings describe datasets at scale — more than 91,000 copies of plaintiffs’ works in mid-training sets, 160,000+ unique publisher works in the “Project Mango” set — plus a paywall workaround one OpenAI researcher described as a “hack,” met with “ah nice” from Greg Brockman.
Why it matters. The fair-use defense of frontier training turns on not substituting for the original work’s market, and the strongest substitution numbers in this case are the defendant’s own — Microsoft’s deck, quantifying 93% fewer clicks. Discovery just put words inside the labs (“the largest theft of labor in human history”) on the record in the case that will decide what unlicensed training data is worth, which is the bill underneath every thread this column has run on the data lane. For the operator nothing about your rate card changes today, but the price of unlicensed web-scale ingestion just acquired a precedent-shaped anchor, and training-time licensing — already the open lane’s quieter advantage — just picked up the defendant’s own receipt.
Source: techcrunch.com, arstechnica.com
What I’m watching
Whether the court admits the click-through and paywall-circumvention evidence into the summary-judgment record — that ruling, not the trial, is the event that actually reprices frontier training data. And whether either lab, with “theft” and “existential” on the record, fast-tracks the licensing deals this filing makes cheaper to sign than to fight.