Microsoft Exec Calls AI ‘Labor Theft,’ Records Show

The unsealed filings say OpenAI’s training dataset contained more than 91,692 data points from copyrighted works produced by The New York Times, Daily News and the Center for Investigative Reporting.
The filings reportedly include a Microsoft memo warning that generative AI could “significantly disrupt the jobs of the very people who generated the data used to train the underlying models.”
Microsoft allegedly received the full dataset used to train GPT-3 and used it to assess how OpenAI’s models could be incorporated into Microsoft products; the documents also reference Microsoft supplying training data to OpenAI through initiatives called “Project Taxi” and “Project Mango.”
The Times’ legal motion argues that the executives’ statements undermine the defendants’ fair-use defense; it quotes Hecht as saying that allowing the defense to prevail could “make a complete mockery of the idea of ‘fair use.’”
The filings reportedly describe efforts by OpenAI to bypass paywalls without detection in order to maximize access to scraped material.
Court filings unsealed in The New York Times' copyright lawsuit reveal damning internal admissions from Microsoft and OpenAI executives. Decrypt reports that Microsoft director Brent Hecht called AI training data scraping "the largest theft of labor in human history," while OpenAI leader Greg Brockman warned that generative AI poses an "existential threat" to publishers. The documents show OpenAI's dataset contained over 91,692 copyrighted articles from The Times, Daily News, and other outlets — all used without permission to train competing commercial AI products.
The Times argues these internal quotes demolish Microsoft and OpenAI's fair-use defense. A decision on whether the case proceeds to trial is expected by 2027. Tom's Hardware notes the revelations center on whether tech companies can legally scrape news articles to train their AI systems without paying publishers or getting consent.
Internal Microsoft memos show the company knew scraping would harm journalists. The documents warn that generative AI could "significantly disrupt the jobs of the very people who generated the data used to train the underlying models," according to Tom's Hardware. Microsoft received the full GPT-3 training dataset and used it to evaluate how OpenAI's models could fit into Microsoft products. The company also supplied training data to OpenAI through secret projects called "Project Taxi" and "Project Mango."
Court documents reveal OpenAI deliberately worked around paywalls to grab more copyrighted articles. Yahoo Tech reports that executives worried the scraping could create a "doom loop" — where AI products trained on news eventually replace the journalists who created that news. OpenAI's stated goal was to maximize access to scraped material, showing deliberate intent to harvest content at scale without detection or compensation.
Brent Hecht went further than calling scraping "theft." According to the Times' legal motion, Hecht said that allowing Microsoft and OpenAI's fair-use defense to succeed would "make a complete mockery of the idea of 'fair use.'" His statements directly contradict the defendants' core legal argument. Engadget notes these admissions from the executives themselves — not critics or competitors — give the Times enormous credibility in court.
The Times is seeking summary judgment, meaning a judge could rule in favor of the plaintiff without a full trial. A win could force tech companies to pay publishers for training data or stop scraping entirely. Times Now News reports that Hecht's internal comments frame the legal debate as a battle over whether "theft of labor" becomes acceptable when dressed up as AI training. The case decision expected in 2027 will likely set the standard for how tech companies train AI models.
Publishers
21
Articles
28
Reach
49