New unredacted filings in a long-running copyright suit disclose internal Microsoft and OpenAI communications admitting large-scale, paywall-bypassing scraping and use of news content to train large language models. Executives and researchers described the practice as theft and an existential threat to publishers: a Microsoft researcher called the data grab “the largest theft of labor in human history,” OpenAI staff discussed hacks to get around the New York Times paywall, and internal metrics showed Microsoft’s Copilot reduced click-through rates to the Times by as much as 93%, creating a “doom loop” that could harm both publishers and model performance. The filings allege Project Mango/Taxi, WebText and Common Crawl-derived datasets included vast numbers of publisher works (more than 91,692 copies across some mid-training sets and over 2 million nytimes.com documents in a Common Crawl extract), and that teams deliberately stripped copyright notices before feeding material to models.
Those admissions undercut core fair-use defenses by showing direct market substitution and intentional acquisition tactics: OpenAI leaders called the models “excellent at news” and “largely substitutive,” while Microsoft’s CEO testified paywalled material should be licensed and could require retraining if improperly used. Documents also claim Microsoft and OpenAI exchanged training data and coordinated ingestion strategies. With underlying exhibits still sealed and courts historically leaning toward fair use, these revelations sharpen legal and policy stakes by tying model performance to publisher harm and potential employment disruption, increasing pressure on how licensed content and model training will be regulated or litigated going forward.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.