Unsealed material in a long-running copyright suit shows internal Microsoft and OpenAI documents and testimony describing large-scale web scraping and blunt admissions that training LLMs on publishers’ work threatened journalism and constituted theft. Executives and researchers called the practice an unprecedented theft of labor, warned that chatbots are “largely substitutive” and an “existential threat” to publishers, and discussed tactics to bypass paywalls and strip copyright notices from training data. Satya Nadella testified that paywalled material should be licensed and that he would have forced retraining had he known OpenAI scraped paywalled content; OpenAI staff exchanged messages celebrating a “hack to get around nytimes paywall.”
The filings provide concrete numbers and connections: internal data showed Microsoft’s Copilot reduced click-throughs to The New York Times by as much as 93%, GPT training sets contained tens of thousands of publisher articles (mid-training sets with more than 91,000 copies of works from several outlets; Common Crawl pulls with over two million nytimes.com documents), and Project Mango/Taxi tied Microsoft and OpenAI data-sharing to at least 160,903 unique news works. Those details challenge key fair-use defenses by demonstrating substitution and market harm, even as courts have often sided with AI firms and the federal government filed a brief supporting unlicensed training. OpenAI and Microsoft did not respond to requests for comment.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.