hn.today

Microsoft exec called AI scraping 'the largest theft of labor in human

techcrunch.com13 points3 comments
Screenshot of Microsoft exec called AI scraping 'the largest theft of labor in human

New unredacted filings in a long-running copyright suit disclose internal Microsoft and OpenAI communications admitting large-scale, paywall-bypassing scraping and use of news content to train large language models. Executives and researchers described the practice as theft and an existential threat to publishers: a Microsoft researcher called the data grab “the largest theft of labor in human history,” OpenAI staff discussed hacks to get around the New York Times paywall, and internal metrics showed Microsoft’s Copilot reduced click-through rates to the Times by as much as 93%, creating a “doom loop” that could harm both publishers and model performance. The filings allege Project Mango/Taxi, WebText and Common Crawl-derived datasets included vast numbers of publisher works (more than 91,692 copies across some mid-training sets and over 2 million nytimes.com documents in a Common Crawl extract), and that teams deliberately stripped copyright notices before feeding material to models.

Those admissions undercut core fair-use defenses by showing direct market substitution and intentional acquisition tactics: OpenAI leaders called the models “excellent at news” and “largely substitutive,” while Microsoft’s CEO testified paywalled material should be licensed and could require retraining if improperly used. Documents also claim Microsoft and OpenAI exchanged training data and coordinated ingestion strategies. With underlying exhibits still sealed and courts historically leaning toward fair use, these revelations sharpen legal and policy stakes by tying model performance to publisher harm and potential employment disruption, increasing pressure on how licensed content and model training will be regulated or litigated going forward.

Read on techcrunch.com3 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Security

Hacking OpenAI

Hacking OpenAI

Hackers exploited two vulnerabilities, including a heap buffer overflow and an SSO misconfiguration, to access OpenAI employees' ChatGPT accounts and internal repositories. The attack demonstrated how quickly security flaws can be chained together to compromise sensitive systems within 72 hours. (hacktron.ai)

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.