Executives at OpenAI and Microsoft admitted in recently unsealed court materials that large language models were trained on appropriated human-created content and are actively destroying the commercial web. Internal documents and depositions describe the training corpus as “an astonishing theft” and “the largest theft of labor in human history,” and warn of a “doom loop” in which AI products cannibalize clicks and revenue from the sites and creators that supplied their data. Lawyers compiling these admissions filed them in a summary-judgment submission to challenge industry claims that model training is fair use or fundamentally transformative.
The filings include concrete examples and blunt quotes: a Microsoft analysis showing Bing traffic to news sites dropped by more than 90% after content was used; engineers and executives describing hacks to bypass paywalls; statements that models substitute for the labor that defines culture and do not route economic value back down the supply chain; and admissions that users won’t click links even when provided. Taken together, the evidence argues that current LLM deployment relied on large-scale unconsented copying and now poses an existential threat to writers, artists, and news businesses while contradicting public legal defenses.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.