hn.today

Microsoft, OpenAI lose fight to hide internal docs admitting scraping is theft

arstechnica.com34 points7 comments
Screenshot of Microsoft, OpenAI lose fight to hide internal docs admitting scraping is theft

A recently unsealed motion in a lawsuit brought by news publishers led by the New York Times reveals internal Microsoft and OpenAI documents showing executives privately warned that mass-scraping news for large language models was effectively theft and posed an existential, substitutive threat to journalism. Microsoft scientist Brent Hecht described the practice in stark terms, and OpenAI staff including Nick Turley warned of a “doom loop” in which chatbots would cannibalize news traffic and undermine the supply of reliable reporting. Company data cited in the filings show dramatic drops in click-through rates from search or chatbot referrals - reported declines as large as 83-93% for some plaintiffs and 51-94% for others - while CEO Satya Nadella acknowledged under oath that AI can divert users away from publishers when it provides answers directly.

Publishers documented tests in which chatbots reproduced lengthy verbatim excerpts when asked for summaries, bullet points, or bias ratings, and are asking the court to rule on instances of substantial overlap rather than wait for trial on all claims. The motion also accuses the defendants of repurposing datasets improperly - Microsoft allegedly sold Bing-acquired data for model training, and OpenAI used a 1.8 million-article New York Times dataset tied to noncommercial terms despite internal recognition that was inappropriate. Microsoft and OpenAI counter that the products are transformative and that isolated employee comments do not reflect company positions. Publishers argue that only a ruling requiring licensing will prevent an industry-wide “free‑riding” prisoners’ dilemma that could starve journalism and destabilize responsible AI.

Read on arstechnica.com7 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Security

Hacking OpenAI

Hacking OpenAI

Hackers exploited two vulnerabilities, including a heap buffer overflow and an SSO misconfiguration, to access OpenAI employees' ChatGPT accounts and internal repositories. The attack demonstrated how quickly security flaws can be chained together to compromise sensitive systems within 72 hours. (hacktron.ai)

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.