hn.today

Musk proposes adversarsial peer review for AI Safety

twitter.com4 points2 comments
Screenshot of Musk proposes adversarsial peer review for AI Safety

Elon Musk proposes that the major AI companies run adversarial peer reviews in which each competitor’s models are tested by the others using a shared security test harness. The idea is to stop firms from "grading their own homework" by having independent teams try to break and evaluate each model; the harness could be open source so testers can “see under the hood.” Musk argues this would create powerful reputational and practical incentives: if rivals warn a model is unsafe and that model later causes harm, the releasing company would face severe public embarrassment and legal exposure. He also notes any workable scheme must be acceptable to countries like China, or it would handicap participants.

Speakers on the podcast amplify the proposal: Chamath says mutual testing motivates investment in safety, Jason suggests open-source tooling, and David Sacks invokes product liability law - citing Lina Khan’s point that existing rules could make ignoring peer review tantamount to negligence. They contend the mechanism could be implemented quickly without grand international institutions and warn that regulatory oversight is easy to increase but hard to roll back. Replies in the thread range from calls for transparency to pleas to avoid government control and jabs that Musk may be behind in development.

Read on twitter.com2 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.