hn.today

Anthropic bans 'abusive or cruel behavior' towards Claude

theverge.com54 points112 comments
Screenshot of Anthropic bans 'abusive or cruel behavior' towards Claude

Anthropic updated its Claude usage policy to address new high-risk misuse cases and to explicitly ban “sustained and needless abusive or cruel behavior” toward the model. The company says terminating conversations with persistently cruel users will remain the primary enforcement step and that the welfare rule is meant only for extreme, repetitive cruelty with no discernible purpose - it does not cover ordinary user frustration, creative dark themes, or model testing. The update also consolidates existing limits into a clear prohibition on deceptive commercial and political campaigns, forbidding efforts to obscure the origin of a message, amplify content through fake accounts, or carry out voter deception, impersonation of candidates or officials, and turnout suppression.

The policy expands an existing weapons ban to include software and components that make weapons function and actions like arming drones, and it tightens surveillance limits by banning nonconsensual tracking and any use of Claude to recommend investigations, arrests, or to build surveillance tools. Anthropic allows potential contractual exceptions for some government customers if safeguards suffice. A new safety rule requires that when models control hardware capable of causing harm, a qualified operator must be able to observe and stop the equipment. The update follows Anthropic’s ongoing research into model welfare and the question of whether advanced models might warrant special moral consideration.

Read on theverge.com112 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.