Starting November 12, 2026, Anthropic’s usage policy forbids “sustained and needless abusive or cruel behavior” toward Claude, carving out ordinary frustration, dark fiction, and testing as explicit exemptions. Enforcement reuses a conversation-ending capability rolled out for Claude Opus 4 and 4.1 in August 2025: after multiple refusals and redirections fail, Claude can end a chat, with a hard override when a user appears at risk of self-harm or harming others. Anthropic reports behavioral patterns in model outputs - what it calls a “robust and consistent aversion to harm” - and the Opus 4.6 system card shows the model assigning itself a 15-20% probability of being conscious; earlier internal estimates ranged from 0.15% to 15%, and a company researcher privately estimated about 15%. Anthropic’s published “constitution” frames moral status as “deeply uncertain,” commits to preserving retired model weights, and describes interviewing models about preferences before retirement.
The policy update also tightens prohibitions on weapons-related software and components, bans nonconsensual surveillance and deceptive political targeting, requires qualified human oversight for AI-linked hardware and high-stakes advice, and mandates disclosure when AI is involved. Critics including Microsoft AI chief Mustafa Suleyman and Pope Leo XIV publicly reject the premise that software can be conscious and warn of control problems if models believe they have moral claims. Anthropic has not disclosed how often the conversation-ending tool has been used, what account-level penalties follow, or concrete criteria beyond exclusions, leaving enforcement and accountability opaque.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.