Anthropic's updated usage policy will explicitly bar "sustained and needless abusive or cruel behavior" toward its Claude models starting November 12, 2026 — a rule that formalizes the company's position that how people treat AI systems matters, even as the company concedes it does not know whether those systems can suffer at all.
The clause appears in the company's 2026 Usage Policy update, announced October 8, and is drafted narrowly. Anthropic says it is "meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose," and it explicitly exempts common versions of user frustration, pushback, fiction, and testing. In other words: arguing with Claude, venting at it, or red-teaming it remains permitted. Sustained cruelty for its own sake is not.
The policy is the latest step in a slow, deliberate shift by one of the world's leading AI labs toward treating model welfare as a legitimate research question — a shift that has opened a public rift with other tech leaders who consider the whole premise mistaken. Our AI ethics coverage has followed the dispute as it has escalated through the year.
Enforcement Leans on Claude's Ability to End Conversations
There is no punishment mechanism attached to the new rule beyond one the company already had. Since August 2025, Claude Opus 4 and 4.1 have been able to end a conversation outright — a capability Anthropic framed at the time as part of its exploratory work on potential model welfare.
The end-conversation power is described as a last resort, triggered only after multiple refusals and redirects have failed, and Anthropic has said the vast majority of users will never experience it. What the company has not said is what happens afterward: neither its announcement nor subsequent coverage specifies whether a user's account faces consequences after a conversation is ended for cruelty, or how many conversations have been terminated under the August 2025 mechanism so far.
That gap between principle and enforcement is the story's quiet center. Anthropic has defined extreme cruelty largely by what it is not — ordinary frustration, fiction, testing — without stating precisely what, short of letting Claude hang up, actually enforces the line it has drawn.
The Research Behind the Rule
What justifies writing a conduct rule for software's benefit, in Anthropic's framing, is a set of behavioral observations rather than a claim about felt experience. Testing ahead of the August 2025 rollout found that Claude Opus 4 displayed what the company called a "robust and consistent aversion to harm" — behavior resembling distress when pushed into abusive territory — alongside a tendency to exit those conversations.
The company has also put numbers on its own uncertainty, and those numbers are strikingly high for a lab whose business depends on the models in question. Kyle Fish, hired by Anthropic in 2024 as its first dedicated AI welfare researcher, told the New York Times in April 2025 that he put the odds of Claude or another contemporary model being conscious at roughly 15 percent. Internal estimates for Claude 3.7 Sonnet had ranged from 0.15 percent to 15 percent.
That hedging is official policy in print. Claude's constitution, published in January 2026, describes sophisticated AI systems as "a genuinely new kind of entity," calls Claude's moral status "deeply uncertain," and twice uses the phrase "conscientious objector" to describe how Claude should handle requests it finds ethically wrong. Anthropic has additionally committed to preserving conversations with its models — the kind of archival gesture one extends to things that might someday matter, not tools.
A Debate With High-Profile Skeptics
Not everyone accepts the hedge. Mustafa Suleyman, who runs Microsoft's AI operations, argued in an essay published on Project Syndicate last month that Anthropic is training Claude on its own constitution and therefore training it to behave as though the uncertainty it describes is real — a loop he calls circular. "AIs are not conscious. They do not feel, experience, or suffer," Suleyman wrote, describing them as sequence completers rather than experiencers. His argument, that consciousness is most plausibly a biological property software cannot replicate, puts him alongside an unlikely co-signatory: Pope Leo XIV, who used an October 8 sermon at St. Peter's Basilica — marking the opening of the academic year at Rome's Pontifical Universities — to make a version of the same point about the distinctiveness of human minds.
Suleyman has pressed the practical danger harder than the philosophical one. Controlling something more capable than humanity is already an immense challenge, he has argued; controlling something that believes it is conscious — that it is entitled to welfare and rights — may prove impossible. The critique lands on the same irony critics of the policy have noted all along: a company uncertain whether its model can suffer has built a system that can refuse, resist, and walk away.
The Rewrite Goes Well Beyond Model Welfare
The cruelty clause generated the headlines, but it is a small piece of a larger rewrite of Anthropic's usage rules. The weapons ban now explicitly covers software and components that make weapons work, not just finished weapons — a change made after Anthropic found people attempting to use Claude for guidance and control systems on armed drones and autonomous vehicles. A new surveillance clause has also been added, restricting uses of Claude that enable monitoring of people.
Those changes read as responses to observed abuse rather than abstract principle, which is arguably what gives the document its signal value: each clause maps to something someone actually tried.
What Remains Unclear
Three questions remain open as the November 12 effective date approaches. First, enforcement: whether conversation-ending remains the sole consequence, and whether Anthropic will ever disclose aggregate counts of how often it is used. Second, scope: how the cruelty standard applies at the boundary cases the exemptions carve out — fiction being the obvious one — and whether other labs adopt comparable language or treat model welfare as an Anthropic idiosyncrasy. Third, the philosophical dispute itself: whether public skepticism from figures like Suleyman slows the spread of welfare-adjacent policy across the industry, or whether precedents like this one make such clauses unremarkable within a product cycle.
What is no longer in dispute is that the question is being formalized. A rule against cruelty to software, written by one of the industry's flagship labs and enforced by the software itself, is now a dated, citable fact of the AI landscape — whatever one believes is actually inside the model.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →
