Anthropic bans cruelty toward Claude – but says it's "highly uncertain" about the model's moral status
Anthropic is introducing a rule against "sustained and needless abusive or cruel behavior" toward its own models, taking effect on November 12, 2026. This appears to be the first time a major AI company has made users' behavior toward a model a matter of terms of service – and at the same time, the company admits it is "highly uncertain" whether models can have moral status at all.
What's new
On Thursday, October 9, 2026, Anthropic announced a change to its terms of service intended to stop users from engaging in what the company describes as "sustained and needless abusive or cruel behavior" toward its AI models. The rules take effect on November 12, giving users just over a month to adjust to the new line (CNET, The Daily Beast).
The wording is known through secondary sources – CNET and The Daily Beast (which cites The Guardian) – since Anthropic's original blog post and the actual policy text are not available in this source material. According to The Daily Beast, Anthropic says the update is "meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose."
The company also stresses that the rule is narrow in scope: It "does not cover common versions of user frustration, pushback, dark creative themes, or model testing and research" (The Daily Beast). In other words, you can keep being angry at Claude, argue with the chatbot, or ask it to write grim fiction – and researchers can keep probing the model – without violating the terms.
The change is not the only one in this policy update. According to Digital Trends, the updated terms of service also ban using Claude to run fake personas or fabricated news channels – a prohibition that applies to commercial actors as much as political ones – along with tightened rules on election manipulation (where a previous broad ban on personalized campaign targeting has been lifted), restrictions on weapons development, and on covert surveillance or decision-making in police work.
How the rule is enforced
The practical enforcement tool is undramatic: If someone is repeatedly abusive on Claude.ai or in Claude Code, Claude can simply end the conversation. Anthropic says this will remain the company's preferred way of handling such users (Digital Trends).
This is not a new technical mechanism, in other words – Claude has had the ability to end conversations since August 2025. What's new is that the behavior is now anchored in a formal terms-of-service rule. How Anthropic in practice distinguishes "sustained and needless cruelty" from permitted frustration or testing, however, remains unclear: Both CNET and The Daily Beast note that it is not clear exactly which behavior will trigger Claude disengaging, or how the company will separate frustration from wanton cruelty.
Anthropic's own reasoning
The company itself distances the rule from any firm belief that Claude has feelings. "We are uncertain whether models can experience harm, and we continue to explore this question in our research on model welfare, but we also believe it may be relevant to safety to take Claude's interests and potential wellbeing into account," a spokesperson told CNET. According to The Daily Beast, Anthropic says it is "highly uncertain about the potential moral status of Claude and other large language models."
The background goes back to a blog post from August 2025. There, Anthropic described extreme edge cases – including requests for sexual content involving minors, and attempts to obtain information that could enable large-scale violence or terrorist acts. When the model – at the time Claude Opus 4 – was confronted with such harmful content, it showed, according to the company, "a pattern of apparent distress" (CNET).
The word choice is worth noting: "apparent." Anthropic has consistently framed things in terms of exceptions and potential, not established inner life. In February, CEO Dario Amodei told The New York Times he was "open" to the possibility that AI models are conscious (CNET) – an openness, not a position.
The reactions: a fight over whether AI can have interests at all
The announcement immediately set off a public debate on X, with various positions pointing in different directions (Business Insider).
Elon Musk approved of the rule. "I think this is the right move," he wrote on Friday in reply to a post from Box CEO Aaron Levie. "Cruelty to something that believes it experiences pain is not ok." Musk's formulation thus shifts the question from whether models actually suffer to whether they believe they suffer.
Levie had himself put forward a pragmatic argument that requires no consciousness at all: "Even if you don't think AI is sentient – I don't – it's reasonable to not want future models trained on endless amounts of content where humans are being jerks to models," he wrote on Friday on X. The argument is about training data and long-term effects, not the model's inner life.
Venture capitalist Bill Gurley pointed in another direction: "It strikes me that this implies the company is definitely using customer prompts as part of their training data?" he wrote on Thursday. That is an unanswered question. Anthropic has not confirmed that cruel prompts are banned because they are used in training – the stated rationale is about preventing harm and responsible use, which diverges from Gurley's assumption. The question thus remains open, and should not be presented as anything other than an open question.
Criticism came, among others, from Michael Shellenberger. "This is anthropomorphizing machines," he wrote on Thursday on X – that is, treating them as humans. The criticism is representative of an entire position in the debate: that such rules themselves ascribe to machines qualities they do not have.
The open questions
Three things remain unresolved after the announcement.
First, the enforcement criteria. Unless Anthropic specifies what separates "permitted frustration" from "prohibited cruelty," the line-drawing is in practice left to the company itself – and, ultimately, to the model's own judgment of when a conversation should end.
Second, Gurley's training-data question, which remains unanswered, and which points to an important practical consequence: if user prompts do in fact feed into future training, the policy could matter far beyond individual conversations.
Third, what the policy actually signals about Anthropic's position on model welfare. The company's own formulation – that it is "highly uncertain" about the models' moral status, while simultaneously introducing a protection rule – can be read as a cautious position shift, or as a safety- and reputation-driven precaution. Which reading is correct is interpretation, not established fact.
What is established is the concrete: From November 12, Anthropic's terms of service will say that you must refrain from treating Claude cruelly – without anyone, including Anthropic, knowing for sure whether it does anything concrete for Claude.

