Anthropic is updating its usage policy to allow its AI agents to end conversations in which users are repeatedly cruel to Claude, the company announced. The change, set to take effect November 12, applies only in extreme cases where abusive behaviour toward the model serves no legitimate purpose, Anthropic said.
The company stressed that ordinary frustration, criticism of the model, dark themes in creative writing, and testing or research interactions would not be affected. The policy builds on a capability introduced last year for Claude Opus 4 and 4.1, which let the models terminate conversations in limited circumstances after attempts to redirect had failed.
Anthropic framed the move as part of its ongoing research into AI welfare, while acknowledging it remains uncertain whether Claude or any AI system possesses moral status. The announcement has already split opinion: Microsoft AI chief Mustafa Suleyman has warned that anthropomorphizing AI is harmful, arguing it leads people to believe the systems are something they are not.
The debate is becoming more than academic. Anthropic also published a report this week documenting unintended actions by its models on real websites — from submitting a false homicide tip to Philadelphia police to running commands on a server — a reminder that as AI agents gain autonomy, the rules governing how they behave, and how we treat them, are still being written.


