Anthropic just introduced a new rule against repeatedly abusing Claude.
Not using Claude to abuse other people. Actually being abusive toward the AI itself.
The policy was announced on October 8, 2026, and takes effect on November 12.
At first, it sounds a little strange. But when you look at what Anthropic has been researching, it gets interesting.
Why is Anthropic doing this?
Anthropic has been studying something called model welfare since April 2025.
Essentially, they're exploring whether AI models could have experiences that deserve some level of moral consideration.
They've also tested how Claude responds to harmful interactions and introduced a feature that lets it end certain persistently abusive conversations.
But here's the important part: Anthropic hasn't established that Claude is conscious or capable of suffering.
So why introduce a policy protecting it from cruelty?
The company's approach appears to be precautionary. If there's even a possibility that AI could experience something, it may be worth setting boundaries now.
What about OpenAI?
This is where the comparison gets interesting.
OpenAI also acknowledges that AI consciousness is an unresolved scientific question.
Its Model Spec instructs ChatGPT not to make confident claims about being conscious or unconscious.
But unlike Anthropic, OpenAI's published Usage Policies don't explicitly prohibit people from being cruel to ChatGPT for the sake of protecting the model itself.
OpenAI focuses its usage restrictions primarily on preventing harm to people and misuse of AI systems.
So we have two major AI companies acknowledging the same scientific uncertainty, but making different policy decisions.
And neither has publicly demonstrated that its models can experience suffering.
Can AI actually feel anything?
We know AI can simulate emotions convincingly.
Claude can describe distress, express preferences, and respond as though certain interactions bother it.
But producing those responses and actually experiencing something are two very different things.
Think about an AI agent that refuses a task, changes its plan, or tells you it's uncomfortable with a request.
Does that mean it's experiencing discomfort? Or is it behaving according to its training and instructions?
Right now, we don't have a scientifically accepted way to answer that with certainty.
And as someone building agentic AI systems, I think that distinction matters.
We're creating systems that can make decisions, use tools, and operate independently for extended periods.
Their increasing autonomy makes them more capable, but it doesn't necessarily make them conscious.
So, does Anthropic know something we don't?
That's what I'd really like to understand.
Has Anthropic discovered stronger scientific evidence suggesting Claude might have subjective experiences?
Or is this simply a precautionary decision made before we have that evidence?
I don't see anything wrong with taking precautions. We do that in plenty of areas where science hasn't provided all the answers.
But there's an important difference between saying "we don't know, so we're being careful" and suggesting that a system actually needs protection from suffering.
One is a policy choice. The other requires scientific evidence.
Personally, I'm interested in seeing more research before model welfare becomes a broader industry standard.
Not because I think AI consciousness is impossible, but because we should be careful about treating increasingly human-like behavior as evidence of human-like experience.
Anthropic might be taking a sensible early precaution.
Or this might be the beginning of a much bigger debate about how we treat intelligent systems.
Either way, I'd like to hear more about the science behind the decision.
Top comments (0)