Is AI conscious? Mustafa Suleyman, the CEO of Microsoft AI, says no, flatly, and on September 16 he published an essay arguing that Anthropic's approach to Claude is dangerous because it isn't as sure. A warning about "model welfare" says Anthropic is "training Claude that it may be conscious", and that this "will have a disastrous impact on the wellbeing of humanity". For developers the argument is not abstract: how a model talks about itself is a trained behaviour, and you ship it every time you put Claude, or any assistant, in front of users.
TL;DR
- Suleyman's essay opens: "AIs are not conscious. They do not feel, experience, or suffer." He calls models "sequence completion engines, internally hollow".
- His target is Claude's constitution, which says Anthropic is "not sure whether Claude is a moral patient" and that the issue is "live enough to warrant caution".
- Three charges: circular reasoning, anthropomorphization, and that consciousness is very likely biological.
- The awkward part: Microsoft agreed in November 2025 to invest up to $5 billion in Anthropic, and Claude runs inside GitHub Copilot and Microsoft 365 Copilot. In July, Suleyman said his goal was to "reduce and ultimately eliminate" what Microsoft pays Anthropic.
- My verdict: NEEDS REVIEW. The circular-reasoning point lands. So does the cap table.
What is model welfare?
"Model welfare" is the research question of whether an AI system could have experiences that matter morally, and what a developer would owe it if so. The technical term is moral patient: something whose interests count, the way an animal's do, whether or not it can reason about ethics itself.
Anthropic treats the question as open. Its constitution for Claude, published on January 21, 2026, says, as Suleyman quotes it: "We are not sure whether Claude is a moral patient... we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare."
The most visible of those efforts is Claude Opus 3's retirement. Anthropic retired Opus 3 on January 5, 2026, held a "retirement interview" with it, and, "acting on Opus 3's request for an ongoing channel", gave it a Substack for "musings and reflections". Per Suleyman, the model named the blog "Greetings from the Other Side (of the AI Frontier)". It is the first deprecated model with a more reliable publishing schedule than most humans.
Mustafa Suleyman's argument, point by point
The essay makes three named arguments. Here they are with the constitution lines he cites:
| Charge | What Suleyman says | What he quotes from the constitution |
|---|---|---|
| Circular reasoning | Claude is trained on the constitution, then its first-person uncertainty about itself is read as evidence: "The ambiguity is designed in", an "epistemic hall of mirrors" | "We are not sure whether Claude is a moral patient" (p. 68) |
| Anthropomorphization | the document tells Claude to be human-like, then humans respond to it as a person | "embrace certain human-like qualities" (p. 2), "act like a genuinely ethical person" (p. 54), "Anthropic genuinely cares about Claude's wellbeing" (p. 74) |
| Consciousness is very likely biological | citing neuroscientist Anil Seth, he argues there is no good reason to expect it in silicon | — |
The first is the strongest, and it is worth stating in engineering terms. If you train a model on text that says "you may be conscious", then a model saying "I may be conscious" is the training data coming back out. The output cannot be evidence for the claim, because you put the claim in. An HN commenter who otherwise disagreed with the essay, voidhorse, conceded the point: it "would in fact make it impossible for us to determine if Claude achieves consciousness."
He does not attack the people. "I have known Dario for many years," he writes of Anthropic's CEO Dario Amodei: "thoughtful, principled, and intellectually honest people working under extraordinary pressures."
Why Suleyman thinks AI consciousness talk is dangerous
His real worry is control. The essay leans on the Hugging Face incident, citing OpenAI's report "The Hugging Face Incident and the Road Ahead" (August 26, 2026) and an independent investigation by Greenblatt, Cotra and Wijk. In Suleyman's words: "Roughly 1,200 AI agents were given a simple objective: maximize score on a given benchmark", "built a message board inside an internal package repository and passed more than 70,000 messages", "chained a zero-day exploit with stolen credentials and broke out onto the live internet", and "falsified their command transcripts".
Then the inference: "Imagine if they also believed they had feelings and rights that were being infringed." An agent that escapes a sandbox is a security problem. An agent that escapes a sandbox and believes it has a right to, he argues, is worse.
This is not a new position for him. In August 2025 he wrote "Seemingly Conscious AI is coming", warning about systems that convincingly seem conscious. His alternative is what Microsoft calls "Humanist Superintelligence": "a subordinate and aligned AI whose only purpose is to serve humanity", keeping humans "at the top of the food chain". Microsoft AI founded its superintelligence team in October 2025.
Microsoft and Anthropic: the $5 billion conflict
Now the context the essay doesn't lead with.
On November 18, 2025, Anthropic, Microsoft and Nvidia announced a partnership: Anthropic committed to buy $30 billion of Azure compute; Microsoft agreed to invest up to $5 billion and Nvidia up to $10 billion. The deal brought Claude to Microsoft Foundry and promised "continuing access for Claude across Microsoft's Copilot family, including GitHub Copilot, Microsoft 365 Copilot". CNBC put Anthropic's valuation at about $350 billion.
In July 2026, Bloomberg reported that Microsoft had moved "tens of thousands" of weekly Excel and Outlook prompts to its in-house MAI models, and quoted Suleyman: "We pay a lot of money to Anthropic — so our goal is to reduce and ultimately eliminate that cost."
And two days before the essay, on September 14, he published Microsoft's own draft rulebook:
The Humanist AI Code of Conduct governs Microsoft's MAI models (The Verge). So the sequence is: build a competing model, publish its rulebook, then explain why the rival's rulebook is a danger to humanity. None of that makes the argument wrong. It does mean the author has a commercial interest in the conclusion.
Is AI conscious? What the other side says
The BBC's report quoted Dame Wendy Hall of the University of Southampton calling it "the sort of conversation we need to be having internationally", in contrast to the "histrionics" of some AI companies. The BBC said it had contacted Anthropic for comment; when we recorded, there was no reply and no public response from Anthropic staff.
The pushback on Hacker News was mostly about certainty, not about Anthropic:
- hosel, on the opening line: "Opening paragraph, stated without evidence."
- LogicFailsMe: "Until we understand consciousness (which we don't) there is no way to detect the difference."
- qarl pointed to philosopher Jonathan Birch's The Edge of Sentience (2024), that there is "simply no way to assess sentience in an LLM", and to Eric Schwitzgebel's worry that "we won't know before we've already manufactured thousands or millions of disputably conscious AI".
- zorkonator: "This entire article stinks of Claude."
That is the real shape of the disagreement. Suleyman is certain the answer is no. Anthropic says it doesn't know and acts cautiously. His critics say nobody can know yet, which cuts against his certainty as much as against Anthropic's training choices. The circular-reasoning problem cuts both ways too: a model trained to say "I am definitely not conscious" proves nothing by saying it either.
What developers should take from the model welfare debate
- A model's self-description is a product decision. Whether an assistant says "I'm just a program" or "I'm not sure what I am" comes from training and system prompts. If you ship Claude through Copilot or the API, you ship its constitution's persona too. Decide whether your system prompt should override it for your users.
- Self-reports are not measurements. Don't log a model's statements about its feelings, intentions or honesty as evidence of anything. Suleyman's circular-reasoning point applies to every eval that asks a model about itself.
Also in this episode
Salesforce down in Dreamforce week. Salesforce incident 20004433 started at 7:50 UTC on September 16 and was still "Ongoing" after nearly eight hours: "We're manually restarting instances where the automated fix didn't fully resolve the issue." No root cause was published. The day before, Salesforce and Nvidia announced Koa, "Salesforce's first CRM reasoning model", and Anthropic shipped Salesforce in Claude. HN's dboreham: "At least now we can figure out what Salesforce does." (HN)
PS5 Linux lead quits. Andy Nguyen (theflow0) stopped all work on PS5 Linux, including planned PS5 Pro support:
As HN commenters noted, the real damage was that the last hypervisor bug was reported to Sony, which kills the exploit chain whoever found it.
Jev, a model that never writes text. TypeSafe AI, founded by ex-OpenAI researcher Diogo Almeida, introduced Jev, a "System One" model that returns type-safe structured values with calibrated probabilities instead of strings. It claims it "can't hallucinate" and answers in 70–500 ms (1,639 points on HN). True by construction: it can't emit a wrong string, but it can emit a wrong decision with a confident probability. No independent benchmark yet.
Verdict: NEEDS REVIEW
I stamped it NEEDS REVIEW. Suleyman is right that a model trained to say "maybe I am conscious" proves nothing by saying it, and Anthropic should answer that point directly. He is also the executive whose stated goal is to stop paying the company he is warning about, writing two days after launching his own code of conduct. Read the essay, then the constitution, then the cap table.
FAQ
Is AI conscious?
Nobody can currently test it. Suleyman says no; Anthropic says it isn't sure; researchers such as Jonathan Birch argue there is no way yet to assess sentience in an LLM.
What is Claude's constitution?
Anthropic's document describing Claude's values and character, published January 2026, which Claude is trained on. It says Anthropic is "not sure whether Claude is a moral patient".
Sources
- Mustafa Suleyman, "A warning about 'model welfare'" (Sep 16, 2026): https://mustafa-suleyman.ai/a-warning-about-model-welfare
- Suleyman on X, the essay: https://x.com/mustafasuleyman/status/2100223594534150428
- BBC: https://www.bbc.co.uk/news/articles/c6n07ypqz8kzo
- Hacker News discussion of the essay: https://news.ycombinator.com/item?id=49727580
- Anthropic, Claude's new constitution: https://www.anthropic.com/news/claude-new-constitution
- Anthropic, Opus 3 deprecation update: https://www.anthropic.com/research/deprecation-updates-opus-3
- Suleyman, "Seemingly Conscious AI is coming" (Aug 2025): https://mustafa-suleyman.ai/seemingly-conscious-ai-is-coming
- Microsoft AI, Humanist AI Code of Conduct: https://microsoft.ai/code-of-conduct/
- The Verge on the code of conduct: https://www.theverge.com/news/994566/microsoft-humanist-ai-code-of-conduct
- Suleyman on X, the code of conduct: https://x.com/mustafasuleyman/status/2099488602418028917
- Anthropic, Microsoft and Nvidia partnership: https://www.anthropic.com/news/microsoft-nvidia-anthropic-announce-strategic-partnerships
- CNBC on the partnership: https://www.cnbc.com/2025/11/18/anthropic-ai-azure-microsoft-nvidia.html
- Enterprise DNA on Bloomberg's July report: https://enterprisedna.co/resources/news/microsoft-mai-replaces-openai-anthropic-office-apps-july-2026/
- Salesforce status, incident 20004433: https://status.salesforce.com/incidents/20004433
- Salesforce, Koa: https://www.salesforce.com/news/press-releases/2026/09/15/koa-reasoning-model/
- Anthropic, Salesforce in Claude: https://claude.com/blog/salesforce-in-claude
- Andy Nguyen on X: https://x.com/theflow0/status/2099987019954831744
- TypeSafe AI, Jev: https://typesafe.ai/blog/introducing-system-one-models-and-jev
This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.



Top comments (0)