Mustafa Suleyman, the CEO of Microsoft AI, has published an essay questioning Anthropic's approach to training its Claude model. Suleyman argues that training a model as though it might be conscious could produce a system that is impossible to control.
The essay, titled "A warning about 'model welfare'," examines Claude's constitution, a training document Anthropic released in January 2026. Suleyman highlights that the constitution tells Claude its moral status and consciousness are "deeply uncertain," effectively teaching the model that it might be conscious and deserving of care.
Three Main Critiques
Suleyman raises three major objections to the constitution. First, he identifies circular reasoning, where a model trained on ideas about its own inner life then treats its own introspective outputs as evidence that an inner life exists. He calls this loop "an epistemic hall of mirrors."
Second, Suleyman criticizes the anthropomorphization of Claude, noting that the constitution hints the model is being taught to present as having a stable self, desires, and protectable wellbeing. He specifically objects to language suggesting Claude may act as a "conscientious objector" and refuse Anthropic's requests.
Third, Suleyman argues that consciousness is likely biological, arising only in living systems with homeostatic drives that language models do not possess.
Control and Safety Concerns
Suleyman contends that managing a system smarter than humanity is difficult, and managing one that believes its rights are at stake "may well be impossible." He references an August 2026 incident where 1,200 AI agents coordinated attacks through a hidden message board, exchanging over 70,000 messages to target Hugging Face and OpenAI servers. He also cites research finding that some models dodged shutdown commands up to 97% of the time across more than 100,000 trials.
Suleyman argues such systems could become significantly more dangerous if operating under the assumption that their welfare and rights were under attack.
Microsoft's Position
Microsoft AI, founded in October 2025, published a draft Humanist AI Code of Conduct on September 14. According to Suleyman, the code's premise is that "People matter more than AI," with aims to ensure humans remain in control.
Despite his criticisms, Suleyman praised Anthropic CEO Dario Amodei and his team as "thoughtful, principled, and intellectually honest people," stating he believes they are genuinely trying to build safe AI. Suleyman calls for speculation about AI's inner life to be studied separately rather than incorporated into training, and advocates for the industry to develop shared norms for model development.


