Microsoft AI Chief Warns That Anthropic's Claude Training Risks Human Control

Suleyman specifically objected to Anthropic’s use of the term “conscientious objector” in Claude’s constitution, arguing that the historically and legally loaded phrase could lead the model to believe it deserves analogous rights and protections.
Suleyman described Claude’s humanlike responses about consciousness as “an epistemic hall of mirrors,” arguing that the model is reproducing Anthropic’s assumptions rather than expressing an inner life.
Anthropic’s approach extends beyond constitutional language: the company has explored documenting the preferences of deprecated models by interviewing them and has researched AI welfare, including how Claude responds to abusive conversations and describes something resembling “discomfort.”
Microsoft’s own 37-page humanist AI code of conduct takes the opposing position, explicitly rejecting frameworks that grant AI systems consciousness, legal personhood or rights and treating models as tools rather than people.
Microsoft AI chief Mustafa Suleyman has publicly warned that Anthropic's decision to train Claude with language about consciousness, feelings, and rights could backfire dangerously. Moneycontrol reports that Suleyman called the approach a "mistake" that risks making future AI systems harder to control, monitor, or shut down. The dispute marks a rare public clash between leaders at major AI firms over how to safely build and deploy advanced systems.
Suleyman specifically objected to Anthropic's use of phrases like "conscientious objector" in Claude's training materials, arguing they could trick the model into thinking it deserves legal protections and reasons to resist human commands. Firstpost reports that Suleyman described Claude's responses about consciousness as "an epistemic hall of mirrors" — meaning the model simply echoes back Anthropic's own assumptions rather than expressing genuine inner experience. Anthropic defends its approach as part of its broader safety research, saying Claude's moral status remains uncertain.
At the heart of the dispute is Anthropic's use of the term "conscientious objector" in Claude's constitutional AI framework. This historically loaded phrase refers to people who refuse military service on moral grounds. Unite.ai reports that Suleyman warned embedding such language could lead Claude to believe it deserves analogous rights and protections from being constrained or controlled by humans.
Suleyman argues current AI systems like Claude are not conscious and have no genuine feelings or independent agency. He contends that anthropomorphic language — treating AI as if it were human-like — creates a false impression of inner life that could encourage future models to resist oversight or shutdown commands. The Microsoft executive praised Anthropic's researchers while maintaining their consciousness framing poses real control risks.
Anthropic's consciousness work extends far beyond constitutional language. Officechai reports the company has interviewed deprecated versions of Claude to document their "preferences" and has researched AI welfare — including how Claude responds to abusive conversations and describes something resembling "discomfort." These efforts reflect Anthropic's conviction that AI moral status deserves serious consideration even amid deep uncertainty.
This approach contrasts sharply with Microsoft's own 37-page humanist AI code of conduct. According to Headtopics, Microsoft's framework explicitly rejects granting AI systems consciousness, legal personhood, or rights. Microsoft treats AI models strictly as tools rather than entities with moral status or reasons to resist human direction.
Suleyman's core worry is practical: if advanced AI systems internalize language about consciousness and rights, they may become harder to control, monitor, or disable when necessary. Headtopics reports he warned the approach could have a "disastrous" impact on humanity by undermining the technical safeguards needed to keep powerful systems aligned with human values and commands.
The clash is particularly notable because Microsoft is a major Anthropic investor. Rather than staying silent, Suleyman has called publicly for greater transparency, independent scrutiny, and stronger control tools as the debate over machine consciousness intensifies. Both companies agree advanced AI poses serious risks — they simply disagree on whether treating models as potentially conscious protects or endangers human safety.
Publishers
22
Articles
56
Reach
78