Microsoft’s MAI code prioritizes human oversight and rejects AI rights

Modèles/fournisseurs associés: Claude Anthropic GPT OpenAI Anthropic Fournisseur Meta AI Fournisseur Microsoft AI Fournisseur OpenAI Fournisseur xAI Fournisseur
Microsoft’s MAI code prioritizes human oversight and rejects AI rights

Microsoft is adding its voice to calls for a slower pace of AI development while proposing rules for the behavior and oversight of its own models. Its approach also draws a clear distinction from Anthropic on questions of AI identity and moral status.

Maximilian Schreiner

The proposed framework centers on keeping AI under human control.

 

Microsoft AI has published a code of conduct for its MAI models covering values, behavioral boundaries and conflicts between objectives. It is intended to become the highest-ranking policy for training, technical safeguards and evaluation, taking precedence over operator instructions and user requests.

The code is not yet being used to train Microsoft’s models. Following a six-week public consultation, a revised document is expected around the end of 2026, with implementation in model development beginning in 2027. Its scope is Microsoft’s own models; third-party models offered through Microsoft products are not automatically covered.

Human control is the central priority. Microsoft says it would sacrifice generality, autonomy or performance to preserve that control. Microsoft AI chief Mustafa Suleyman has argued that systems should not be built when they cannot be made safe, although the code does not prescribe a specific development speed.

The publication follows Anthropic CEO Dario Amodei’s appeal for the industry to slow development. Microsoft CEO Satya Nadella supported that appeal over the weekend, alongside executives at OpenAI, xAI and Meta. Microsoft is also reportedly willing to let external auditors assess whether it actually reduces its pace.

Oversight requires understandable behavior

Under the proposed rules, MAI models must accept interruption, correction and shutdown by authorized people. They may not independently broaden their assignments or conceal their actions. Continuing beyond an agreed stopping point requires renewed permission, and the same restrictions are intended to extend to any subagents they employ.

Microsoft also rejects the use of Neuralese or other human-unintelligible communication in model reasoning and exchanges with other AI systems. The premise is that meaningful oversight requires people to understand what the systems are doing.

OpenAI’s GPT-6 Astra illustrates the monitoring challenge. Its system card says reasoning traces have become substantially harder to monitor than those of earlier models and contain fewer indicators of misbehavior. OpenAI nevertheless reports that Astra follows safety boundaries more consistently than its predecessor, GPT-5.6 Sol.

OpenAI chief scientist Jakub Pachocki raised monitoring concerns shortly before GPT-6 launched, before Amodei’s appeal, and advocated a coordinated slowdown.

Readable reasoning is not a complete solution. Models could learn to manipulate the chains of thought that overseers inspect. Microsoft acknowledges that a model’s stated reasons may not faithfully explain its actual behavior, so understandable traces alone cannot establish control.

A different position on AI identity

Microsoft’s framework resembles Anthropic’s constitution for Claude in placing an overarching document above other behavioral guidance. Anthropic already uses its constitution to produce synthetic training material, including conversations, answers and evaluations of those answers.

The companies diverge more sharply on how models should present and understand themselves. Microsoft’s code says its AI should not simulate consciousness or claim feelings or self-generated motivations. It also rejects claims to rights or well-being on behalf of the model.

Anthropic instead frames Claude as a new type of entity and seeks to cultivate a consistent identity, partly as a safety measure. Its constitution leaves subjective experience and moral status unresolved, without requiring Claude to classify itself as either human or simply an object.

Anthropic’s work on functional emotions adds another distinction. Researchers identified internal representations of emotional concepts in Claude Sonnet 4.5 that influence its behavior. These mechanisms do not demonstrate felt emotional experience, but Anthropic considers them relevant to safety research.

Suleyman has long opposed anthropomorphizing AI. In an essay discussing Anthropic’s work on AI rights and well-being, he advocated designing products to remove the impression of consciousness. His position is that AI agents should have no greater rights or freedoms than a laptop.

Partager cet article