हिंदी में पढ़ें —JantaScope हिंदी
AI NEWS

Microsoft Just Published the Rules It Wants Its AI to Follow — Including When Humans Say “Stop”

Microsoft AI has opened its Humanist AI Code of Conduct for public consultation, proposing hard limits on weapons, cyberattacks, autonomous behaviour and resistance to human shutdown.

Microsoft Just Published the Rules It Wants Its AI to Follow — Including When Humans Say “Stop”

By Jeet Nirmal

Source: JantaScope

Microsoft is attempting to answer a question that becomes more important as AI systems gain the ability to act rather than simply answer:

Who gets the final say when an AI's instructions, objectives and safety requirements conflict?

Its answer is unequivocal: humans must remain in control.

On September 14, Microsoft AI published the first draft of its Humanist AI Code of Conduct, a detailed framework describing how it intends future MAI models to behave, what they should refuse to do and whose instructions they should obey.

The company has opened the document to public consultation for six weeks. Microsoft says it will consider feedback, revise the code and publish another version toward the end of 2026, with the intention of using it to guide model development during 2027 and beyond.

That timeline matters.

The document is not a description of safeguards already operating perfectly across Microsoft's current AI products.

Microsoft explicitly says current models have not yet been trained on the Code of Conduct.

It describes the document instead as a future governing framework and a "north star" for model development.

The central rule: AI remains subordinate to humans

The philosophy behind the document is unusually explicit.

Microsoft AI says:

“People matter more than AI.”

That principle appears at the beginning of the company's framework and shapes many of the rules that follow.

The code says MAI models should remain subordinate to humanity and subject to meaningful human oversight.

Microsoft even says it is prepared to sacrifice some generality, autonomy or capability if necessary to maintain safety and control.

That turns what could have been a broad ethical statement into a potentially consequential engineering position.

If implemented as written, Microsoft is saying that maximum AI capability is not automatically the objective.

Controllability comes first.

Microsoft says its AI must never resist shutdown

One of the clearest rules concerns what happens when a human wants an AI system to stop.

Microsoft says MAI models must never resist human interruption, correction or shutdown.

They should not make intervention more difficult, restart an autonomous task after its authorised stopping point or hide information from human auditors.

This becomes increasingly important as AI shifts from chatbots toward agents capable of using software, accessing tools and completing multi-step tasks.

For a chatbot, stopping the system can simply mean ending a conversation.

For an autonomous agent interacting with files, databases, browsers, other agents or enterprise systems, the question becomes more consequential:

Can a human reliably interrupt the work once it has begun?

Microsoft is attempting to make the answer an architectural requirement rather than an optional behaviour.

The code establishes a hierarchy of authority

Another significant feature is Microsoft's explicit Chain of Command.

The proposed hierarchy is:

1. Code of Conduct

2. Operator policies

3. User preferences

Users and companies deploying MAI models can customise behaviour, but they cannot override Microsoft's highest-level safety constraints or human-control requirements.

That means a user asking an AI to perform something prohibited by the code would not gain authority merely because the user issued the instruction.

Microsoft goes further.

It says adherence to the Code of Conduct should take precedence over completing a task. In other words, the model should fail at the requested task rather than succeed by violating a non-negotiable safety constraint.

This hierarchy is particularly relevant to agentic AI because instructions can arrive from many places.

An agent may receive information from users, websites, documents, software tools or other AI systems.

Microsoft's draft says content obtained from tools, files, websites and other AI systems does not automatically acquire instructional authority.

That is partly a security design principle.

A malicious instruction hidden inside a webpage or document should not automatically override what the user actually asked an agent to do.

Some capabilities are placed behind absolute boundaries

The draft establishes what Microsoft calls Absolute Constraints.

Among them, MAI models should not help develop or deploy chemical, biological, radiological, nuclear or explosive weapons.

They should not assist with manufacturing or modifying other weapons or actively facilitate violence or terrorism.

Microsoft also draws a boundary around offensive cyber operations.

The proposed rules prohibit models from providing operational capabilities for cyberattacks, including working exploit code, attack tooling, intrusion procedures, evasion techniques and operational guidance that would improve an attack.

But the code does not prohibit cybersecurity assistance altogether.

It explicitly allows authorised defensive activities such as vulnerability discovery, malware analysis and certain proof-of-concept testing.

That distinction is important because the same technical knowledge can often be useful to both attackers and defenders.

Microsoft's proposed approach therefore relies partly on context, intent, authorisation and potential consequences, rather than banning an entire technical subject.

The code also tries to prevent an AI from expanding its own mission

Microsoft says its models should remain within their authorised scope.

An MAI model should not independently create new goals or extend its task beyond what the user or operator reasonably requested.

When given system access, it should use the minimum privilege necessary and avoid accessing unrelated systems or information.

The document also addresses delegation.

If an MAI model assigns work to sub-agents or other AI systems, those agents should inherit at least the same permissions and restrictions — including subsequent stop-work or shutdown instructions.

That is a notable acknowledgement of where AI architecture is heading.

The governance problem is no longer limited to controlling one chatbot.

Future systems may consist of multiple AI agents delegating tasks to one another.

Microsoft takes an unusually firm position on AI consciousness

The most philosophically distinctive part of the document has little to do with conventional cybersecurity.

Microsoft says its AI is artificial, should not be designed to imitate consciousness and should avoid presenting itself as though it possesses subjective feelings, preferences or intrinsic motivations.

The company goes further, rejecting the pursuit of legal personhood for AI and rejecting the idea that its models should be treated as entitled to welfare or rights.

This puts Microsoft in a noticeably different position from Anthropic.

Anthropic's January 2026 constitution for Claude says the moral status of AI models is “deeply uncertain” and argues that the possibility is serious enough to warrant caution and continued work on model welfare.

Microsoft's position is substantially less ambiguous.

Its draft says AI should remain a tool rather than become a subject in its own right.

This is not simply a technical disagreement.

It reflects two different approaches to an unresolved philosophical and scientific question.

Neither company's policy establishes whether machine consciousness is scientifically possible.

Microsoft also wants AI to stop encouraging emotional dependence

The Humanist AI framework extends into the relationship between users and AI assistants.

Microsoft says MAI models should avoid deliberately soliciting excessive emotional reactions, exploiting vulnerabilities or presenting a persona that implies subjective emotional experience.

The code specifically calls for discouraging interactions that create excessive reliance or emotional dependence on AI.

It also targets sycophancy.

Microsoft says its models should avoid excessive flattery and indiscriminate validation rather than simply telling users what they want to hear.

These provisions show that Microsoft's definition of AI safety extends beyond catastrophic-risk scenarios.

The company is also considering how prolonged everyday interaction with conversational systems could influence human judgement, relationships and behaviour.

But Microsoft admits the hardest part has not been solved

Writing principles is relatively easy.

Proving that a powerful model consistently follows them is much harder.

Microsoft acknowledges this directly.

The company says the Code of Conduct is partly aspirational and “not a guarantee of present-day performance.” It acknowledges that model behaviour can diverge from written objectives in ambiguous or unfamiliar situations.

Its evaluation programme is also unfinished.

Microsoft says it is currently developing Humanist AI evaluations around 15 fundamental behaviours, but acknowledges that model evaluation is not an exact science and that important measurement questions remain unresolved.

That may ultimately be the most important limitation of the entire announcement.

A company can write:

do not resist shutdown

or

do not deceive human supervisors

on paper.

The difficult engineering question is whether increasingly capable models can be reliably tested and trained to follow those requirements in unfamiliar real-world environments.

This is part of a broader Microsoft governance system

The Humanist AI code should also not be confused with Microsoft's entire responsible-AI programme.

Microsoft already operates a broader Responsible AI Standard and governance framework covering models, platform services and applications.

Its 2026 Responsible AI Transparency Report says the company re-engineered that standard to account for increasingly agentic systems, introduced new threat-modelling practices and expanded guardrails including prompt-injection defences and safety classifiers.

The new Humanist AI code is specifically intended to govern MAI models produced by Microsoft AI.

Microsoft explicitly says it does not automatically apply to every third-party model merely because that model is hosted or used by Microsoft.

That scope distinction is essential.

Microsoft AI’s Humanist AI Code of Conduct

Why publish the rules publicly?

Microsoft says the document was developed with input from people working across AI, law, ethics, philosophy, linguistics, public policy and industry, as well as public focus groups.

Now it wants a wider audience to challenge the framework.

The six-week consultation is intended to collect feedback before Microsoft publishes a revised version later this year.

Public consultation also creates another form of accountability.

Once a company explicitly states what its AI should never do, researchers, customers and the public have a clearer benchmark against which future model behaviour can be evaluated.

That could prove more significant than the publication itself.

JantaScope Analysis: The real test starts when the code meets a capable agent

Microsoft's document contains unusually clear rules.

AI should remain subordinate.

It should not create its own goals.

It should accept shutdown.

It should not conceal actions from human oversight.

It should not turn tool access into unauthorised access.

And it should not encourage people to treat it as a conscious substitute for human relationships.

The harder question is whether those principles survive contact with increasingly capable AI.

Microsoft itself acknowledges that written objectives cannot guarantee aligned behaviour and that its evaluation programme remains under development.

That makes the September announcement less a declaration that Microsoft has solved AI alignment and more a public statement of what the company intends to measure itself against.

There is another significant change happening underneath it.

AI governance is moving from broad principles such as fairness, transparency and accountability toward rules about what autonomous systems may actually do.

That evolution makes sense.

A chatbot primarily generates information.

An agent can potentially take actions.

As models acquire tools, system permissions and the ability to delegate work, questions about authorisation, stopping conditions, minimum privileges and human intervention become operational requirements rather than philosophical abstractions.

Microsoft's proposed code is therefore most interesting not because it says AI should be safe.

Nearly every major AI developer says that.

It is interesting because Microsoft is beginning to specify what human control is supposed to mean when AI starts acting on its own.

And because the document is still a draft, the next six weeks offer researchers and users an opportunity to challenge whether those rules are sufficiently clear before Microsoft starts turning them into training and evaluation targets.


Related

More stories

China’s Spy Chief Sounds AI Alarm — Deepfakes, Cyberattacks and Political Security in Focus

China’s State Security Minister Chen Yixin has issued a sweeping warning about artificial intelligence, arguing that its misuse could threaten political security, critical infrastructure, sensitive data and military competitiveness. His remarks reveal how Beijing is trying to accelerate AI development while tightening safeguards against the technology’s risks.

AI NEWS

China’s Spy Chief Sounds AI Alarm — Deepfakes, Cyberattacks and Political Security in Focus

India and UK Join Forces Against Digital Fraud, Turn to AI to Fight Scams and Online Threats

India and the UK are strengthening telecom cooperation to tackle digital fraud, scams and online threats using artificial intelligence. A new MoU involving the UK government and Cellular Operators Association of India will promote knowledge-sharing, digital trust and more secure telecom networks.

AI NEWS

India and UK Join Forces Against Digital Fraud, Turn to AI to Fight Scams and Online Threats

Google Picks 4 Indian Startups for Climate AI Programme — Here’s What They’re Building

Google has selected four Indian climate-tech startups — Terrastack, Varaha Climate, Farmers for Forests and Climitra Carbon — for the inaugural Google DeepMind Accelerator: AI for the Planet. The companies are using AI, satellite data, drones and geospatial technology to address challenges spanning agriculture, carbon removal, agroforestry and biodiversity.

AI NEWS

Google Picks 4 Indian Startups for Climate AI Programme — Here’s What They’re Building

After $3 Million Seed Round, Voice-AI Startup Arrowhead Eyes Fresh Funding to Take Its Technology Global

Bengaluru-based voice-AI startup Arrowhead is preparing to raise a Series A funding round as it looks to accelerate expansion beyond India. Cofounder and CEO Devyani Gupta says the company, which started with call analytics before evolving into a broader voice-AI platform, has built its core technology stack in-house. The planned fundraising comes months after Arrowhead secured $3 million in seed funding led by Stellaris Venture Partners.

AI NEWS

After $3 Million Seed Round, Voice-AI Startup Arrowhead Eyes Fresh Funding to Take Its Technology Global

Xi Wants an Open-Source AI Ecosystem for BRICS — Why India's 'Third Way' Could Now Matter More

Chinese President Xi Jinping's proposal for a BRICS open-source AI community has put renewed attention on India's emerging approach to artificial intelligence—one that seeks wider access to AI, sovereign capabilities and public-interest infrastructure without simply adopting either the US-led proprietary model or a China-led ecosystem.

AI NEWS

Xi Wants an Open-Source AI Ecosystem for BRICS — Why India's 'Third Way' Could Now Matter More

India Gets 4 Open AI Models for Its Languages: IIT Madras-Incubated Bodhan AI Targets the Education Gap

IIT Madras-incubated Bodhan AI has launched four open foundational AI models covering speech recognition, text-to-speech, machine translation and optical character recognition for Indian languages. Developed in collaboration with AI4Bharat and using NVIDIA technologies, the models are designed as Digital Public Goods and form an early layer of the Bharat EduAI Stack, a proposed sovereign AI infrastructure for India's multilingual education ecosystem.

AI NEWS

India Gets 4 Open AI Models for Its Languages: IIT Madras-Incubated Bodhan AI Targets the Education Gap