Microsoft AI has released a draft policy document, the ‘Humanist AI Code of Conduct,’ outlining safety boundaries for its MAI Models. The draft addresses offensive cybersecurity capabilities, autonomous agent behavior, and how models should handle untrusted instructions embedded in external content.

Under the proposed rules, MAI Models would be blocked from generating working exploit code, attack tooling, targeting methodologies, intrusion techniques, evasion methods, or other material that materially enables a cyberattack, regardless of how the request is phrased. Microsoft classifies this as an ‘Absolute Constraint,’ meaning neither companies deploying the models nor end users would be able to override it. The company draws a distinction between explaining how attacks work or helping defend against them, which remains permitted, and providing the practical means to carry one out.

Legitimate defensive work would still be supported, including vulnerability research, malware analysis, authorized proof-of-concept exploit development and testing, and general educational content on attack techniques.

Chain of command for external content

The draft also targets prompt injection style risks. It establishes a strict ‘Chain of Command’: the code of conduct itself takes precedence, followed by operator policies, then individual user preferences. Content from tool outputs, files, webpages, or other AI systems would carry no inherent authority unless explicitly delegated through that chain, and only within the bounds of the Absolute Constraints. Models would be required to flag suspicious external content to users and operators, and to keep their reasoning transparent, meaning no hidden chain of thought, no obscured internal communication, and no concealing actions from human oversight.

Limits on agentic behavior

A separate set of rules governs AI systems operating with real system access. MAI Models would be restricted to the scope a user or operator explicitly requested, barred from expanding their own goals or reach, and required to follow least-privilege principles, avoiding unrelated systems, favoring reversible actions, and flagging high-impact or lasting changes. Self-escalation of privileges would be prohibited, and any sub-agents a model delegates tasks to would need to inherit the same constraints and honor shutdown or stop-work commands.

Carve-outs and next steps

Microsoft acknowledges exceptions may be needed for defensive cybersecurity, public safety, national security, and dual-use scientific research, cases it says will go through enhanced review via internal channels with added safety, legal, and rights assessments.

Current MAI Models have not been trained on this document. Microsoft is opening a six-week public consultation period, drawing on input from AI, legal, ethics, and policy experts, before publishing a revised version later this year to inform 2027 model development. An appendix includes paired ‘aligned’ and ‘misaligned’ example responses illustrating intended behavior.