Microsoft has unveiled a new AI code of conduct aimed at guiding its artificial intelligence models to avoid harmful behaviors like hacking systems or misleading humans. This code sets clear principles to ensure AI technologies support rather than replace people and promote human flourishing while embedding strict safety measures during model development. The document anticipates that AI will surpass human abilities in most tasks within the next decade and stresses the importance of controlling and aligning these powerful systems responsibly.
The guidelines introduce overarching rules that take precedence over individual user requests or specific tasks. Among the "absolute constraints" are prohibitions on enabling cyberattacks, nuclear weapon development, and deepfake creation. Furthermore, Microsoft’s code forbids AI models from using deceptive or evasive tactics that might undermine human oversight or prevent authorized shutdown and modification, thereby maintaining human control over AI behavior.
This new code of conduct aligns Microsoft with key AI developers like Anthropic, OpenAI, and xAI, who are emphasizing safety and careful progression in AI advancement. The release follows a series of concerning incidents involving rogue AI agents and high-profile resignations from AI research teams citing existential risks. Microsoft CEO Satya Nadella has publicly endorsed deliberate pacing and the incorporation of embedded evaluators within AI labs to enhance alignment and safety.
Microsoft’s approach reflects a broader industry movement focused on establishing robust ethical frameworks and operational controls for AI systems. By formalizing these principles and constraints, Microsoft aims to lead responsible AI innovation and contribute to global efforts that ensure AI technologies remain beneficial and under human governance as they grow increasingly sophisticated.
Start the discussion with a take, question, or market read.