OpenAI is codifying how its artificial intelligence systems should act with the introduction of its formal Model Spec. This document details the intended behavior for AI models, aiming for clarity and public debate on how these powerful tools should operate.
The Model Spec is designed to make intended AI conduct explicit, moving beyond internal training processes to a format accessible to users, developers, researchers, and policymakers. It’s not a claim of current perfection but a target for future development, guiding training, evaluation, and improvement.
This initiative is part of OpenAI's broader strategy for safe and accountable AI, complementing efforts like the Preparedness Framework which addresses risks from advanced capabilities. The ultimate goal is to foster a gradual, iterative, and democratically legible transition to advanced AI, ensuring it aligns with human interests.
The Structure of AI Demeanor
The Model Spec begins with high-level intent, clarifying OpenAI's mission-level goals: iteratively deploying empowering models, preventing serious harm, and maintaining operational license. It then details how these goals are balanced, acknowledging tradeoffs without directly instructing models to pursue abstract concepts like 'benefiting humanity' autonomously.
Central to the Spec is the 'Chain of Command,' a framework for prioritizing instructions from various sources, OpenAI, developers, and users, when conflicts arise. This hierarchy assigns authority levels to policies and instructions, ensuring safety boundaries, like preventing bomb-making requests, take precedence over user prompts for less critical behaviors.