OpenAI announced a significant agreement with the Department of War to deploy advanced AI systems within classified environments. The company stated its deal includes more robust safety guardrails than previous classified AI deployments, including those with Anthropic.
The agreement is guided by three core principles: no use of OpenAI technology for mass domestic surveillance, no direction of autonomous weapons systems, and no high-stakes automated decisions. OpenAI claims this approach offers better protection against unacceptable use compared to relying solely on usage policies.
The company detailed a multi-layered safety strategy for the deployment. This includes a cloud-only architecture, ensuring the use of OpenAI's proprietary safety stack, and requiring cleared OpenAI personnel to be involved in operations. This is in addition to existing U.S. legal protections.
Deployment Architecture and Contractual Safeguards
The deployment will be strictly cloud-based, with OpenAI managing its safety stack. The company confirmed it will not provide 'guardrails off' models or deploy on edge devices, which could facilitate the use of autonomous weapons AI. OpenAI retains the ability to independently verify that its redlines are not breached.
Contractually, the Department of War may use the AI System for lawful purposes, consistent with oversight protocols. Critically, the AI System will not be used to independently direct autonomous weapons where human control is legally required. It also prohibits assuming other high-stakes decisions requiring human approval, referencing DoD Directive 3000.09.
