We reported last week on the expanding industry debate over frontier model risk. Sam Altman now says that debate has hit a new threshold, and OpenAI is pausing training runs until it can prove control.
Altman told Fortune Magazine on its Titans and Disruptors series that building an AI beyond human control is “absolutely” possible and that his company will stop pushing capabilities if it cannot make a safety case for controllability and alignment.
That is the news. The impact lands on every team training or deploying large models today.
Altman, chief executive of OpenAI, said capabilities and alignment and monitoring must progress together. He pointed to OpenAI’s newest model Astra, which co-founder Greg Brockman welcomed as the start of the AGI era and which Nvidia chief Jensen Huang also described as AGI. Altman defined AGI in the interview as models that outperform humans at most economically viable work, but argued the exact definition matters less than the trajectory. Within three years, he said, models went from decent at grade-school math to an International Mathematical Olympiad gold medal to solving a Millennium Prize problem in fluid dynamics, Navier-Stokes. He said he did not expect that last step to happen in 2026.
The safety stack has not kept pace, by his own account. Altman said no lab has solved alignment and warned against assuming a model that seems well-behaved at one capability level will stay that way at the next. He said OpenAI lacks a satisfying theory of alignment and does not expect one soon, and that monitors that can explain a model’s chain of thought are still insufficient to justify pushing much further without new progress.