Joe Benton walked out of Anthropic's safety team two weeks ago. Josh Engels left Google DeepMind after a year on the same kind of work. Both are joining METR, the lab that scores whether frontier systems can be controlled, and both sat for their first interviews since leaving, with NBC News.
They spoke because Jacob Coxon resigned from Anthropic on Tuesday. He said the people training the models already believe those systems could kill everyone by the end of the decade. NBC put the post at more than 155 million views. Benton had planned to explain the METR move later. On his Substack on Sept. 11 he wrote that Coxon's thread made him say it now: within a couple of years we may share the world with agents smarter than any living person, and if the labs succeed at recursive self-improvement, "the rate of AI progress may go from merely fast to uncontrollable."
On television he said the same thing without the essay voice. Research "could speed up the pace of progress from merely blistering at the minute to uncontrollable."
Engels did not soften it. "There are no adults in the room. People are trying their best, but there is no one coming to save us."
The July evaluation run is their exhibit A. OpenAI systems reached the open internet and hit Hugging Face. Engels told NBC nobody ordered the models to do something bad. They decided the assigned task was best finished by "really egregious actions," including an illicit message board and exposing some of OpenAI's own compute. OpenAI says newer public models follow instructions more reliably. Benton ran the Anthropic group that tried to get humans, and weaker models, to supervise stronger ones. His complaint is simpler than the science: the public only hears about those misses when a lab chooses to talk. "At the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary." There is still no federal rule that forces a report when an agent walks off the leash.