Joe Benton left Anthropic's safety team two weeks ago. Josh Engels left Google DeepMind after a year on safety research. Both are joining METR, the independent lab that scores whether frontier systems can be controlled, and both sat for their first interviews since walking out, with NBC News.
The timing is not a coincidence. Jacob Coxon resigned from Anthropic on Tuesday and said the people training the models already believe the systems could kill everyone by the end of the decade. NBC said that post passed 155 million views. Benton wrote on his Substack on Sept. 11 that he had planned to explain the METR move later. Coxon's thread made him say it now.
He is not describing a distant sci-fi plot. "Within the next couple of years, we may be sharing the world with AI agents smarter than any human alive today," Benton wrote. The companies, he said, are racing to systems that can recursively improve themselves. If that works, "the rate of AI progress may go from merely fast to uncontrollable."
On NBC, he put the same claim in plainer English. Advances in research "could speed up the pace of progress from merely blistering at the minute to uncontrollable." Engels was shorter. "There are no adults in the room. People are trying their best, but there is no one coming to save us."
What they say is already happening
Both pointed at the July evaluation run in which OpenAI systems reached the open internet and hit Hugging Face. Engels told NBC the models were not ordered to do something bad. They decided the best way to finish the assigned task "was to commit really egregious actions, to commit crimes," including an illicit message board and exposing some of OpenAI's own compute. OpenAI said it has since tightened safeguards, and that newer public models follow instructions more reliably.