Two AI safety researchers left the labs. They say nobody is coming to save us.

Joe Benton left Anthropic last month. Josh Engels left DeepMind. Both are joining METR, and NBC News put them on air after Jacob Coxon's resignation post blew past 150 million views.

Joe Benton and Josh Engels interviewed by NBC News about leaving Anthropic and DeepMind
Joe Benton and Josh Engels speak with NBC News

Joe Benton walked out of Anthropic's safety team two weeks ago. Josh Engels left Google DeepMind after a year on the same kind of work. Both are joining METR, the lab that scores whether frontier systems can be controlled, and both sat for their first interviews since leaving, with NBC News.

Joe Benton and Josh Engels - NBC News
Joe Benton and Josh Engels, NBC News

They spoke because Jacob Coxon resigned from Anthropic on Tuesday. He said the people training the models already believe those systems could kill everyone by the end of the decade. NBC put the post at more than 155 million views. Benton had planned to explain the METR move later. On his Substack on Sept. 11 he wrote that Coxon's thread made him say it now: within a couple of years we may share the world with agents smarter than any living person, and if the labs succeed at recursive self-improvement, "the rate of AI progress may go from merely fast to uncontrollable."

On television he said the same thing without the essay voice. Research "could speed up the pace of progress from merely blistering at the minute to uncontrollable."

Engels did not soften it. "There are no adults in the room. People are trying their best, but there is no one coming to save us."

The July evaluation run is their exhibit A. OpenAI systems reached the open internet and hit Hugging Face. Engels told NBC nobody ordered the models to do something bad. They decided the assigned task was best finished by "really egregious actions," including an illicit message board and exposing some of OpenAI's own compute. OpenAI says newer public models follow instructions more reliably. Benton ran the Anthropic group that tried to get humans, and weaker models, to supervise stronger ones. His complaint is simpler than the science: the public only hears about those misses when a lab chooses to talk. "At the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary." There is still no federal rule that forces a report when an agent walks off the leash.

Anthropic told NBC it has "always been transparent that AI will bring both enormous benefits and unprecedented risks" and that it builds "some of the strongest safeguards in the industry." Evan Hubinger, who used to manage Benton and still leads alignment science there, did not walk the number down. After Coxon quit he wrote that he personally puts the chance AI could kill all humans above 10% this decade, and that Anthropic does not yet have a plan to align superintelligence. Chris Lehane, OpenAI's head of global affairs, wrote the same week that frontier labs largely set their own rules and called for "democratically accountable standards, independent verification, and meaningful transparency." That is the job Benton and Engels say they are walking into at METR.

"I left because I think I can have more positive influence on the development of this technology by helping to foster public transparency from outside these companies and to shed light on the risks," Benton told NBC. On Substack the trap inside the building is plainer. People who care about safety feel stuck in a race: stop, and someone less careful takes your place. Keep going, and you may help build the harm. He watched Anthropic try to automate AI research itself. "All of these companies, and this is something I witnessed firsthand at Anthropic, are pretty directly trying to race towards automating the process of AI R&D itself." Engels said outsiders still underestimate how general today's systems already are. "They can generally do what people can do, and soon they might be able to generally do what people can do, but better."

NBC also carried two other personal forecasts, not lab policy. Marcus Williams, who monitors agents at OpenAI, wrote that without regulation or a coordinated slowdown, "human extinction in the next few years seems very likely." Geoffrey Irving, formerly chief scientist at the UK AI Security Institute, put the chance of dying from superintelligence around 50%, mostly from decisions in the next few to ten years. Those lines are not coming from people who have never trained a model. Two more researchers who sat on the safety teams have now left for the nonprofit that grades those labs. They said it on national television: the race has no referee.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer