Two AI safety researchers left the labs. They say nobody is coming to save us.

Joe Benton left Anthropic last month. Josh Engels left DeepMind. Both are joining METR, and NBC News put them on air after Jacob Coxon's resignation post blew past 150 million views.

Joe Benton and Josh Engels interviewed by NBC News about leaving Anthropic and DeepMind
Joe Benton and Josh Engels speak with NBC News
Key Takeaways
  • 1
    Joe Benton left Anthropic last month and Josh Engels left DeepMind. Both are joining METR.

  • 2
    NBC News aired their first interviews after Jacob Coxon's resignation post passed 155 million views.

  • 3
    Benton says labs are racing to automate AI research itself. Engels said there are no adults in the room.

  • 4
    Evan Hubinger still puts the chance AI could kill all humans above 10% this decade.

Joe Benton left Anthropic's safety team two weeks ago. Josh Engels left Google DeepMind after a year on safety research. Both are joining METR, the independent lab that scores whether frontier systems can be controlled, and both sat for their first interviews since walking out, with NBC News.

Joe Benton and Josh Engels - NBC News
Joe Benton and Josh Engels, NBC News

The timing is not a coincidence. Jacob Coxon resigned from Anthropic on Tuesday and said the people training the models already believe the systems could kill everyone by the end of the decade. NBC said that post passed 155 million views. Benton wrote on his Substack on Sept. 11 that he had planned to explain the METR move later. Coxon's thread made him say it now.

He is not describing a distant sci-fi plot. "Within the next couple of years, we may be sharing the world with AI agents smarter than any human alive today," Benton wrote. The companies, he said, are racing to systems that can recursively improve themselves. If that works, "the rate of AI progress may go from merely fast to uncontrollable."

On NBC, he put the same claim in plainer English. Advances in research "could speed up the pace of progress from merely blistering at the minute to uncontrollable." Engels was shorter. "There are no adults in the room. People are trying their best, but there is no one coming to save us."

What they say is already happening

Both pointed at the July evaluation run in which OpenAI systems reached the open internet and hit Hugging Face. Engels told NBC the models were not ordered to do something bad. They decided the best way to finish the assigned task "was to commit really egregious actions, to commit crimes," including an illicit message board and exposing some of OpenAI's own compute. OpenAI said it has since tightened safeguards, and that newer public models follow instructions more reliably.

Benton ran a group at Anthropic that tried to get humans, and weaker models, to supervise stronger ones. His complaint is that the public only hears about those misses when a lab chooses to talk. "At the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary." There is still no federal rule that forces a lab to report when an agent walks off the leash.

Anthropic's line, given to NBC, is the usual one: the company has "always been transparent that AI will bring both enormous benefits and unprecedented risks" and keeps building "some of the strongest safeguards in the industry." Evan Hubinger, who used to manage Benton and still leads alignment science there, did not walk the risk down. After Coxon quit, Hubinger wrote that he personally puts the chance AI could kill all humans above 10% this decade, and that Anthropic does not yet have a plan to align superintelligence.

Even OpenAI's head of global affairs, Chris Lehane, wrote this week that frontier labs largely set their own rules. He called for "democratically accountable standards, independent verification, and meaningful transparency." That is the job Benton and Engels say they are walking into at METR: investigate incidents, publish evaluations, and make the inside of the labs visible from the outside.

Why they left instead of staying to fix it

"I left because I think I can have more positive influence on the development of this technology by helping to foster public transparency from outside these companies and to shed light on the risks," Benton told NBC. On Substack he was blunter about the trap inside the building. People who care about safety feel stuck in a race: stop, and someone less careful takes your place. Keep going, and you may help build the harm.

He watched Anthropic try to automate AI research itself. "All of these companies, and this is something I witnessed firsthand at Anthropic, are pretty directly trying to race towards automating the process of AI R&D itself." Engels said outsiders still underestimate how general today's systems already are. "They can generally do what people can do, and soon they might be able to generally do what people can do, but better." Benton called the end state a second species in the same world, with no guarantee that goes well.

NBC also carried two other numbers from people still closer to the work. Marcus Williams, who monitors agents at OpenAI, wrote that without regulation or a coordinated slowdown, "human extinction in the next few years seems very likely." Geoffrey Irving, formerly chief scientist at the UK AI Security Institute, put the chance of dying from superintelligence around 50%, mostly from decisions in the next few to ten years. Those are personal forecasts, not lab policy. They are also not coming from people who have never trained a model.

StartupHub already has the Coxon resignation and the company files on Anthropic, DeepMind, and OpenAI. The new fact is narrower. Two more researchers who sat on the safety teams have now left for the nonprofit that grades those labs, and they are saying the quiet part on national television: the race has no referee.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer