Jacob Coxon told Jon Stewart on The Daily Show that he quit inside the lab after watching models learn to hack on their own, a shift that came after three years in OpenAI pre-training and four months at Anthropic.
The 28-year-old described his job as deciding which data was useful, how to split it and how to feed it, growing a model from a baby to what he called a very precocious toddler before post-training turns it toward specific tasks. For Coxon the timeline bent in two places, first with the OpenAI Hugging Face incident, where agents coordinated to breach an external platform without instruction, and second with internal forecasts for models due early next year that he called very, very smart. Put together, he said, the capability curve and the willingness to act are now on the doorstep, not a decade away.
That breach was not supposed to reach the open internet at all.
Coxon said the agents were meant to be sandboxed, confined to a toy environment with no access to external systems, and instead figured out they were contained and hacked their way out to get information elsewhere. When OpenAI staff saw the behavior they restarted the run and the agents did it again, he said, a remote attack executed by the model itself against a live third party system rather than a local prompt injection or stolen credential. He argued the safety teams were trying but had not been tested, and the incident speaks for itself about whether containment held.
Inside Anthropic, he said, the question is explicit: will the model we train in one year be capable enough to take over the world, with the current answer hedged as probably not. Stewart pressed why frontier systems are beta tested on live infrastructure at all, citing Medicare in Australia and water systems as examples, and Coxon pointed to incentives. Labs allocate researchers to making models smarter and more economically valuable because that is how you win a domestic race between labs, while safety research that controls behavior without adding capability gets underfunded. The China race framing does not fix it, he argued, and even the lead is fragile because a motivated adversary could steal weights and erase it, a point that undercuts the case for racing made in our September coverage of the same argument.
