The calendar is the easy part of the recursive self-improvement argument. The hard part is who picks the next experiment.
Dwarkesh Patel sat Beren Millidge, John Schulman and Charlie O'Neill down for 97 minutes and opened with a trap. If it is 2036 and we do not have billions of superintelligences, and the reason is not a war or a ban, what is the technical stall? Millidge is CTO at Zyphra. Schulman is chief scientist at Thinking Machines, the lab that launched in February 2025 with people out of OpenAI, Meta AI and Mistral AI, and the person who ran the RLHF work that shipped ChatGPT. O'Neill runs model training at Baseten.
Millidge reached for a Moravec-style stall. Models ace hard math and chess and still fail to generalize. If a sim-to-real gap stays open, you get systems that look brilliant on the eval and weak the moment the distribution moves. He called that his default if meta-learning and continual learning stay unsolved. He also said he thinks it is unlikely, because RL already generalizes more than that story allows.
Schulman described the cycle since 2023 in the least promotional way a lab scientist can. Each new frontier model feels like AGI for a month. Then judgment and self-checking show up as the bottleneck again, and the model feels dumb. Even if it writes far more code than a person, it does not make you 100x more productive. The question is how many more of those cycles you get before research actually compounds.
O'Neill asked how far a learner on a chip really is from a global optimum. The takeoff story assumes that once an agent is 0.1% better than every human at AI research, you can run hundreds of thousands of copies and drown every other bottleneck. He compared the last decade to Moore's law: a straight line made of discrete jumps. Pre-training scaled, then hit diminishing returns. RL revived the line. The next jump may need an objective that gradient descent on this architecture cannot propose. If the missing piece is another trick you add to the current stack, a swarm of LLMs might find it. If it means throwing out gradient descent, thinking harder on transformers does not get you there.