# Beyond RLHF: The Future of AI Automation _Former OpenAI researcher Diogo Almeida argues that the current RLHF-based AI era is limited to assistance, and true automation requires a new approach._ **Published:** 2026-08-01 **Source:** https://www.startuphub.ai/ai-news/ai-research/2026/beyond-rlhf-the-future-of-ai-automation --- Diogo Almeida, a former key figure at OpenAI and co-author of influential papers on GPT-4 and ChatGPT, presented a compelling argument at the AI Engineer World's Fair, suggesting that the AI industry is poised to move beyond the era defined by RLHF. Almeida, now leading TypeSafe AI, posited that while RLHF has been instrumental in creating conversational AI assistants like ChatGPT, its inherent design for human preference optimization limits its utility for true automation. RLHF DilemmaDriver current AI excels at assistance, struggles with true automation tasksFrom the article 8 mentionsAlmeida, now leading TypeSafe AI, posited that while RLHF has been instrumental in creating conversational AI assistants like ChatGPT, its inherent design for human preference optimization limits its utility for true automation.Human Preference LimitsDriverRLHF optimizes for human feedback, not autonomous task completionFrom the article 4 mentionsAlmeida, now leading TypeSafe AI, posited that while RLHF has been instrumental in creating conversational AI assistants like ChatGPT, its inherent design for human preference optimization limits its utility for true automation.Current AI: AssistanceContextRLHF-based models like ChatGPT are good conversational assistantsFrom the article 6 mentionsHe elaborated that RLHF's primary goal is to please the human user, which is ideal for assistance but not for autonomous tasks where error-free execution is paramount.leads toDiogo Almeida's ArgumentCoreformer OpenAI researcher advocates moving beyond RLHF for automationFrom the articleDiogo Almeida, a former key figure at OpenAI and co-author of influential papers on GPT-4 and ChatGPT, presented a compelling argument at the AI Engineer World's Fair, suggesting that the AI industry is poised to move beyond the era defined by RLHF.advocatesFuture: True AutomationEffectrequires a new approach beyond human preference optimizationFrom the article 3 mentionsLike, why are the like, the building blocks of software actually still the same?" He proposed that the AI industry should focus on redesigning the AI stack for reliability and automation, moving towards a future where AI can handle complex, rote tasks autonomously.Smarter SoftwareContextAI needs to perform complex tasks without constant human oversightFrom the article 4 mentionsHe cited Garry Tan's phrase, "We're entering the golden age of just-in-time software," but cautioned that this could be a double-edged sword, leading to more software generation rather than inherently smarter software.TypeSafe AI's VisionOutcomeAlmeida's company aims to build AI for genuine autonomous automationFrom the article 3 mentionsAlmeida revealed that his current venture, TypeSafe AI, is working on this very problem. ## The RLHF Dilemma: Assistance vs. Automation Almeida opened his talk by acknowledging the rapid advancements in AI, noting how benchmarks are consistently surpassed. However, he highlighted a dichotomy in current AI applications: tasks that are 'too good to be true' (like instruction following and chatbots) versus those that are 'too bad to be useful' (like customer service requiring human oversight or data entry). He argued that the simplest explanation for this divide lies in the core design of RLHF. "Today's AI, everything inherited from RLHF, is incredible at the human in the loop stuff, but not for automation tasks," Almeida stated. He elaborated that RLHF's primary goal is to please the human user, which is ideal for assistance but not for autonomous tasks where error-free execution is paramount. The overpromising nature of current AI, he contended, is a feature of RLHF, designed to optimize for engagement rather than calibrated performance. ## The Limitations of Human Preference Optimization Almeida explained that RLHF works by collecting human preferences and then optimizing models based on those preferences. This process, while effective for creating helpful and engaging AI assistants, inherently leads to a gap between human preference and actual task results. He illustrated this with an anecdote about sending ChatGPT fart sound effects and asking for a musical critique, where the AI provided an elaborate, albeit nonsensical, response, demonstrating its bias towards generating a human-pleasing output. He emphasized that the goal for automation is not to mimic human preferences but to execute tasks correctly and reliably, ideally operating in the background without human intervention. This fundamental difference, he argued, explains why current AI is adept at conversational tasks but falters in domains requiring precision and autonomy. ## The Future: True Automation and Smarter Software Almeida’s core thesis is that the next frontier for AI is true automation, which will require a new approach beyond RLHF. He suggested that the current SaaS landscape, despite AI advancements, has seen little fundamental change, with chatbots often being a superficial addition. He cited Garry Tan's phrase, "We're entering the golden age of just-in-time software," but cautioned that this could be a double-edged sword, leading to more software generation rather than inherently smarter software. "What I want is smarter software," Almeida declared. "Why can't like B2B SaaS just be more expressive? Like, why are the like, the building blocks of software actually still the same?" He proposed that the AI industry should focus on redesigning the AI stack for reliability and automation, moving towards a future where AI can handle complex, rote tasks autonomously. ## TypeSafe AI's Vision Almeida revealed that his current venture, TypeSafe AI, is working on this very problem. Their core question revolves around what would happen if the AI stack were redesigned for reliability and automation. He expressed excitement about this work, stating, "I think it's one of the most satisfying things I've worked on." He also hinted at an upcoming announcement, noting that "the original scaling laws were incorrect." The talk concluded with a call to action for those interested in building smarter software to sign up for TypeSafe's mailing list or careers page, and to follow him on Twitter for further insights. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.