AI Pioneers Debate Transformer's Future, Urge New Architectures

AI pioneers Jerry Tworek and Rohan Anil discuss the limitations of current Transformer architectures and the need for new models that can learn from real-world experience.

Jerry Tworek and Rohan Anil discussing AI architecture
Sequoia Capital
Visual TL;DR
AI pioneers debateCore
Jerry Tworek and Rohan Anil discuss current AI development trajectory
Transformer architecture limitsDriver
From the article 4 mentionsSpeaking on a recent podcast, the duo, who boast an impressive pedigree from leading AI research labs like OpenAI, Gemini, and Google Brain, suggested that the Transformer architecture, while foundational to recent AI advancements, may be approaching its limits.
Core Animation foundersCore
Tworek and Anil from OpenAI, Gemini, Google Brain backgrounds
From the article 2 mentionsIn a candid discussion that offers a glimpse into the future of artificial intelligence, the founders of Core Animation, Jerry Tworek and Rohan Anil, have articulated a contrarian view on the current trajectory of AI development.
Architecture is bottleneckDriver
the architecture itself is now the primary constraint to smarter systems
From the article 7 mentions"I think at this moment, what the bottleneck is to better models and to smarter systems is the architecture itself," Tworek stated.
Need new architecturesEffect
urge new models that can learn from real-world experience
From the article 9 mentionsThey emphasized the need for a baseline level of computational power to even reveal the capabilities of new architectures, a hurdle that often prevents innovative ideas from emerging.
Beyond static dataContext
models need to learn at test time, not just from pre-training
Future AI advancementsOutcome
new architectures will enable more capable and intelligent AI systems
From the article 2 mentionsIn a candid discussion that offers a glimpse into the future of artificial intelligence, the founders of Core Animation, Jerry Tworek and Rohan Anil, have articulated a contrarian view on the current trajectory of AI development.
Contents(4)

In a candid discussion that offers a glimpse into the future of artificial intelligence, the founders of Core Animation, Jerry Tworek and Rohan Anil, have articulated a contrarian view on the current trajectory of AI development. Speaking on a recent podcast, the duo, who boast an impressive pedigree from leading AI research labs like OpenAI, Gemini, and Google Brain, suggested that the Transformer architecture, while foundational to recent AI advancements, may be approaching its limits.

AI Pioneers Debate Transformer's Future, Urge New Architectures - Sequoia Capital
AI Pioneers Debate Transformer's Future, Urge New Architectures, Sequoia Capital

The Bottleneck of Architecture

Tworek, formerly VP at OpenAI, expressed his belief that the AI community has become exceptionally adept at training large models using established methods like pre-training and reinforcement learning at scale. However, he posited that the architecture itself is now the primary constraint. "I think at this moment, what the bottleneck is to better models and to smarter systems is the architecture itself," Tworek stated. He criticized the prevalent trend of optimizing existing Transformer models for cost and efficiency, arguing for a shift towards enhancing their fundamental capabilities and expressiveness.

Anil, a former pre-training lead at Gemini and a key figure in AI research at Google Brain, echoed this sentiment, drawing parallels to human learning. He contrasted the iterative, trial-and-error nature of learning through experience, akin to playing football, with the deep conceptual understanding required in mathematics. Both, he noted, are forms of learning from experience, but vastly different.

Beyond Static Data: Learning at Test Time

A significant point of discussion revolved around the limitations of current models in adapting to the messy and dynamic nature of the real world. Tworek shared his personal disappointment stemming from the realization that scaling up reinforcement learning, which he once believed was the direct path to Artificial General Intelligence (AGI), hadn't fully solved real-world tasks. He observed that benchmarks often mirror training data, failing to capture the true complexity of real-world use cases.

This disconnect, they argued, necessitates models that can learn continuously and adapt at test time, learning directly from users and their specific data. "My conclusion is we need to have models that learn at test time. We need to have models that learn with users on their data, on their real-world tasks, on the real world distribution," Tworek explained. He highlighted the limitations of current adaptation methods like in-context learning and fine-tuning, pointing to issues like catastrophic forgetting and data efficiency constraints.

The Transformer's Legacy and the Path Forward

Tworek acknowledged the immense value and scalability that Transformers brought to AI, enabling the current era of large language models. He noted that while technically other architectures like LSTMs could have been scaled, Transformers proved more economically viable and performed better. "The majestic thing about Transformer... is that Transformers are economically valuable. The training them, the cost of training them is lower than the revenue that they generate," he said.

However, he stressed that this success shouldn't preclude exploration of new avenues. The founders believe that much architectural research has been conducted at too small a scale, hindering the discovery of potentially superior models. They emphasized the need for a baseline level of computational power to even reveal the capabilities of new architectures, a hurdle that often prevents innovative ideas from emerging.

The Case for Core Animation

Addressing the question of why they chose to launch a company rather than pursue this research within a large lab, the founders pointed to market dynamics. They believe that major labs, focused on scaling profitable existing technologies like Transformers to maintain market dominance, are less incentivized to explore alternative, potentially disruptive architectures. Meanwhile, smaller labs often try to replicate the success of larger ones, creating a lack of diversity in research paths.

Core Animation aims to fill this gap by focusing on fundamental architectural research and developing models capable of learning more effectively and over longer horizons, potentially leading to more adaptable and truly intelligent systems. StartupHub.ai data indicates that while Codex, a related technology from OpenAI, scores 47/100, the Transformer architecture itself scores 35/100, suggesting that while Transformers have been impactful, there is indeed room for architectural evolution.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer