World Models: The Key to AGI?

Ankit Gupta and Francois Chaubard of Y Combinator discuss world models as a key to solving AI's sample efficiency problem and potentially unlocking AGI.

World Models: The Key to AGI?
YC
Visual TL;DR
AGI QuestDriver
From the article 3 mentionsThe quest for Artificial General Intelligence (AGI) hinges on solving a fundamental challenge: sample efficiency.
Challenges & ScalingDriver
addressing current hurdles and scaling issues for real-world applications
From the article 3 mentionsThe challenge of non-differentiability in complex environments, where the actions of other agents (like other cars on the road) are unknown, further complicates model-based approaches.
Sample Efficiency GapDriver
AI needs thousands of data points, humans learn with a handful of tries
From the articleThe quest for Artificial General Intelligence (AGI) hinges on solving a fundamental challenge: sample efficiency.
World Action ModelsCore
path forward involves integrating action planning with predictive world models
From the article 9+ mentionsThe conversation highlights the emergence of 'world action models' (WAMs), which jointly model state and action distributions.
World ModelsCore
core concept proposed by Y Combinator partners to bridge the efficiency gap
From the article 9+ mentionsThey delve into the concept of 'world models' as a potential breakthrough, exploring the motivations, mathematics, and real-world applications that could unlock AGI.
Human IntuitionContext
From the article 4 mentionsHow can AI models learn new tasks and skills rapidly from limited data, mirroring human intuition?
Optimal Control MathContext
mathematical foundations underpin how world models predict and plan actions
Unlock AGIOutcome
world models could be the breakthrough needed to achieve Artificial General Intelligence
From the article 3 mentionsThey delve into the concept of 'world models' as a potential breakthrough, exploring the motivations, mathematics, and real-world applications that could unlock AGI.
Contents(7)

The quest for Artificial General Intelligence (AGI) hinges on solving a fundamental challenge: sample efficiency. How can AI models learn new tasks and skills rapidly from limited data, mirroring human intuition? This critical problem is the focus of a recent discussion featuring Ankit Gupta, General Partner at Y Combinator, and Francois Chaubard, Visiting Partner at Y Combinator. They delve into the concept of 'world models' as a potential breakthrough, exploring the motivations, mathematics, and real-world applications that could unlock AGI.

Companies working on this

StartupHub profiles of the companies this article names, with funding and a one-liner from our database.

World
$193M
A humanity verification project and financial network providing anonymous proof of personhood for AI.
Tether
$500K
The stablecoin pegged 1:1 to the US dollar, facilitating blockchain transactions and payments.
Polymarket
$15.0B
A blockchain-based prediction market platform for trading shares on real-world event outcomes.
Circle
$2.8B
Circle provides a full-stack platform for the internet financial system, offering stablecoins like USDC and EURC, tokenized funds, and services for minting, FX, and payments.

The Sample Efficiency Gap

Gupta frames the core challenge as optimizing for 'intelligence per watt' and 'intelligence per sample.' He highlights that while humans can master new skills with just a handful of tries, current state-of-the-art AI models often need tens of thousands of data points. This inefficiency is a major hurdle, particularly in complex domains like robotics and self-driving cars, where real-world data collection is expensive and time-consuming.

The full discussion can be found on YC's YouTube channel.

AI Can't Learn The Way Humans Do - This Could Fix That - YC
AI Can't Learn The Way Humans Do - This Could Fix That, from YC

The discussion draws a parallel to human learning, where innate genetic encoding and learned experiences contribute to an implicit 'world model' within the brain. This internal model allows for prediction and planning, a capability that current AI models often lack in an explicit form.

World Models: The Core Concept

The central thesis revolves around 'world models', AI systems that aim to build an internal representation of how the world works. These models learn the transition function, essentially predicting the next state (St+1) given the current state (St) and an action (ut). This predictive capability is key to improving learning efficiency.

The presenters use the example of Newtonian physics, which serves as a perfect, albeit simplified, world model. This allows for precise predictions of object trajectories without needing to collect millions of real-world data points. This concept is further illustrated with NASA's asteroid interception plans and SpaceX's rocket landing systems, both relying on sophisticated internal models of physics to guide actions.

Human Intuition and World Models

The conversation highlights how humans naturally build and utilize world models. Even mental rehearsal, like imagining shooting a basketball, can lead to significant improvement, demonstrating the power of internal simulation. Neuroscientist Shaw Duckman's theory that the neocortex's expansion was driven by the need for better world modeling underscores this point.

The discussion contrasts the implicit world understanding in natural language models with the explicit need for such models in physical domains like robotics. While LLMs can perform surprisingly intelligent tasks through pattern matching in vast text data, this understanding can break down in scenarios requiring direct interaction with the physical world.

The Mathematics of Optimal Control

The presenters break down the problem of controlling a drone as a practical example. The state vector includes position and velocity, and the goal is to reach a target state. The transition function, governed by physics (F=MA), allows for precise prediction. This leads to the concept of 'model predictive control,' where a model is used to optimize actions over time to minimize a loss function, such as deviation from a target or energy expenditure.

However, the introduction of an adversary, like another drone trying to intercept, transforms the problem into a stochastic and non-differentiable one. This is where traditional optimization breaks down, necessitating approaches like reinforcement learning (RL). The video touches upon various RL techniques like value iteration, policy iteration, DQN, and actor-critic methods, all aimed at modeling these complex, non-differentiable processes.

Challenges and Scaling Issues

The discussion then pivots to the limitations of current approaches, particularly those used in games like Chess and Go. While AlphaGo achieved remarkable success, its scalability is hampered by the combinatorial explosion of states and the computational cost of planning. In real-world applications like self-driving cars and robotics, the state space is effectively infinite, and the need for real-time decision-making is paramount.

The challenge of non-differentiability in complex environments, where the actions of other agents (like other cars on the road) are unknown, further complicates model-based approaches. This forces a reliance on RL, which, while powerful, is often described as 'brutal' and sprawling due to the variety of algorithms and the difficulty in modeling stochastic processes.

The video emphasizes that the success of AlphaGo was contingent on a small action space and a deterministic environment. Real-world scenarios, like the stock market or venture capital, are constantly changing, making a fixed world model insufficient. The need for real-time, adaptive models is clear.

The Path Forward: World Action Models

The conversation highlights the emergence of 'world action models' (WAMs), which jointly model state and action distributions. This approach aims to overcome the computational expense of sampling world models and then passing those actions back into the system. By having a single model output both the action and the next state, WAMs promise greater efficiency.

The discussion concludes by outlining the progression of environments from Chess to Go, then to self-driving cars and robotics, each presenting increasing complexity in state and action spaces. The core takeaway is that world models, by enabling predictive simulation and learning from fewer samples, might be the crucial missing piece in the pursuit of AGI.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer