In an era where large language models dominate headlines, a quietly audacious bet on "world models" is positioning itself as the next frontier in artificial intelligence. This vision, championed by Pim de Witte, CEO of General Intuition (GI), and backed by Khosla Ventures' largest seed investment since OpenAI, posits that spatial-temporal foundation models, trained on a unique trove of human gameplay data, will redefine how AI interacts with the physical and simulated worlds.
Pim de Witte recently spoke with Swyx, Editor of Latent Space, at General Intuition's offices, delving into the foundational technology, the strategic advantage of their data, and the expansive future applications of world models. The discussion illuminated why this approach, distinct yet complementary to LLMs, represents a profound shift in AI's capabilities, moving beyond mere content generation to active, intuitive understanding.
At its core, General Intuition is building agents that learn to perceive and act within environments, mimicking human intuition. Unlike traditional video models that merely predict the next likely frame, world models are tasked with a far more complex challenge. "What world models do is they actually have to understand the full range of possibilities and outcomes... and based on the action that you take, generate the next state," de Witte explained. This action-conditioned generation is crucial, enabling AI to not just observe but also to interact and anticipate consequences within dynamic environments.
The bedrock of GI's innovation is the colossal dataset amassed from Medal, de Witte's previous venture. Medal, a game clipping platform with 12 million users, has accumulated an astounding "3.8 billion clips of the best moments and actions in games, resulting in one of the most unique and diverse datasets of peak human behavior." This treasure trove of "episodic memory for simulation" provides an unparalleled resource for training AI. Crucially, this data is privacy-preserving, mapping actions to visual inputs and game outcomes without revealing personal user data, a prescient design choice that became a goldmine for world model development.
GI's agents are purely vision-based, operating on a "frames in, actions out" paradigm. De Witte demonstrated models that, without any reinforcement learning (RL) or fine-tuning, and seeing "no game state," predict actions from raw pixels. These agents exhibit "incredibly human-like" behaviors, even making the same mistakes or performing actions like gamers checking a scoreboard, a testament to the fidelity of their imitation learning.
