Fei-Fei Li on Spatial AI: Marble, $1.23B, and the Language Gap

World Labs has raised $1.23 billion and shipped Marble since February 2026. Here is what Fei-Fei Li told Bloomberg Tech about why language models miss half the AI problem.

7 min read
Fei-Fei Li, World Labs CEO, spatial intelligence and Marble, 2026
Fei-Fei Li speaking at the AI for Good Global Summit 2017.· Photo by ITU Pictures, via Wikimedia Commons (CC BY 2.0)

World Labs, the spatial intelligence startup co-founded by Fei-Fei Li, shipped its first commercial product to general availability in February 2026 alongside a $1 billion growth round, bringing total capital raised to $1.23 billion. Four months on, speaking at Bloomberg Technology Summit 2026 in San Francisco on June 4, Li offered her clearest public account yet of why she left Stanford to build it.

Marble's First Five Months: Gaming, VFX, and a $200M Autodesk Bet

Marble, World Labs' flagship model, generates explorable 3D environments from multimodal inputs: text prompts, single images, video clips, or spatial sketches. The output is dual-format. Each generated world renders as Gaussian splats for visual exploration and as collision meshes that physics engines can operate on directly, allowing game studios, VFX pipelines, and robotics simulators to import the output without a conversion step.

The company has disclosed paying customers in four verticals: gaming, visual effects, robotics training, and architectural design, per TechCrunch reporting from November 12, 2025, at Marble's limited beta launch. The most concrete public signal of commercial momentum is Autodesk's $200 million lead investment in the February 2026 growth round. Autodesk, the dominant design-software vendor for entertainment and architecture, confirmed a product integration agreement under which the two companies will embed Marble's capabilities into Autodesk's toolchain, beginning with entertainment workflows including virtual film production, according to TechCrunch's February 18, 2026 report.

The broader round, which totalled $1 billion, also included NVIDIA, AMD, Andreessen Horowitz, Emerson Collective, Fidelity Management and Research, and Sea. The round valued World Labs at approximately $5 billion, per reporting from mlq.ai citing fundraising discussions. Combined with the company's $230 million Series A in September 2024, World Labs has raised $1.23 billion across two rounds in under 18 months of operations.

Bar chart showing World Labs fundraising: $230M Series A in September 2024 and $1B Growth Round in February 2026
World Labs capital raised by round. Sources: TechCrunch (Sep 2024), TechCrunch (Feb 2026), mlq.ai.

'Wordsmiths in the Dark': Li's Argument for Spatial Intelligence

At Bloomberg Technology Summit 2026, speaking with Bloomberg's Emily Chang on June 4, Li described what she sees as the structural limit of large language models. "Leading AI technology such as large language models have begun to transform how we access and work with abstract knowledge, yet they remain wordsmiths in the dark; eloquent but inexperienced, knowledgeable but ungrounded," Li said, in remarks carried by Bloomberg.

The critique is not a call to abandon language models. Li has consistently framed spatial intelligence as the complementary layer sitting beneath them. "I believe spatial intelligence is as critical as, and complementary to, language intelligence," she stated in February 2026 at the time of the funding announcement. Her Substack essay "From Words to Worlds" extends the argument: "Language enables machines to talk about the world, and the World Model will enable machines to finally understand, imagine, reason, and interact with the world."

The practical implication is architectural. Language models produce text from text. World models, as Li defines them, operate in three-dimensional, physics-respecting space, accepting perceptual inputs and producing environments that downstream systems (robots, simulation platforms, game engines) can inhabit. The distinction has particular weight for robotics: a robot learning to navigate a factory floor requires a spatial model of that environment, not a textual description of it. Li's investor list reflects this; NVIDIA and AMD, the primary suppliers of compute for simulation workloads, both participated in the February round.

Li's framing also stands in contrast to Meta's approach to world models. Yann LeCun, Meta's chief AI scientist, has staked his lab's research programme on JEPA, a non-generative architecture that learns by predicting representations in abstract latent space rather than rendering visible outputs. Li's approach is explicitly generative and three-dimensional. As this column covered last week, both researchers claim their architecture is the path toward physical-world AI; neither has yet demonstrated a planner-layer system that acts reliably in novel environments at scale.

Doughnut chart showing World Labs 1 billion dollar growth round composition: Autodesk 200M lead and remaining 800M from NVIDIA AMD a16z Fidelity and others
World Labs $1B Growth Round (Feb 2026): Autodesk led with $200M; remaining participants include NVIDIA, AMD, Andreessen Horowitz, Fidelity, Emerson Collective, and Sea. Source: TechCrunch, February 18, 2026.

What Comes After the Renderer: The Planner Layer

In June 2026, Li published a taxonomy on her Substack assigning every system calling itself a "world model" to one of three functions: renderer, simulator, or planner. Marble sits at the renderer-simulator intersection, producing both visual output (Gaussian splats) and physically consistent geometry (collision meshes). The planner layer, which would allow an AI agent to select and execute actions within a generated world, remains an open research problem across the industry, including at World Labs.

The taxonomy also functions as a product roadmap. By naming the planner as a third, adjacent function, Li signals that Marble's current commercial footprint is phase one of a longer architecture. World Labs' co-founders reinforce where it is headed: Justin Johnson, a computer vision researcher formerly at the University of Michigan; Christoph Lassner, who built 3D human body reconstruction systems at Synthesia; and Ben Mildenhall, a co-inventor of NeRF (Neural Radiance Fields) at Google Research. The team's backgrounds span neural rendering, physical simulation, and computer vision, covering the renderer and simulator layers in depth. The planner layer requires integration with reinforcement learning and robotics systems, an area World Labs has not yet publicly detailed.

Li is scheduled to join Geoffrey Hinton and Andrew Ng on the main keynote stage at Ai4 2026 in Las Vegas on August 4, 2026, where the announced agenda covers frontier AI research, human-centred innovation, and governance challenges, per the Ai4 2026 conference announcement from December 2025. Li also co-founded AI4ALL, a nonprofit working to expand access to AI education, which continues in parallel with World Labs. The combination reinforces a recurring theme in her public remarks: that the societal stakes of physical AI require broad public literacy alongside technical execution.

Bar chart showing World Labs cumulative capital raised: 0 at founding in October 2023, 230 million after Series A in September 2024, 1.23 billion after Growth Round in February 2026
World Labs cumulative capital raised, October 2023 to February 2026. Sources: TechCrunch (Sep 2024 Series A), TechCrunch (Feb 2026 Growth Round).

What It Means

World Labs has moved from stealth to $1.23 billion in capital and paying customers across four industry verticals in 18 months. Li's public framing positions spatial intelligence as the physical layer that grounds AI capability in the real world, complementing language models rather than competing with them. The Autodesk partnership, the NVIDIA and AMD participation in the round, and Marble's dual-format output (visual and physical) all signal a company building toward robotics and industrial simulation rather than a pure creative or entertainment play. The June Bloomberg Technology Summit remarks give Li's commercial roadmap its clearest intellectual frame yet: the goal is machines that do not just describe the world, but understand and act within it.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.