Fei-Fei Li's $5B World Labs and the Spatial AI Race No One Agrees On

World Labs raised $1B at a $5B valuation in 2026. Here is how Fei-Fei Li's spatial AI approach diverges from LeCun's JEPA, DeepMind's Genie 3, and Nvidia's Cosmos.

6 min read
Fei-Fei Li, World Labs spatial AI vs peers, 2026
Fei-Fei Li speaking at AI for Good 2017.· Photo by ITU Pictures (CC BY 2.0), via Wikimedia Commons.

World Labs raised $1 billion in February 2026 at a $5 billion valuation, anchored by a $200 million strategic investment from Autodesk, according to Bloomberg. That single round made Fei-Fei Li's startup the best-funded pure-play generative world-model company in the West. It also sharpened a debate that has been building for two years: what a "world model" should actually do.

Marble and the generative bet

World Labs launched its first commercial product, Marble, in November 2025, then opened it to general availability alongside the $1 billion round in February 2026. Marble takes a single image, a short video, a text prompt, or a 360-degree panorama as input and returns an explorable 3D environment with six-degrees-of-freedom navigation, all running in the browser. The underlying architecture renders output as Gaussian splats for visual exploration and as collision meshes that physics engines can operate on directly, according to Fast Company.

The company has paying customers across four verticals: gaming, visual effects, robotics training, and architectural design. Autodesk's $200 million strategic investment connects directly to the last two: the CAD giant gains early access to a generative pipeline that can produce simulation environments and architectural walkthroughs from images or descriptions. Nvidia, AMD, Andreessen Horowitz, Fidelity Management, and Emerson Collective also participated in the round.

In a November 2025 essay published on her Substack newsletter, Li framed the ambition precisely: "Spatial Intelligence is the scaffolding upon which our cognition is built." She described spatial intelligence as something that "will transform how we create and interact with real and virtual worlds, revolutionizing storytelling, creativity, robotics, scientific discovery, and beyond." Her phrase for the shift, "words to worlds," has become the shorthand for what she argues AI must now do.

Where LeCun draws a different line

The phrase "world model" also anchors the thesis at AMI Labs, the Paris-based research startup that Yann LeCun launched after departing Meta. AMI Labs raised $1.03 billion in a seed round at a $3.5 billion pre-money valuation in March 2026, reported as the largest seed round in European history. But the two companies use the same two words to describe something architecturally opposite.

LeCun's core framework, the Joint Embedding Predictive Architecture (JEPA), is deliberately non-generative. It predicts in compressed representation space and explicitly avoids producing pixels or photorealistic output. The argument, laid out in his 2022 paper "A Path Towards Autonomous Machine Intelligence," is that pixel reconstruction is computationally expensive overhead for a system whose job is to plan and act, not to display. As covered in this column on August 4, AMI Labs has begun producing formal proofs that JEPA-trained representations converge in ways that text-only transformers do not.

Li's approach moves in the opposite direction. Marble's outputs are meant to be seen, navigated, and exported by humans. The resolution to the apparent conflict is a use-case split: world models built for humans to see and create with favor the generative school, because the output is the product. World models built for a robot to plan with favor JEPA, because the relevant output is a decision and photorealistic pixels are waste. The two companies are running in different lanes of the same broad race.

The corporate dimension: DeepMind and Nvidia

Both Li and LeCun also compete with division-scale efforts inside large companies. Google DeepMind released Genie 3 for real-time interactive 3D world generation, expanding a model line that began as a text-to-interactive-environment research project. Nvidia's Cosmos platform, a set of physics-aware world foundation models aimed at physical AI applications, surpassed two million downloads by January 2026, according to Nvidia. Cosmos generates physics-aware video of predicted future environment states, targeting robotics and autonomous-vehicle training pipelines.

World Labs's differentiator in that context is persistence and editability. Marble environments can be navigated, modified, and re-exported, making them closer to a 3D content-creation platform than a one-shot video generation tool. That distinction maps to the Autodesk partnership more cleanly than it maps to Cosmos or Genie 3, both of which optimize for simulation throughput over creative workflow.

StartupHub.ai data shows World Labs counts 150 employees versus 10 at AMI Labs, a 15-to-1 ratio that maps to the two companies' contrasting structures. World Labs is building product and distribution infrastructure; AMI Labs is running as a research-intensive lab where a small team produces papers and architecture proofs before scaling headcount.

What it means

Spatial AI is not one race but at least three parallel bets, defined by the intended audience of the model's output. World Labs is building for human creators in games, film, architecture, and robotics simulation. AMI Labs is building for autonomous agents that need to plan within a compressed world model. And Nvidia's Cosmos is building for the industrial training pipelines that need physics-aware video at scale. The $5 billion valuation on World Labs reflects investor confidence that the generative path has the clearest near-term commercial surface, with Autodesk's anchor investment and four active customer verticals as supporting evidence.

Sources

Editorial standards: every claim is sourced. Tips: [email protected]

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.