Yann LeCun says LLMs can't model the real world

Yann LeCun tells Unsupervised Learning LLMs are not a path to human-like intelligence and that his new lab AMI will scale world models and JEPA instead.

Yann LeCun appeared on Unsupervised Learning: With Jacob Effron to draw a hard line between useful language models and true intelligence. He used the chance to argue for his new company, AMI Labs.

Yann LeCun says LLMs can't model the real world
Yann LeCun says LLMs can't model the real world

His bet is simple.

LeCun told host Jacob Effron that large language models are great at handling language, code and math, but they're not a route to human-like or even animal-like intelligence. He introduced AMI, short for Advanced Machine Intelligence, as built around the Joint Embedding Predictive Architecture, or JEPA, which he pioneered at Meta and scaled into a world model for the physical world. The motto is AI for the real world, a contrast to the tidy realm of text. Reality, he said, is high-dimensional, continuous, noisy and messy, and harder to learn than language.

LeCun shared the 2018 Turing Award with Yoshua Bengio and Geoffrey Hinton and led Meta’s FAIR team for ten years. He said the decision to leave became clear at the end of last year. Llama 1, built inside FAIR, looked promising in early 2023 and pushed Meta to create a new GenAI group to turn it into Llama 2, 3 and 4. Llama 4 disappointed him, he said, and triggered a reorganization under Mark Zuckerberg. The bigger shift, he argued, was strategic: Meta had fallen behind and refocused almost entirely on catching up in LLMs. His work on JEPA and world models kept backing from Zuckerberg and CTO Andrew Bosworth, but the rest of the company went the other way. He left to develop the research into products Meta wasn’t pursuing, with AMI Labs headquartered in Paris and an office in New York, deliberately outside Silicon Valley where he sees herd behavior and everyone digging the same trench.

LeCun warned that ‘world model’ has turned into a buzzword and split the field into two camps. He dismissed vision-language-action models, or VLAs, which try to turn vision and language straight into robot actions using LLM-style autoregressive prediction. He said that approach is now widely seen as failing because it’s unreliable and demands too much task-specific data. In his view, a world model is narrower and more essential. It lets an agentic system forecast the consequences of its own actions, then plan a sequence by search and optimization instead of token-by-token prediction. LLMs, he noted, do neither; they can’t predict consequences and they don’t plan by search.

JEPA is LeCun’s answer to predicting without reconstructing pixels. He pointed out that humans don’t forecast the exact pixel pattern of a sliding or tipping water bottle; they think at an abstract level. JEPA works the same way in latent space: one encoder views a corrupted input, another encodes the target, and a predictor tries to match the two representations. He said this non-generative route came from a realization about five years ago that generative methods had stalled for images and video. Variational autoencoders and masked autoencoders such as BERT for text and MAE for images at FAIR turned out to be compute-heavy disappointments, learning little beyond identity and denoising. In contrast, joint embedding approaches like DINO, DINOv2, I-JEPA and V-JEPA learned richer representations precisely because they didn’t try to generate pixels.

The skeptic’s question is whether latent prediction actually delivers data efficiency in the wild. LeCun pointed to robotics demos that look impressive but rely on huge amounts of imitation data from teleoperation, handheld grippers or hand tracking, plus fine-tuning in simulation. He said that makes them brittle, demanding new data for every task. By contrast, a world model should generalize zero-shot to new tasks through planning, much like a 17-year-old learns to drive in about 20 hours while industry has burned millions of hours of driving data without reaching level-five autonomy. He noted the logic is clean but unproven at scale. AMI has not released public benchmarks, customer deployments, compute or funding details, and LeCun acknowledged a domestic robot is still several years away despite many companies claiming otherwise. Generative world models from Google with Genie and efforts to generate synthetic video for simulation from Nvidia and others offer a competing bet that imperfect physics is good enough to close the data gap, a path LeCun rejects as unnecessary if sample efficiency improves.

JEPA is being extended beyond Meta's image/video work to end-to-end autonomous driving as AD-E2E-JEPA, showing independent validation of the architecture in September 2026.

For now, AMI remains a thesis with early results LeCun says he’s confident in, though it hasn’t been opened for independent verification. Until the lab shows that planning in latent space beats imitation learning on real hardware, using less data and across tasks, the idea that LLMs are useful but the wrong track will stay a principled dissent rather than a product lead.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.