Yann LeCun said AMI Labs, his new research venture, has nothing to ship yet-six months after he launched it. In an interview with La Tribune Événements, he explained the team has grown to between 50 and 60 researchers across Paris, New York, Montreal and Singapore. They’re in talks with potential partners while building what he calls “world models.”
LLMs are monohulls, he said, using the image of a sailing boat to explain why he stepped away from the LLM race. You can keep refining a single hull with better materials and ballast, he noted, but two or three hulls with foils that lift you out of the water get you there faster. LLMs, in his telling, are the monohull.
They tokenize text into discrete symbols and train to predict the next symbol, producing a kind of distillation of human knowledge that is grammatically fluent and useful for accelerating access to information. But they aren’t truly intelligent because they can’t predict the consequences of actions. The flying trimaran is JEPA.
JEPA-the Joint Embedding Predictive Architecture-was first proposed by LeCun in 2022. Over 2,000 papers have followed, covering everything from medicine and imaging to video interpretation, autonomous driving and robotics. The method doesn’t try to generate every pixel of the next video frame. Instead, it shows the system a short clip and asks it to predict what happens next in an abstract representation that discards unpredictable detail. His example: a self-driving car on a windy day. The model should learn to track cars and pedestrians, not the random flutter of leaves. A generative model forced to reconstruct pixels wastes capacity on noise. JEPA learns to keep only what it can predict.
That distinction is where LeCun draws the line with current product work. Companies have bolted vision encoders and tool use onto LLMs to make agentic systems that can output actions instead of text-for cinema or desktop automation. But those systems are brittle because they can’t simulate outcomes before acting. World models, he said, are meant to do exactly that planning. He described them as an automated way to build digital twins. A traditional twin for a turbojet requires years of physics equations for fluids, thermodynamics and vibration. A world model would learn a phenomenological twin from sensor streams. For something as complex as a human cell, where no one can write the equations, he said learning is the only path.