Unified Embodied AI: Pelican-Unified 1.0

Pelican-Unified 1.0, the first unified embodied foundation model, achieves SOTA performance by integrating VLM, reasoning, and generation, proving unification enhances rather than compromises specialist strengths.

Diagram illustrating the unified architecture of Pelican-Unified 1.0
Pelican-Unified 1.0's unified approach to embodied AI.
Visual TL;DR
Fragmented AI ModelsDriver
From the article 4 mentionsThe pursuit of truly intelligent embodied agents has long been hampered by the need to train disparate, specialized models for perception, reasoning, and action.
Inefficiency & LimitsDriver
From the articleThis fragmentation leads to inefficiencies and limits the holistic capabilities of AI systems.
Pelican-Unified 1.0Core
From the article 3 mentionsThe introduction of Pelican-Unified 1.0 marks a significant departure, presenting the first embodied foundation model built on the principle of unification.
Unified VLMCore
single visual-language model maps diverse inputs to shared semantic space
From the article 5 mentionsPelican-Unified 1.0 leverages a single Visual-Language Model (VLM) to serve as a unified understanding and reasoning module.
Chain-of-Thought ReasoningCore
From the article 4 mentionsCrucially, it also performs autoregressive chain-of-thought reasoning, generating task- and action-oriented sequences in a single pass.
Simultaneous OptimizationEffect
From the articleThis unified approach allows for the backpropagation of language, video, and action losses into the shared representation, enabling simultaneous optimization of understanding, reasoning, imagination, and action, rather than relying on isolated expert systems.
SOTA PerformanceOutcome
achieves state-of-the-art performance by integrating perception, reasoning, generation
From the article 2 mentionsContrary to the intuition that unification might lead to diluted capabilities, Pelican-Unified 1.0 demonstrates that this paradigm can preserve and even enhance specialist performance.
Unification EnhancesOutcome
proving unification enhances rather than compromises specialist strengths
From the article 2 mentionsContrary to the intuition that unification might lead to diluted capabilities, Pelican-Unified 1.0 demonstrates that this paradigm can preserve and even enhance specialist performance.

The pursuit of truly intelligent embodied agents has long been hampered by the need to train disparate, specialized models for perception, reasoning, and action. This fragmentation leads to inefficiencies and limits the holistic capabilities of AI systems. The introduction of Pelican-Unified 1.0 marks a significant departure, presenting the first embodied foundation model built on the principle of unification.

Unifying Perception, Reasoning, and Imagination

Pelican-Unified 1.0 leverages a single Visual-Language Model (VLM) to serve as a unified understanding and reasoning module. This VLM maps diverse inputs, scenes, instructions, visual contexts, and action histories, into a shared semantic space. Crucially, it also performs autoregressive chain-of-thought reasoning, generating task- and action-oriented sequences in a single pass. This unified approach allows for the backpropagation of language, video, and action losses into the shared representation, enabling simultaneous optimization of understanding, reasoning, imagination, and action, rather than relying on isolated expert systems.

Specialist Strength Without Compromise

Contrary to the intuition that unification might lead to diluted capabilities, Pelican-Unified 1.0 demonstrates that this paradigm can preserve and even enhance specialist performance. A single checkpoint of the model achieved impressive results across multiple domains: 64.7 on eight VLM benchmarks (outperforming comparable-scale models), a first-place ranking of 66.03 on WorldArena, and 93.5 on RoboTwin (second-best among action methods). These findings underscore the efficacy of the unified approach in consolidating complex AI capabilities without sacrificing individual performance.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.