CLAP cross-embodiment action-conditioned video generation

CLAP trains action-conditioned video world models across human and robot video and matches single-embodiment models on DROID with zero-shot generalization.

5 min read
Concept diagram for CLAP cross-embodiment action-conditioned video generation spanning human and robot video
CLAP unifies end-effector poses, language and latent actions to train video world models across embodiments.
Visual TL;DR
Embodiment lock-inDriver
single-robot video models cannot use heterogeneous data containing rich physics signals
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
CLAP frameworkCore
From the article 5 mentionsCLAP proposes CLAP cross-embodiment action-conditioned video generation to train on internet-scale human and robot video instead.
DROID zero-shot matchOutcome
matches single-embodiment baseline performance on DROID with zero-shot generalization across platforms
Latent actions channelContext
inferred action representations for unlabeled human video, completing the unified three-channel design
From the article 2 mentionsCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
Embodiment lock-inDriver
single-robot video models cannot use heterogeneous data containing rich physics signals
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
Universal physics substrateContext
spatiotemporal dynamics obey the same laws regardless of whether the actor is human or robot
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
CLAP frameworkCore
From the article 5 mentionsCLAP proposes CLAP cross-embodiment action-conditioned video generation to train on internet-scale human and robot video instead.
End-effector poses channelContext
precise robot action control, but limited to embodiment-specific kinematic structures and measurements
From the articleCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
Language instructions channelContext
natural language goals bridging human and robot videos, but lacks fine-grained motor specification
From the articleCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
Latent actions channelContext
inferred action representations for unlabeled human video, completing the unified three-channel design
From the article 2 mentionsCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
Unlabeled video curriculumCore
training schedule turns raw internet human video into implicit physics priors without manual annotation
From the articleIt first learns foundational physical priors across unlabeled video data using latent actions.
DROID zero-shot matchOutcome
matches single-embodiment baseline performance on DROID with zero-shot generalization across platforms
Contents(4)

Single-embodiment video world models cannot scale. CLAP proposes CLAP cross-embodiment action-conditioned video generation to train on internet-scale human and robot video instead.

According to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.

Why Embodiment Lock-In Breaks World Models

State-of-the-art action-conditioned video models are tied to one robot.

That prevents them from using heterogeneous video that already contains rich physics signals.

CLAP is built on the observation that spatiotemporal dynamics obey the same laws regardless of the actor, whether human or robot.

How CLAP Unifies Three Incompatible Action Spaces

Cross-embodiment learning fails because action representations do not align and human video has no action labels at all.

CLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.

Each has a clear limitation alone. Poses are precise but embodiment specific. Language is universal but underspecified. Latent actions are scalable but ungrounded.

Together they cover the tradeoff.

The Curriculum That Turns Unlabeled Video Into Physics Priors

CLAP introduces a curriculum-based recipe to resolve those limitations.

This two-stage design lets the model absorb internet-scale dynamics without requiring action labels everywhere.

So What: From Generalist Pretraining to Specialist Supremacy

CLAP approaches or surpasses state-of-the-art single-embodiment video models on challenging environments like DROID.

The advantage compounds with few-shot adaptation, establishing a new paradigm for training single-embodiment video world models.

The suite spans end-effector, language, and latent conditioning across DROID, Bridge, bimanual YAM robots and G1 humanoids, with all code and models open-sourced.

For founders and investors, this inverts the standard world-model bet. Instead of collecting more DROID or Bridge data to make a better DROID model, build a cross-embodiment generalist on cheap human video and few-shot it down.

It mirrors what happened in language and vision: scale on heterogeneous data, then specialize. Robotics world models are now following the same curve.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.