CLAP cross-embodiment action-conditioned video generation
CLAP trains action-conditioned video world models across human and robot video and matches single-embodiment models on DROID with zero-shot generalization.

Visual TL;DR
single-robot video models cannot use heterogeneous data containing rich physics signals
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
From the article 5 mentionsCLAP proposes CLAP cross-embodiment action-conditioned video generation to train on internet-scale human and robot video instead.
matches single-embodiment baseline performance on DROID with zero-shot generalization across platforms
inferred action representations for unlabeled human video, completing the unified three-channel design
From the article 2 mentionsCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
single-robot video models cannot use heterogeneous data containing rich physics signals
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
spatiotemporal dynamics obey the same laws regardless of whether the actor is human or robot
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
From the article 5 mentionsCLAP proposes CLAP cross-embodiment action-conditioned video generation to train on internet-scale human and robot video instead.
precise robot action control, but limited to embodiment-specific kinematic structures and measurements
From the articleCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
natural language goals bridging human and robot videos, but lacks fine-grained motor specification
From the articleCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
inferred action representations for unlabeled human video, completing the unified three-channel design
From the article 2 mentionsCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
training schedule turns raw internet human video into implicit physics priors without manual annotation
From the articleIt first learns foundational physical priors across unlabeled video data using latent actions.
matches single-embodiment baseline performance on DROID with zero-shot generalization across platforms
Contents(4)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.