CLAP cross-embodiment action-conditioned video generation
CLAP trains action-conditioned video world models across human and robot video and matches single-embodiment models on DROID with zero-shot generalization.
5 min read

Visual TL;DR
single-robot video models cannot use heterogeneous data containing rich physics signals
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
From the article 5 mentionsCLAP proposes CLAP cross-embodiment action-conditioned video generation to train on internet-scale human and robot video instead.
matches single-embodiment baseline performance on DROID with zero-shot generalization across platforms
inferred action representations for unlabeled human video, completing the unified three-channel design
From the article 2 mentionsCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
single-robot video models cannot use heterogeneous data containing rich physics signals
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
spatiotemporal dynamics obey the same laws regardless of whether the actor is human or robot
From the articleAccording to arXiv, the work from Kechen Liu and Ola Shorinwa treats universal physical laws as the common substrate across embodiments.
From the article 5 mentionsCLAP proposes CLAP cross-embodiment action-conditioned video generation to train on internet-scale human and robot video instead.
precise robot action control, but limited to embodiment-specific kinematic structures and measurements
From the articleCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
natural language goals bridging human and robot videos, but lacks fine-grained motor specification
From the articleCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
inferred action representations for unlabeled human video, completing the unified three-channel design
From the article 2 mentionsCLAP reconciles this with three channels: end-effector poses, language instructions, and latent actions.
training schedule turns raw internet human video into implicit physics priors without manual annotation
From the articleIt first learns foundational physical priors across unlabeled video data using latent actions.
matches single-embodiment baseline performance on DROID with zero-shot generalization across platforms
Contents(4)
© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.