The generation of motion-controlled videos has been hampered by an inability to disentangle object dynamics from camera perspectives and a lack of causal reasoning, leading to animations that merely displace pixels rather than simulating realistic reactions. This limitation restricts user control and the plausibility of generated scenes.
Unifying Disentangled Motion and Causal Dynamics
The MoRight framework represents a significant advancement by addressing these core challenges. It introduces a unified approach that disentangles object motion control from arbitrary camera viewpoint adjustments. This is achieved by specifying object motion in a canonical static view and then transferring it to the target camera perspective using temporal cross-view attention. This architectural innovation allows for independent control over scene dynamics and camera framing, a critical step beyond existing methods that conflate these signals.