Marionette: Decoupling Appearance from State

The Marionette world model decouples visual synthesis from geometric state, enabling controllable, long-horizon interactive simulations with explicit state repair.

Diagram illustrating the three-stage architecture of the Marionette world model, showing state prediction, graphics bridging, and observation synthesis.
The Marionette world model separates world state prediction from appearance synthesis.
Visual TL;DR
Current World ModelsDriver
autoregressively generating visual observations, implicitly encoding structured properties like pose and geometry
From the article 5 mentionsCurrent interactive game world models struggle with long-term consistency.
Error AccumulationDriver
leads to compromised controllability and coherence in long-horizon interactive simulations
From the article 2 mentionsThis leads to error accumulation over time, compromising controllability and coherence.
Marionette World ModelCore
From the article 5 mentionsThe Marionette world model introduces a novel paradigm: explicitly modeling the evolving world state and delegating exact geometric computation to a fixed, zero-parameter renderer.
Predict 3D World StateContext
From the article 3 mentionsFirst, a two-stage autoregressive dynamics model predicts an explicit, interpretable 276-dimensional 3D world state.
Graphics BridgeCore
zero-parameter bridge translates predicted state into pose-control videos, computing geometry in closed form
From the article 2 mentionsSecond, a graphics bridge, requiring zero parameters, translates this predicted state into pose-control videos.
Decoupled AppearanceEffect
separates visual synthesis from geometric state, enabling explicit state repair
From the articleRouting appearance through the predicted state incurs no detectable loss in fidelity.
Controllable SimulationsOutcome
enables long-horizon interactive simulations with explicit state repair and uncompromised fidelity
From the article 2 mentionsThe core innovation lies in the explicit world state, which proves directly controllable.
Enhanced UtilityOutcome
improves repairability and consistency for interactive game world models
Contents(3)

Current interactive game world models struggle with long-term consistency. By autoregressively generating visual observations directly, they implicitly encode structured properties like pose and geometry. This leads to error accumulation over time, compromising controllability and coherence. The Marionette world model introduces a novel paradigm: explicitly modeling the evolving world state and delegating exact geometric computation to a fixed, zero-parameter renderer.

Separating Structure from Synthesis

Marionette operates in three stages to achieve this separation. First, a two-stage autoregressive dynamics model predicts an explicit, interpretable 276-dimensional 3D world state. This state encapsulates multi-entity articulated skeletons, metric root trajectories, and rotations. Second, a graphics bridge, requiring zero parameters, translates this predicted state into pose-control videos. Crucially, this bridge computes world-space geometry and occlusion in closed form, offering exactness without neural computation. Finally, a control-conditioned video-diffusion observation model synthesizes photorealistic RGB observations based on these structured controls.

Controllability and Repairability in the State Space

The core innovation lies in the explicit world state, which proves directly controllable. Experiments demonstrate that forcing a mismatched action stream alters root-aligned joint error by 31% across held-out segments, validating the state's responsiveness. Furthermore, long-horizon behaviors are determined and can be repaired within this state representation. Left unconstrained, generated characters drifted apart significantly, and a third of frames exhibited ground penetration. However, by imposing simple rules on the explicit state, such as a terrain collider and a separation cap, ground penetration was reduced by 66%, and character engagement was maintained. These improvements were achieved without any modifications to the observation model, highlighting the power of state-level control.

Uncompromised Fidelity, Enhanced Utility

Routing appearance through the predicted state incurs no detectable loss in fidelity. The system achieves a Frame-to-Frame Video Diffusion (FVD) score of 831, a marginal increase compared to the 799 FVD for recorded pose, suggesting photorealism is preserved. This architectural choice provides the significant benefit of a controllable and repairable world model, without sacrificing the visual quality expected from generative approaches. The Marionette world model offers a promising direction for building more robust and predictable interactive AI systems.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.