The model learns action-conditioned visual dynamics from egocentric, synthetic, and web video. The project reports real-time generation at 720p and 60 frames per second with sub-second control latency. Player and director inputs can jointly influence the generated world, including subjects, events, and environmental changes.
The work supports research in interactive entertainment, embodied simulation, and future-state prediction. The repository releases 14B and 1.3B variants with inference code under CC BY-NC-SA 4.0, superseding the landing page release notice. Long-term memory and faithful physical reasoning remain research challenges, and the internal deployment stack is not released.

