The system reformulates the diffusion transformer as a recurrent, chunk-wise autoregressive process with near-constant memory attention. It combines causal linear attention, learnable attention head pruning, sampling-step distillation, and memory-optimized VAE decoding so a large video model can generate short vertical videos within a practical mobile latency budget.
MobileWan is useful for researchers and product teams exploring local video generation, private mobile media creation, and efficient diffusion deployment. The project reports 5-second 480x832 videos at 16 FPS in about 20 seconds, an 83.79 VBench score, and user-study preference over earlier mobile video generation baselines.


