Key Features

Adapts a 5B-parameter Wan2.2 video diffusion model for mobile deployment.
Uses recurrent distillation to generate video chunk by chunk with near-constant memory.
Combines causal linear attention with transformer-to-RNN-style inference behavior.
Applies learnable attention head pruning with binary gates and sparsity regularization.
Uses sampling-step distillation to reduce the number of diffusion steps.
Optimizes VAE decoding for memory-constrained devices.
Generates 5-second 480x832 videos at 16 FPS with about 20 seconds latency.
Reports state-of-the-art VBench results among mobile video diffusion methods.

The system reformulates the diffusion transformer as a recurrent, chunk-wise autoregressive process with near-constant memory attention. It combines causal linear attention, learnable attention head pruning, sampling-step distillation, and memory-optimized VAE decoding so a large video model can generate short vertical videos within a practical mobile latency budget.


MobileWan is useful for researchers and product teams exploring local video generation, private mobile media creation, and efficient diffusion deployment. The project reports 5-second 480x832 videos at 16 FPS in about 20 seconds, an 83.79 VBench score, and user-study preference over earlier mobile video generation baselines.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner
Zero to AI Engineer Program

Zero to AI Engineer

Skip the degree. Learn real-world AI skills used by AI researchers and engineers. Get certified in 8 weeks or less. No experience required.

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!