The system sends video through Wan 2.1's VAE, bridges its latent space into the Qwen-Image-Edit DiT with two small projections, and applies grid rotary positional encoding across latent frames. A Qwen2.5-VL prompt path supplies instructions, and Wan 2.2 denoising enhancement improves the edited result.
Fine-tuning can use LoRA or full parameters on Ditto-1M source, edited, and instruction triplets. The method supports long-video editing in 45-frame segments and portrait videos with preserved aspect ratio, making it useful for controllable post-production, localization, and creative video transformations.

