Key Features

Instruction-based video editing.
Image DiT reused for video latents.
Wan 2.1 video-VAE integration.
Qwen-Image-Edit transformer backbone.
Qwen2.5-VL instruction prompting.
Wan 2.2 denoising enhancement.
45-frame long-video segments.
LoRA or full-parameter fine-tuning.

The system sends video through Wan 2.1's VAE, bridges its latent space into the Qwen-Image-Edit DiT with two small projections, and applies grid rotary positional encoding across latent frames. A Qwen2.5-VL prompt path supplies instructions, and Wan 2.2 denoising enhancement improves the edited result.


Fine-tuning can use LoRA or full parameters on Ditto-1M source, edited, and instruction triplets. The method supports long-video editing in 45-frame segments and portrait videos with preserved aspect ratio, making it useful for controllable post-production, localization, and creative video transformations.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!