Key Features

Controls hard cuts at user-specified frames.
Supports soft transitions such as cross-fades.
Triggers camera movement at precise times in the clip.
Uses learnable planning tokens injected into a video diffusion transformer.
Applies Fractional Temporal RoPE for frame-level timing in latent space.
Removes planning tokens before decoding to avoid output overhead.
Built on Wan video diffusion model workflows.
Provides public code, model, dataset, and paper links.

The method adds learnable planning tokens to a pre-trained video diffusion transformer and assigns them fractional temporal RoPE positions. These tokens guide denoising toward requested transition events, then are removed before decoding so output shape and inference overhead stay aligned with the base model.


ShotPlan is useful for AI filmmakers, video generation researchers, and tool builders who need multi-shot structure from a single prompt. It helps create videos with frame-accurate cuts, cross-fades, localized camera movement, and more coherent narrative pacing.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!