The method adds learnable planning tokens to a pre-trained video diffusion transformer and assigns them fractional temporal RoPE positions. These tokens guide denoising toward requested transition events, then are removed before decoding so output shape and inference overhead stay aligned with the base model.
ShotPlan is useful for AI filmmakers, video generation researchers, and tool builders who need multi-shot structure from a single prompt. It helps create videos with frame-accurate cuts, cross-fades, localized camera movement, and more coherent narrative pacing.

