The self-contained Diffusers package includes a Qwen2.5-VL planner, Bernini planning weights, and Wan2.2 diffusion components. Compared with renderer-only Bernini releases, v2 adds stronger instruction following and multi-step semantic planning, with a connector warmup and co-training recipe that improves reference-guided editing and image-to-video performance.
The model can be downloaded from Hugging Face and run with the public Bernini repository, including single-GPU inference and a Gradio demo. It is useful for text-to-video, image-to-video, video-to-video, reference-guided editing, and research into planner-renderer video systems.

