Key Features

The model takes a driving video and one repainted frame from that video.
The video stage uses fixed conditioning rather than a user-supplied text prompt.
Motion, timing, and camera behavior are inherited from the driving clip.
No separate pose estimator is needed in the video-generation stage.
Character replacement does not require a separate segmentation-mask channel.
The distilled sampler uses three forward passes in the upstream setup.
It fine-tunes the 33.1B MiniMax-H3 reference-to-video Transformer.
The model card links community ComfyUI nodes for the workflow.

The model is a 33.1B full fine-tune of MiniMax-H3 reference-to-video architecture, distilled to three forward passes. Because the edited reference already matches the source pose and framing, the video stage does not need a separate pose estimator, segmentation model, face tracker, or text encoder.


Viggle-Animate suits animation and character-replacement workflows where a prepared reference image is available. The release includes weights, examples, and links to community ComfyUI integration. Reported speed measurements use a B200 GPU; the MiniMax-H3 community license governs usage and should be reviewed for deployment.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!