The model is a 33.1B full fine-tune of MiniMax-H3 reference-to-video architecture, distilled to three forward passes. Because the edited reference already matches the source pose and framing, the video stage does not need a separate pose estimator, segmentation model, face tracker, or text encoder.
Viggle-Animate suits animation and character-replacement workflows where a prepared reference image is available. The release includes weights, examples, and links to community ComfyUI integration. Reported speed measurements use a B200 GPU; the MiniMax-H3 community license governs usage and should be reviewed for deployment.

