The recommended checkpoint combines four-step distillation with Video Sparse Attention and prompt-only training. It reuses H3 text encoding and audio and video autoencoders, while the VSA-H3 backend reduces attention cost. Both full weights and a pre-extracted LoRA are provided for the recommended release.
FastH3 is useful for developers experimenting with fast open-weight audiovisual generation on NVIDIA Blackwell systems. Its current preview supports text-to-video-and-audio only; frame and full-reference conditioning remain under development. Reported speedups depend on the sparse backend and specified hardware and should not be generalized to dense attention substitutes.

