Key Features

Preview V1 supports text-to-video-and-audio generation.
The recommended distilled model uses four DiT calls.
The recommended release requires FastVideo VSA-H3 sparse attention kernels.
The release includes full model weights and a pre-extracted LoRA.
One checkpoint supports multiple resolutions, aspect ratios, and durations.
The project reports 15-second 768p video in under 13 seconds on eight B200 GPUs.
First/last-frame and full-reference variants are described as forthcoming.
The authors state dense attention is not a drop-in replacement for the recommended sparse checkpoint.

The recommended checkpoint combines four-step distillation with Video Sparse Attention and prompt-only training. It reuses H3 text encoding and audio and video autoencoders, while the VSA-H3 backend reduces attention cost. Both full weights and a pre-extracted LoRA are provided for the recommended release.


FastH3 is useful for developers experimenting with fast open-weight audiovisual generation on NVIDIA Blackwell systems. Its current preview supports text-to-video-and-audio only; frame and full-reference conditioning remain under development. Reported speedups depend on the sparse backend and specified hardware and should not be generalized to dense attention substitutes.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!