Key Features

Efficient video generation model for up to 720p outputs.
Uses Hybrid Linear-Softmax Attention to reduce quadratic attention cost.
Adds Attention Residuals to reuse softmax anchor features across depth.
Provides 5B and 14B model scale descriptions.
Reports strong VBench performance with low H100 latency.
Uses Sol-Engine optimization for additional inference acceleration.
Links to NVIDIA's Sana GitHub repository and an arXiv paper.
Designed for scalable long and high-resolution video generation.

The model uses Hybrid Linear-Softmax Attention with periodic softmax anchors and Block Attention Residuals to recover expressiveness while keeping long-sequence computation efficient. It is instantiated at 5B and 14B scales and optimized with Sol-Engine for faster inference on H100 hardware.


SANA-Video 2.0 is useful for video generation researchers, infrastructure teams, and developers seeking efficient high-resolution generation. Its architecture shows how careful attention design and system optimization can reduce video diffusion latency without abandoning quality.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!