Key Features

VDN-H3 applies Video DeltaNet to MiniMax H3.
A sliding-window softmax branch retains exact attention around nearby frames.
Bidirectional linear states summarize video outside the local attention window.
First and last frames act as boundary anchors visible across the sequence.
Text prompts influence both the softmax and linear attention branches.
The checkpoint supports first-frame, last-frame, and combined boundary-frame conditioning.
A 14.4-second 768p clip completes in about nine seconds on eight B200 GPUs.
The project links public code and model weights.

The architecture combines sliding-window softmax attention with bidirectional linear memory. First and last frames act as global anchors, while the linear branches handle distant context outside the local window. Text conditions both paths, and the design avoids counting the same video context twice.


Video DeltaNet is useful for researchers and developers optimizing large video models. The project reports a finished 14.4-second, 768p clip in about nine seconds after warm-up on eight B200 GPUs. Those measurements are hardware-specific and do not imply equivalent performance on a consumer graphics card.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!