Key Features

Keyboard-style character actions and camera actions control generated video.
Actions are converted into short English instructions for individual video latent frames.
A directed attention mask binds instructions to corresponding latent frames.
The approach trains rank-32 LoRA adapters on self-attention projections.
It reuses the pretrained text pathway without action-specific trainable modules.
Temporally specified controls can switch camera or character behavior within a sequence.
The experiments test unseen combinations of known action primitives.
The project provides GitHub and Hugging Face release links.

The system translates each action into a short English instruction associated with a video latent frame. A directed attention mask aligns those instructions with their intended times, while a rank-32 LoRA adapts self-attention projections. This reuses the pretrained language pathway without adding a separate trainable action module.


H3-World is useful for studying timed control and compositional generalization in video models. Its examples show different actions from the same initial frame and combinations absent from training. The project links code and weights for experimentation with camera motion, character behavior, and longer generation horizons.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!