The system translates each action into a short English instruction associated with a video latent frame. A directed attention mask aligns those instructions with their intended times, while a rank-32 LoRA adapts self-attention projections. This reuses the pretrained language pathway without adding a separate trainable action module.
H3-World is useful for studying timed control and compositional generalization in video models. Its examples show different actions from the same initial frame and combinations absent from training. The project links code and weights for experimentation with camera motion, character behavior, and longer generation horizons.

