Masked Visual Actions

NEW

Key Features

Uses masked visual actions for unified robot world modeling.
Supports forward modeling from action representations to video.
Supports inverse modeling for recovering action-related structure.
Uses rendered robot control videos as visual conditioning signals.
Works with reference images and text prompts in generation workflows.
Evaluates unseen grippers and unseen embodiments.
Links to public code for finetuning and inference.
Includes direct MP4 videos showing robot control comparisons.

The method uses rendered control videos, reference images, and prompts to generate corresponding real-world robot video behavior. Its project examples include DROID training-domain cases, unseen grippers, unseen embodiments, and comparisons with baseline control representations.


Masked Visual Actions is useful for robotics researchers, world-model developers, and teams building action-conditioned video prediction systems. It provides a practical bridge between robot control signals and generative video models by making action information visible and transferable.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!