Key Features

Human demonstration videos serve as in-context task prompts.
One policy supports both language instructions and human-video prompts.
Video examples can identify the intended instance among similar objects.
Demonstrations convey temporal order for multi-stage manipulation tasks.
The demonstrated in-context approach executes tasks without task-specific parameter updates.
The reported dataset contains 74.2K human-robot pairs across 8.6K tasks.
IFP trains the policy to use human-video context when predicting task evolution.
The project provides code, data, pretraining, and post-training release links.

A causal video-action policy learns from both language and human-video prompts. A data pipeline converts sampled robot trajectories into semantically matched human demonstrations, producing 74,200 pairs across 8,600 tasks. In-context future chunk prediction encourages the policy to use the visual demonstration rather than shortcut through text or recent robot history.


Zero-WAM is useful for research on task generalization without collecting a new set of robot demonstrations for every request. The project demonstrates specifying tasks through internet or on-site human video and executing without parameter updates. Public links provide code, data, and pretrained and post-trained resources.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!