A causal video-action policy learns from both language and human-video prompts. A data pipeline converts sampled robot trajectories into semantically matched human demonstrations, producing 74,200 pairs across 8,600 tasks. In-context future chunk prediction encourages the policy to use the visual demonstration rather than shortcut through text or recent robot history.
Zero-WAM is useful for research on task generalization without collecting a new set of robot demonstrations for every request. The project demonstrates specifying tasks through internet or on-site human video and executing without parameter updates. Public links provide code, data, and pretrained and post-trained resources.

