A control-oriented autoencoder compresses visual input, a single-step visual planner predicts future states, and an inverse dynamics model produces actions. The planner and action components are pretrained separately, then connected through knowledge-aligned selective optimization that filters futures incompatible with the recorded behavior.
The research evaluates direct deployment on 100 tasks across 20 skill groups without per-task fine-tuning. It is relevant to teams studying grounded instruction following and data-efficient robot learning. Public results and demonstrations are available, while the project currently marks its code release as coming soon.

