Key Features

GE-Act combines a control-oriented autoencoder, single-step visual planner, and inverse dynamics model.
Separate pretraining lets visual prediction learn from video without recorded actions.
Knowledge-aligned selective optimization filters predicted futures for compatibility with action supervision.
The experiments scale manipulation data from 300 to 30,000 hours.
Evaluation covers 100 tasks across 20 manipulation skill groups.
Published zero-shot evaluations use pretrained checkpoints without per-task adaptation.
The full pipeline reports 104 milliseconds per action chunk on one RTX 5090.
Experiments test object, color, shape, position, and other language-grounding distinctions.

A control-oriented autoencoder compresses visual input, a single-step visual planner predicts future states, and an inverse dynamics model produces actions. The planner and action components are pretrained separately, then connected through knowledge-aligned selective optimization that filters futures incompatible with the recorded behavior.


The research evaluates direct deployment on 100 tasks across 20 skill groups without per-task fine-tuning. It is relevant to teams studying grounded instruction following and data-efficient robot learning. Public results and demonstrations are available, while the project currently marks its code release as coming soon.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!