Key Features

It selects discrete semantic action units interpreted as local robot motion.
Action symbols are parameter-free; the embodiment interpreter supplies metric magnitudes.
The action vocabulary is shared across real robot embodiments and simulation.
The project demonstrates direct robot control with closed-source frontier VLMs.
Open VLMs can be fine-tuned through the same semantic interface.
GUMI is a graphical manipulation interface for collecting robot demonstrations.
The demonstrated collection workflow uses keyboard actions through a GUI.
The project links public code, trained models, and demonstration data.

The VLM chooses actions such as moving, rotating, grasping, releasing, or finishing. An embodiment-specific interpreter deterministically translates each symbol into local motion. The same interface supports zero-shot use of frontier VLMs and lightweight adaptation of smaller open models for faster deployment.


The project is useful for robot-control research and demonstration collection without specialized teleoperation hardware. Its GUMI interface allows humans or GUI agents to drive the robot by keyboard while recording training pairs. The release links code, models, and datasets for reproducing the approach.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!