The VLM chooses actions such as moving, rotating, grasping, releasing, or finishing. An embodiment-specific interpreter deterministically translates each symbol into local motion. The same interface supports zero-shot use of frontier VLMs and lightweight adaptation of smaller open models for faster deployment.
The project is useful for robot-control research and demonstration collection without specialized teleoperation hardware. Its GUMI interface allows humans or GUI agents to drive the robot by keyboard while recording training pairs. The release links code, models, and datasets for reproducing the approach.

