The 36B-parameter sparse model consumes visual inputs, language, robot state, and previous actions. Training jointly combines video understanding and control across more than 35 robot systems. Its reported data mix includes 100,000 hours of robot experience, one million hours of video, and three trillion multimodal tokens.
The release is useful for teams studying how broad video knowledge can reduce dependence on expensive robot demonstrations. Perceptron provides checkpoints and training and inference code through a LeRobot workflow. The published scaling results describe experimental action-loss tradeoffs rather than a guarantee of task success on every robot.

