Key Features

Video question answering is part of the shared embodied model.
Isaac supports spatial grounding and tracking across observations.
The model can report task progress in addition to generating actions.
Robot state and previous actions complement images, video, and instructions.
Isaac 0.5 is described as a 36B-parameter sparse model.
Training spans more than 35 robot systems.
The release supports policy fine-tuning and use inside planners or controllers.
Checkpoints and training and inference code are released through LeRobot integration.

The 36B-parameter sparse model consumes visual inputs, language, robot state, and previous actions. Training jointly combines video understanding and control across more than 35 robot systems. Its reported data mix includes 100,000 hours of robot experience, one million hours of video, and three trillion multimodal tokens.


The release is useful for teams studying how broad video knowledge can reduce dependence on expensive robot demonstrations. Perceptron provides checkpoints and training and inference code through a LeRobot workflow. The published scaling results describe experimental action-loss tradeoffs rather than a guarantee of task success on every robot.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!