The system combines embodied reasoning, predictions of dynamic visual regions, and action learning. Its embodied reasoner builds on Qwen3-VL-4B and more than five million reasoning samples. World-language-action training then draws on approximately 2,500 hours of real robot interaction data.
The project is useful for researchers developing unified humanoid policies that connect spatial understanding with execution. It demonstrates 64 tasks and supplies linked code, models, and datasets. Actual deployment requires compatible robot interfaces and the task-specific setup described by the release.

