Key Features

A video demonstration provides the main task prompt.
The demonstrated in-context workflow operates without task-specific post-training.
The research evaluates tasks outside the pretraining distribution.
The announcement emphasizes tasks with horizons around ten minutes.
Demonstrations communicate procedural details that short verbal instructions may omit.
One pretrained model interprets different demonstrations at inference time.
Evaluation separates seen versus unseen tasks and short versus long horizons.
The announcement defers deeper training explanations to later posts.

The model uses demonstration video as task context and applies a shared set of pretrained weights to execution. The research examines both whether a task was seen during pretraining and whether it requires a long sequence of actions, separating familiar imitation from harder forms of task and horizon generalization.


S1 is relevant to teams researching how robots can acquire new behaviors without collecting extensive new teleoperation data. The announcement highlights unseen tasks and roughly ten-minute horizons, while deeper training details are reserved for later publications. Public materials are available, but a downloadable model or priced self-service product is not established.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!