The model is a multimodal autoregressive diffusion Transformer pretrained from scratch. Inputs are organized into a shared spatial context that grounds observations at positions in 3D. Camera geometry then guides new views, letting the system preserve known scene content while imagining unobserved regions.
Atlas demonstrates camera-controlled video, reconstruction from one or many images, space-time simulation, and text-driven panoramas. It is intended to power future World Labs products, including Marble. The announcement offers early-access requests rather than establishing a broadly available production service or published subscription price.

