Key Features

A single monocular video provides the visual input.
The pipeline composes pretrained foundation models without training a new task model.
The scene is decomposed into independent labeled mesh instances.
The system estimates metric scale and six-degree-of-freedom pose trajectories.
Direct per-vertex deformation captures non-rigid mesh motion.
Motion is represented without category-specific skeleton rigs.
Contact-aware assembly produces coherent scenes with URDF export for physics engines.
Mixed third-party licenses limit the assembled release to noncommercial research.

The training-free pipeline combines pretrained models for instance discovery, segmentation, object reconstruction, and pose recovery. A render-match-optimize loop estimates scale and motion, while physics-grounded assembly enforces contact and support. Direct vertex deformation represents both rigid and non-rigid behavior without requiring skeleton rigging.


OVOW is useful for research on video-to-simulation conversion and structured world-model datasets. Outputs include watertight instances and URDF-compatible scene assembly. Although its own code is MIT-licensed, bundled dependencies impose additional restrictions, so the assembled release is intended for noncommercial research rather than unrestricted commercial deployment.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!