Key Features

Monocular video to 4D human reconstruction.
Multiview-consistent target-view video generation.
4D Gaussian Splatting output.
3D skeleton geometric conditioning.
Reference Context Packing.
Target Context Routing.
Unknown-camera and in-the-wild robustness.
Public paper, code, and model resources.

The method uses a 3D skeleton for sparse but accurate geometric conditioning, Reference Context Packing to keep reference context within a fixed budget, and Target Context Routing to communicate structure across independently generated view groups. The generated views are then used to train a 4DGS model.


4DAnyone is useful for novel-view human video, digital characters, virtual production, telepresence, sports analysis, and 4D scene reconstruction from footage captured in the wild. Its focus on mild camera motion and unknown intrinsics makes it practical for casual video rather than specialized capture setups.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!