The method uses a 3D skeleton for sparse but accurate geometric conditioning, Reference Context Packing to keep reference context within a fixed budget, and Target Context Routing to communicate structure across independently generated view groups. The generated views are then used to train a 4DGS model.
4DAnyone is useful for novel-view human video, digital characters, virtual production, telepresence, sports analysis, and 4D scene reconstruction from footage captured in the wild. Its focus on mild camera motion and unknown intrinsics makes it practical for casual video rather than specialized capture setups.

