The method operates on top of a frozen WAN text-to-video diffusion transformer. It samples source anchors with Grounded SAM-2, tracks their trajectories, matches them to the target subject with diffusion features, and injects retargeted flow through TransPE, a positional encoding transfer mechanism that rewires attention without training new weights.
Motion4Motion is useful for researchers, animators, and generative video developers who need cross-species or cross-morphology animation without a manually defined rig. The project reports strong benchmark results across animal and human motion transfer and demonstrates applications like transferring human gait onto a table through flow-based attention manipulation.


