Key Features

Transfers motion from a source video to a target subject at inference time.
Avoids skeleton templates, making cross-species and cross-topology transfer possible.
Uses dense point trajectories as topology-agnostic motion flow.
Builds on a frozen WAN text-to-video diffusion transformer.
Uses Grounded SAM-2, point tracking, and diffusion-feature matching to align source and target subjects.
Introduces TransPE to transfer positional encoding into the denoiser attention path.
Requires no additional training for a new target subject.
Reports strong results on animal and human motion-transfer benchmarks.

The method operates on top of a frozen WAN text-to-video diffusion transformer. It samples source anchors with Grounded SAM-2, tracks their trajectories, matches them to the target subject with diffusion features, and injects retargeted flow through TransPE, a positional encoding transfer mechanism that rewires attention without training new weights.


Motion4Motion is useful for researchers, animators, and generative video developers who need cross-species or cross-morphology animation without a manually defined rig. The project reports strong benchmark results across animal and human motion transfer and demonstrates applications like transferring human gait onto a table through flow-based attention manipulation.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner
Zero to AI Engineer Program

Zero to AI Engineer

Skip the degree. Learn real-world AI skills used by AI researchers and engineers. Get certified in 8 weeks or less. No experience required.

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!