The method combines an autoregressive diffusion model with a hybrid motion representation that separates explicit root motion from a compact latent body embedding. A two-stage transformer denoiser generates motion in a streaming fashion while accepting online text prompts, root paths, waypoints, full-body keyframes, sparse joint positions, and rotations as constraints.
ARDY is useful when animators or robotics researchers need controllable motion that can respond as a scene evolves rather than waiting for offline generation. The project links to paper, code, and model resources, and demonstrates text-to-motion, kinematic constraint following, long-horizon goals, and interactive humanoid control.


