UniMo: Unifying Human and Animal Motion Generation
UniMo is pushing 3D motion generation toward a unified setup that can handle both human and animal bodies without building species-specific models. The core trick is to convert parametric skeletons into an unparametric point-cloud representation, which avoids the topology mismatch that makes animals awkward to model alongside humans.
The system also uses dynamic sampling to concentrate more points around active joints, which should matter for motion fidelity where limbs and articulation change quickly. For teams working in animation, virtual production, or game tooling, the practical appeal is obvious: one representation that can generalize across very different rigs is easier to scale than a stack of per-species solutions.
The bigger data contribution is UniML3D, a motion-language dataset spanning human and animal categories with 145,907 motion sequences and 433,388 captions. That is more than 102 times larger than existing animal motion datasets, which helps explain why the model can train and evaluate across a broader range of movement patterns and prompts.
UniMo reports state-of-the-art results on UniML3D plus HumanML3D, KIT-ML, and AnimalML3D. For developers, the interesting takeaway is less about a finished production tool and more about a direction: if unified motion priors keep improving, they could reduce the cost of authoring creature animation, motion search, and text-driven motion generation across mixed character sets.
“unified modeling across species difficult”
- what
- UniMo is a unified point-cloud motion generation framework for humans and animals.
- who
- Zeyu Zhang, Zhiyuan Zhang, Siheng Wang, Yiran Wang, Danning Li, Ian Reid, and Richard Hartley.
- when
- Submitted on 11 Sep 2026.
- impact
- Could reduce the need for separate motion models for human and creature animation workflows.
Promising technical advance for motion generation
Follow motion generation updates
See relevant stories in your personalized news feed.
Discussion