DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation
DC-Motion tackles a familiar problem in text-to-motion: you want both high-level planning and believable micro-motion, but most methods force a compromise. The paper argues that action semantics like intent, phase changes, and timing are naturally discrete, while joint trajectories are continuous, so it factorizes motion into two parts instead of trying to represent everything one way.
The system uses a text-conditioned structure generator to predict discrete structural tokens through iterative masked modeling. A separate diffusion-based residual generator then fills in the continuous motion details conditioned on that structure. In practice, that means the model can plan the “shape” of an action first, then refine the local dynamics afterward.
For game teams, the interesting part is not just the benchmark numbers, but the representation choice. If this holds up beyond paper settings, it could make text-driven animation generation easier to steer, especially for tools that need predictable action beats without losing natural motion. That matters for prototyping, NPC behavior authoring, and any pipeline trying to bridge designer intent with animation output.
The paper was submitted on 28 May 2026 and revised on 6 Jul 2026. It reports strong results on HumanML3D and KIT-ML, claiming better FID and R-Precision than representative diffusion and discrete-token baselines. As usual with motion papers, the real question for developers is whether the structure/detail split survives messy production data, style variation, and runtime constraints.
“action semantics ... are inherently discrete and compositional”
- what
- DC-Motion factorizes human motion generation into discrete structural tokens and continuous residual latents.
- who
- Authors: Hequan Wang, Xuean Chen, Jiaxu Zhang, Zhengbo Zhang, and Zhigang Tu.
- when
- Submitted 28 May 2026; revised 6 Jul 2026 (v2).
- impact
- Could improve text-driven animation control and motion quality for prototyping, NPCs, and animation tools.
Promising control/quality tradeoff for motion generation
Follow AI updates
See relevant stories in your personalized news feed.
Discussion