DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion
arXiv cs.GR details DynaConTalk, a wavelet-constrained diffusion framework for long-form, controllable holistic co-speech 3D motion. The core idea is to move generation into stationary wavelet transform coefficient space, so coarse posture, mid-frequency gesture strokes, and fine facial or hand detail are handled in separate bands instead of being averaged together.
For developers building character animation or AI-assisted performance tools, the practical angle is control. DynaConTalk keeps HuBERT and speaker identity as a base, then selectively injects rhythm, mel, and transcript cues through motion-state- and noise-aware residual gates. That should help avoid the common failure mode where dense audio cues drown out content-specific motion intent.
The system also adds attention pooling, learned depth routing, and a frame-resolution rhythm path to preserve timing, while a signed proposal-consensus update reconciles the conditions during denoising. Matched-noise constraint injection supports history continuation, localized keypose repair, and reference-guided control through the same sampling interface.
On the output side, separate body-hand and facial denoisers feed an inverse SWT stage and a pose-driven root regressor. The paper says experiments cover generation quality, facial accuracy, temporal continuity, and controllable editing, and the code, models, and interactive editing interface are available.
“Diffusion operates in stationary wavelet transform (SWT) coefficient space”
- what
- DynaConTalk is a wavelet-constrained diffusion framework for long-form, controllable co-speech 3D motion
- who
- Authors include Yifei Zhu, Yangyang Cai, Mingyi Shi, Miao Cheng, Lin Gu, Taku Komura, and Yoshifumi Kitamura
- when
- Submitted to arXiv cs.GR on 7 Oct 2026
- impact
- Aims to reduce over-smoothed motion and improve controllability for speech-driven character animation
Promising control and fidelity gains for speech-driven animation
Follow animation updates
See relevant stories in your personalized news feed.
Discussion