Skip to main content
GameDev.net gamedev.net
Research Paper

This is an academic paper or technical research. Key findings may require technical background to fully understand.

Explore Research Radar

PRO Tired of ads? Read GameDev.net ad-free and help keep the community independent with GameDev Pro — $3/month.

arXiv cs.GR
arXiv cs.GR Research
· 1 day, 13 hours ago • Yifei Zhu, Yangyang Cai, Mingyi Shi, Miao Cheng, Lin Gu, Taku Komura, Yoshifumi Kitamura

DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion

Briefing

arXiv cs.GR details DynaConTalk, a wavelet-constrained diffusion framework for long-form, controllable holistic co-speech 3D motion. The core idea is to move generation into stationary wavelet transform coefficient space, so coarse posture, mid-frequency gesture strokes, and fine facial or hand detail are handled in separate bands instead of being averaged together.

For developers building character animation or AI-assisted performance tools, the practical angle is control. DynaConTalk keeps HuBERT and speaker identity as a base, then selectively injects rhythm, mel, and transcript cues through motion-state- and noise-aware residual gates. That should help avoid the common failure mode where dense audio cues drown out content-specific motion intent.

The system also adds attention pooling, learned depth routing, and a frame-resolution rhythm path to preserve timing, while a signed proposal-consensus update reconciles the conditions during denoising. Matched-noise constraint injection supports history continuation, localized keypose repair, and reference-guided control through the same sampling interface.

On the output side, separate body-hand and facial denoisers feed an inverse SWT stage and a pose-driven root regressor. The paper says experiments cover generation quality, facial accuracy, temporal continuity, and controllable editing, and the code, models, and interactive editing interface are available.

“Diffusion operates in stationary wavelet transform (SWT) coefficient space”

— Paper abstract · Core representation choice
Original source
Read on arXiv cs.GR
At a glance
what
DynaConTalk is a wavelet-constrained diffusion framework for long-form, controllable co-speech 3D motion
who
Authors include Yifei Zhu, Yangyang Cai, Mingyi Shi, Miao Cheng, Lin Gu, Taku Komura, and Yoshifumi Kitamura
when
Submitted to arXiv cs.GR on 7 Oct 2026
impact
Aims to reduce over-smoothed motion and improve controllability for speech-driven character animation
Signal Positive

Promising control and fidelity gains for speech-driven animation

Discuss

Follow animation updates

See relevant stories in your personalized news feed.

Sign in to follow

Continue on GameDev.net

Useful next steps related to this story.

Game development news without the noise

One useful weekly briefing. No daily flood.

Sending your confirmation email…

Discussion

Loading comments...