SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching
Audio-driven facial animation has gotten much better at lip sync, but upper-face motion still tends to look frozen or mechanically repetitive. SubtleTalk targets that gap by generating weakly correlated dynamics such as eyebrow movement, eye blinks, and head motion, which are the details that often separate a decent talking head from something that feels alive.
The core idea is to move beyond speech-only conditioning. The system adds explicit controls for prosody, regional intensity, and valence-arousal signals so creators can steer timing, magnitude, and emotional variation instead of relying on a single deterministic prediction from audio. That matters for production because the same line can need very different facial behavior depending on character, scene tone, or performance style.
To handle the fact that upper-face motion is inherently variable, SubtleTalk uses residual flow matching on top of a stable speech-driven motion prior. In practice, that gives the model room to generate stochastic deviations instead of collapsing to the same safe motion every time. The result is meant to preserve lip synchronization while making the rest of the face feel less scripted.
The project also includes SubtleTalk-Face, a large-scale dataset with roughly 3,900 identities and 74 hours of 3D facial animation. It was built through a scalable pseudo-labeling pipeline and includes improved upper-face tracking plus frame-level valence-arousal annotations. For teams working on avatars, virtual production, or character-driven AI assistants, this is a useful signal that controllable...
“weakly correlated dynamics... remain difficult to model faithfully”
- what
- SubtleTalk is a controllable facial animation framework for 3D talking heads that generates weakly correlated dynamics beyond lip sync.
- who
- Chenyang Ding, Shuai Tan, Qunfen Lin, Xinwei Jiang, Zijiao Zeng, and Ye Pan.
- when
- Submitted to arXiv on 3 Aug 2026.
- impact
- Could help teams generate more natural eyebrow, blink, head, and emotional motion without sacrificing speech alignment.
Promising step toward more lifelike, controllable facial animation
Follow facial-animation updates
See relevant stories in your personalized news feed.
Discussion