ETHead: Generating Expressive 3D Facial Animation and Head Movement from Speech
ETHead tackles a familiar bottleneck in speech-driven facial animation: there just isn’t enough high-quality 3D performance data to learn rich emotional motion reliably. The new system focuses on generating both facial expression and head movement from speech, with the goal of matching the emotional content of the audio instead of producing generic lip-sync and idle motion.
The core idea is a self-distillation pipeline that pre-trains a specialized speech encoder on large-scale 2D talking videos. That encoder is then shaped with an emotion-modulated probabilistic masking scheme, which helps align audio features with expressive visual dynamics. In practice, ETHead builds a joint speech-motion latent space that provides explicit supervision for 3D generation, giving the model stronger cues for both expression and head pose.
For developers working on avatars, virtual production, or NPC dialogue systems, the practical appeal is obvious: better expressiveness without needing a massive proprietary 3D capture set. The authors say the motion-aligned speech encoder is transferable, so it could be dropped into other talking-head frameworks as a plug-in style upgrade rather than a full pipeline replacement.
The reported results are strong enough to claim a clear step up over current state of the art, though the exact deployment constraints, runtime cost, and production integration details haven’t been disclosed yet. Even so, the direction is important for teams trying to bridge the gap between audio-driven animation research and shippable character performance.
“Generating expressive 3D talking heads solely from speech remains a significant challenge”
- what
- ETHead generates expressive 3D facial animation and head movement directly from speech.
- who
- Jiu-Cheng Xie, Jiwang Zheng, Yongkang Xia, Jian Xiong, Chi-Man Pun, Hao Gao, and Feng Xu.
- when
- Submitted to arXiv on 3 Aug 2026.
- impact
- Could improve speech-driven character animation and be reused in other talking-head systems.
Promising quality boost for speech-driven character animation
Follow facial-animation updates
See relevant stories in your personalized news feed.
Discussion