Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic Priors
Learn2Chat reframes conversational character animation as a layering problem instead of a single end-to-end dyadic generator. The core idea is to keep the speech-to-motion knowledge already learned by monologic models, then add a separate interaction signal for the social back-and-forth between speakers.
The framework uses a Monologic-Anchored Motion Factorization scheme to disentangle intrinsic audio-driven motion from interaction effects. On top of that, a Cross-Attentive Interaction Latent Prediction module maps paired speech streams into interaction latents, which then modulate canonical motion during inference. That makes the system more structured than a fully entangled dyadic model and helps it stay data-efficient.
For game teams building believable NPC conversations, this matters because conversational animation is one of the hardest places to scale realism without exploding content costs. A model that can reuse pretrained monologic backbones and adapt them to two-person exchanges could reduce the amount of paired dialogue data needed while improving sync and perceived naturalness.
The work was submitted to arXiv on July 11, 2026, and evaluated on the DualTalk benchmark, where it claims state-of-the-art quantitative and perceptual results. It is also model-agnostic, so the same interaction layer can be attached to different pretrained motion systems rather than forcing a full pipeline rewrite.
“interaction modulation over pretrained monologic motion priors”
- what
- Learn2Chat introduces a dyadic talking-head framework that modulates pretrained monologic motion with learned interaction latents.
- who
- Zikai Huang, Siyue Chen, Xuemiao Xu, Haoxin Yang, Cheng Xu, Yihong Lin, and Shengfeng He.
- when
- Submitted to arXiv on July 11, 2026.
- impact
- Could help teams generate more natural two-person dialogue animation with less paired training data and better reuse of existing motion models.
Promising reuse of existing motion models with better conversational realism.
Follow animation updates
See relevant stories in your personalized news feed.
Discussion