Skip to main content
GameDev.net gamedev.net
2 sources covering this story

GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling

A new co-speech gesture system aims to make digital humans move more naturally while cutting inference from dozens of steps to a faster flow-matching pipeline. GestureLSM models body-region interactions with spatial and temporal attention, then adds latent shortcut learning to improve quality. The result is state-of-the-art performance on BEAT2 with a practical speedup for real-time embodied agents.

First reported 1 month, 1 week ago • animation ai digital-humans gesture-generation
Want to discuss what this means for developers?
Open discussion
PRIMARY SOURCE
arXiv cs.GR arXiv cs.GR

GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling

1 month, 1 week ago Read source
arXiv cs.GR arXiv cs.GR 1% match

EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

1 month, 1 week ago Read source