A Benchmark for Spatially Grounded Gesture Generation
arXiv cs.GR details a benchmark for spatially grounded gesture generation that tries to solve a familiar problem for conversational avatars: a gesture can look natural while still pointing at the wrong thing. The work introduces a dataset of roughly 2,000 pointing-annotated clips captured from naturalistic VR dialogue, each paired with ground-truth 3D referents.
The benchmark is designed to separate three concerns that are often blended together in evaluation: temporal alignment, spatial grounding, and perceived naturalness. That matters for anyone building NPCs, social VR characters, or embodied assistants, because a model that wins on motion realism may still fail at the core communicative task.
The paper also includes a flow-matching baseline, MM-Conv-Flow, and compares it with an independent retrieval-based system and human motion capture. One notable result is that geometric grounding can surpass human pointing without improving perceived naturalness, which is a useful warning for teams relying on generic motion metrics.
For developers, the practical takeaway is straightforward: referential gesture quality needs separate metrics, not a single “good gesture” score. If your pipeline touches animation generation, avatar interaction, or multimodal dialogue, this benchmark gives you a more realistic way to test whether gestures actually support communication.
“no common framework exists for evaluating whether generated gestures indicate their intended referent”
- what
- A benchmark for spatially grounded gesture generation was introduced
- who
- Anna Deichler, Rishabh Dabral, Fethiye Irmak Dogan, Anindita Ghosh, and Jonas Beskow
- when
- Submitted to arXiv on 2 Oct 2026
- impact
- Helps evaluate whether generated pointing gestures identify the correct referent, not just look natural
Useful evaluation advance, but exposes hard quality gaps
Follow AI updates
See relevant stories in your personalized news feed.
Discussion