The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset
The fourth GENEA Challenge has put speech-driven gesture generation under a much harsher microscope. Using the Seamless Interaction dataset of dyadic conversations, organizers ran four separate human-evaluation studies to disentangle motion quality, speech alignment, interlocutor response, and semantic meaning instead of blending them into a single score.
That matters for anyone building conversational avatars, NPCs, VTubers, or virtual assistants: a gesture system can look smooth in isolation and still fail the moment timing, listening, or meaning enters the picture. The challenge collected over 23,000 votes from 869 test-takers, giving the field a much stronger read on where current models actually stand.
The numbers suggest the gap is still wide. Filtered dataset segments beat every challenge submission on motion realism by 68-95% in pairwise comparisons, while motion-capture clips set a 62% ceiling for speech alignment and 65% for appropriateness in dyadic response. The best submission reached 32% alignment, but most hovered near input-independent behavior; in the semantic task, the top system managed only 8% appropriateness even though humans could identify the matching transcript 79% of the time.
The challenge also introduces a new semantic gesture task built on the Grounded Gestures subset, plus a text-mismatching evaluation method. The collected votes and generated outputs will be released publicly, which should make this a useful benchmark for teams working on embodied AI, animation synthesis, and interactive character systems that need more than idle motion.
“collected over 23,000 votes from 869 test-takers”
- what
- The fourth GENEA Challenge evaluated five speech-driven gesture-generation systems with four large-scale human studies on the Seamless Interaction dataset.
- who
- Organizers and authors include Rajmund Nagy, Silvia Arellano García, Hendric Voss, Mihail Tsakov, Taras Kucherenko, Youngwoo Yoon, and Gustav Eje Henter.
- when
- The preprint was submitted on 11 Aug 2026.
- impact
- Results show current gesture models still struggle with alignment, dyadic response, and semantic expressiveness, despite decent-looking motion in some cases.
Useful benchmark, but current systems still fall short
Follow gesture-generation updates
See relevant stories in your personalized news feed.
Discussion