A large, crowdsourced evaluation of gesture generation systems on common data: The GENEA Challenge 2020

Full-Body Interaction & Embodied InputHuman Pose & Activity RecognitionConversational ChatbotsUI/UX DesignersAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

A Large, Crowdsourced Evaluation of Gesture Generation Systems on Common Data: The GENEA Challenge 2020

Paper Information

  • Subject Area: Speech-driven gesture generation, human-computer interaction evaluation
  • Keywords: Gesture generation, conversational agents, evaluation framework, data-driven, user study, human-computer interaction, dataset standardization, motion generation, multimodal communication

Research Background and Problem

  • Problems or Challenges:

    1. Gestures play a crucial role in speech interaction, but the technology for automatic gesture generation is still underdeveloped.
    2. Current research in the field of gesture generation is difficult to compare due to differences in datasets, evaluation standards, and visualization methods used.
    3. The lack of a common dataset and standardized evaluation framework hinders the further development and comparison of state-of-the-art techniques.
  • Significance: Automatic gesture generation is a key technology for building conversational agents (ECAs) capable of natural interaction. Compared to relying solely on speech, multimodal interaction combining gestures and speech can enhance the naturalness of communication and the efficiency of information transfer.

  • Research Motivation and Related Work:

    1. Data-driven methods are increasingly becoming the mainstream in the field of speech gesture generation, but there is a lack of reliable benchmarking.
    2. Drawing inspiration from challenges in other fields (e.g., the Blizzard Challenge in speech synthesis and the CLIC challenge in computer vision), the authors aim to establish the first challenge and benchmark in the field of gesture generation.
    3. The goal is to provide standardized training data and evaluation methods to enable fair and systematic comparison of different generation methods.

Solution

  • Method or Solution: The authors organized the "GENEA Challenge 2020" gesture generation competition:

    1. Provided a unified dataset: the Trinity Gesture Dataset, which includes speech and motion capture data.
    2. Defined a standardized evaluation framework: gestures were visualized using a unified 3D virtual character, and evaluations were conducted through large-scale user studies.
    3. Combined two evaluation criteria: the "human-likeness" of gestures and their "speech appropriateness."
  • Innovations:

    1. Controlled variables unrelated to gesture generation methods (e.g., dataset, evaluation criteria, and visualization).
    2. Provided open-source code, processed datasets, and standardized visualization and evaluation procedures.
    3. Conducted the first joint evaluation of multiple state-of-the-art gesture generation methods and made the generated results publicly available for future research.
  • Implementation Steps:

    1. Participants built models using the provided common training dataset.
    2. Generated corresponding gesture sequences for the challenge test data and submitted them for centralized evaluation.
    3. Assessed each method's "human-likeness" and "speech appropriateness" through large-scale user studies.

Research Outcomes

  • Specific Outcomes:

    1. Evaluated nine gesture generation systems, including five competition entries, two baseline methods, and two human motion reference conditions.
    2. Conducted large-scale, standardized user studies (with a total of 250 participants) to provide clear quantitative rankings of the gesture generation quality of each system.
    3. Released test data, generated results, and related experimental outcomes under the same standards for future research comparisons.
  • Advantages and Comparisons:

    1. Overall, the competition entries outperformed previous baseline methods, indicating progress in the field's generation quality.
    2. Different models showed varying performance in "human-likeness" and "speech appropriateness," suggesting that generation models can be optimized for specific metrics.
  • Experimental Results:

    1. Natural human motions scored significantly higher in human-likeness compared to other generation methods.
    2. Mismatched motion (motions unrelated to the speech) scored higher in speech appropriateness than most generation methods, indicating that current methods struggle to produce gestures strongly correlated with speech.
  • Limitations and Future Directions:

    1. The dataset includes only a single actor's English monologues, lacking support for multilingual and multi-scenario corpora.
    2. Gesture generation is limited to upper-body movements, excluding full-body or facial dynamics.
    3. The challenge did not separate semantic relevance from rhythmic relevance in evaluations, which may affect the accuracy of assessments.
    4. Future work could explore extending to highly interactive meeting scenarios and more complex multilingual datasets.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/57992/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3397481.3450692
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Full-Body Interaction & Embodied Input, Human Pose & Activity Recognition, Conversational Chatbots
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers