SGToolkit: An Interactive Gesture Authoring Toolkit for Embodied Conversational Agents

Agent Personality & AnthropomorphismHuman-Robot Collaboration (HRC)Game Developers & DesignersMusicians, DJs & Sound Designers

Title of the Paper

SGToolkit: An Interactive Gesture Authoring Toolkit for Embodied Conversational Agents

Paper Information

  • Research Area: Human-Computer Interaction, Artificial Intelligence, Virtual Character Animation
  • Keywords: Speech gestures, social behavior, gesture authoring, toolkit, virtual agents, neural generative models

Research Background and Problem

  • Identified Issues or Challenges: While non-verbal behaviors are crucial for embodied artificial agents such as social robots and virtual characters, existing keyframe animation and motion capture methods are costly and unsuitable for large-scale speech gesture scenarios. Additionally, the quality of automatic generation methods is limited, making it difficult for designers to make modifications.
  • Significance: Appropriate gesture actions can reveal character intentions, capture audience attention, and enhance the affinity between human-computer interaction roles.
  • Research Motivation: Existing toolkits (e.g., USC's Virtual Humans Toolkit and SoftBank Robotics' NAOqi) rely on predefined rules, with limited quality in speech gesture generation and lack of platform independence. Furthermore, the absence of interactive gesture editing features prevents gestures from accurately expressing designers' intentions.

Solution

  • Proposed Solution: Development of SGToolkit, a speech-driven gesture generation toolkit that integrates automatic generation and manual control, supporting fine-grained posture control and coarse-grained style control.
  • Innovations:
    • Combines data-driven models with user manual control to deliver high-quality gesture outputs.
    • Provides a platform-independent toolkit supporting both speech text and audio inputs.
  • Implementation Steps and Key Technologies:
    1. Design a neural generative model architecture that combines automatic gesture generation based on speech content with user input controls (posture and style).
    2. Train the model using the TED dataset, which includes 97 hours of TED talk videos and 3D human pose data.
    3. Develop an intuitive interactive interface with user participation, integrating web applications and database support to ensure flexibility and usability of the toolkit.

Research Outcomes

  • Specific Results:
    • SGToolkit efficiently generates speech-driven gestures, allowing users to easily adjust gestures based on automatically generated drafts.
    • User studies indicate that compared to manual keyframe animation, the toolkit significantly improves generation efficiency and gesture quality under time constraints.
    • Experiments show that the trained model accurately follows user posture and style controls while also supporting fully automated generation.
  • Advantages:
    • Generated gestures surpass manual keyframe animation in anthropomorphism and speech consistency within limited timeframes.
    • Supports gesture adjustments based on user intent, offering greater flexibility than purely automatic generation methods.
    • Provides a complete open-source API interface compatible with various platforms.
  • Experiment and Evaluation Results:
    • User studies and crowdsourced quality evaluations demonstrate that gestures generated by the toolkit achieve high levels of speech appropriateness and human-likeness, comparable to real human movements.
    • Under interactive conditions, users rated the tool highly (intention expression: M=4.7) and found the generated gestures better suited to speech content (M=4.7).
  • Limitations and Future Directions:
    • Output gestures may slightly deviate from user posture controls, especially when unreasonable postures are input. Future improvements could include post-processing functions to optimize such cases.
    • The toolkit currently does not support facial expressions and hand movements. It is recommended to improve the dataset or integrate existing generation modules to achieve comprehensive motion generation.
    • Further expansion of application scenarios in social robots and virtual agents is suggested, enhancing the model's multimodal adaptability.

Conclusion

SGToolkit combines automatic and interactive gesture generation, providing developers with an effective solution for more precise control while significantly reducing the time and cost of gesture creation. It has broad application potential in human-computer interaction, gaming, and education. Future improvements include comprehensive gesture generation for better completeness and optimization of user requirements.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/61398/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3472749.3474789
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Agent Personality & Anthropomorphism, Human-Robot Collaboration (HRC)
work
Professions
Game Developers & Designers, Musicians, DJs & Sound Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers