Giving Robots a Voice: Human-in-the-Loop Voice Creation and open-ended Labeling

Agent Personality & AnthropomorphismSocial Robot InteractionSoftware Engineers & DevelopersUI/UX DesignersAI/ML Researchers & Engineers

Title of the Paper

Giving Robots a Voice: Human-in-the-Loop Voice Creation and Open-ended Labeling

Paper Information

  • Subject Area: Human-Computer Interaction, Robot Persona Design, Voice Synthesis
  • Keywords: Human-Computer Interaction, Robot Voice, Voice Generation Tools, Open-ended Labeling, Voice Prediction, Text-to-Speech Synthesis

Research Background and Problem

  • Problem or Challenge: Aligning robot appearance with voice is critical for improving robot usability. However, current research explores limited dimensions of robot voice, and existing voice types lack diversity. Additionally, generating voices suitable for various robot appearances remains a technical and methodological challenge.
  • Significance: Robot voice is a core component of human-computer interaction, capable of conveying intent, personality, and emotions. A mismatch between robot appearance and voice can lead to discomfort, a sense of eeriness, or even resistance, negatively impacting user experience and functionality.
  • Motivation and Related Work:
    • Existing studies emphasize the relationship between visual appearance and voice but are limited in analytical dimensions and dataset size.
    • State-of-the-art Text-to-Speech (TTS) synthesis technologies can produce realistic human voices but lack tools tailored for synthesizing voices for diverse robot types.
    • Voice selection in human-computer interaction is still considered an expert task, with manual operations being inefficient and prone to subjective bias.

Solution

  • Method and Solution:
    • A five-step method is proposed to create and predict robot voices:
      1. Develop a Voice Creation Tool: Integrate the latest TTS technology with traditional signal processing to cover a spectrum from highly synthetic to natural voices.
      2. Match Voices with Robot Appearance: Employ a human-in-the-loop voice generation paradigm (Gibbs Sampling with People, GSP), allowing participants to iteratively adjust parameters to find the most suitable voice for each robot image.
      3. Identify Robot Perception Attributes: Use literature review and an adaptive labeling mining method (Sequential Transmission Evaluation Pipeline, STEP-Tag) to extract perception labels associated with robot personas.
      4. Rate Attribute Dimensions: Conduct multidimensional evaluations of robot appearance images and voice samples.
      5. Predict New Robot Voices: Use a matching model to predict suitable voices for new robots.
  • Innovations:
    • Combines cognitive science and machine learning to provide a novel framework for developing voice generation tools.
    • Offers open voice generation and prediction tools for engineers.
    • Systematically summarizes the perceptual attributes between robot personas and voices, enabling data-driven matching and prediction.
  • Implementation Steps and Key Techniques:
    • Voice Generation: Use TTS technology to generate natural voices, combining audio effects (e.g., pitch variation, distortion) to adjust voice characteristics.
    • Label Mining: Create open-ended labels for robot images or voices using STEP-Tag, converting participants' perceptions into measurable datasets.
    • Matching Model: Predict suitable voices based on experimental data and image ratings.

Research Outcomes

  • Specific Results:
    • Designed and validated a robot voice creation tool covering a wide range of voice types.
    • Created a large dataset comprising 175 robots and their matched voices, spanning multiple application scenarios.
    • Developed an online interface tool for engineers to customize robot voices and predict voices for new robots.
    • Demonstrated that perceptual dimensions can effectively predict suitable voices for unseen robots.
  • Advantages:
    • Systematically addresses multidimensional voice matching from both technical and cognitive perspectives.
    • Resolves the alignment issue between robot appearance and voice while developing practical application tools.
  • Experimental and Evaluation Results:
    • Validated with 2,505 participants, showing significant improvement in voice matching scores after multiple iterations.
    • Independent validation experiments demonstrated high accuracy in predicting new robot voices.
    • Robustness tests of the dataset showed results are generalizable across datasets.
  • Limitations and Future Directions:
    • Did not consider the impact of robot motion or dynamic interaction on voice matching.
    • The dataset and methods are primarily based on English and have not been extended to multilingual or non-Western cultural contexts.
    • Voice models do not yet achieve perfect matching, requiring higher-dimensional voice parameterization.
    • Future directions include exploring the perceptual impact of dynamic robot materials (e.g., videos) and long-text voice content, as well as extending research to different cultural and linguistic environments.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147115/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642038
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Agent Personality & Anthropomorphism, Social Robot Interaction
work
Professions
Software Engineers & Developers, UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
3 related papers