A New Uncanny Valley? The Effects of Speech Fidelity and Human Listener Gender on Social Perceptions of a Virtual-Human Speaker

Intelligent Voice Assistants (Alexa, Siri, etc.)Agent Personality & Anthropomorphism

Structured Literature Review and Key Insights

Title of the Paper

A New Uncanny Valley? The Effects of Speech Fidelity and Human Listener Gender on Social Perceptions of a Virtual-Human Speaker

Bibliographic Information

  • Field of Study: The impact of speech synthesis technology on social perceptions of virtual humans
  • Keywords: Virtual humans, social perception, speech synthesis technology, gender effects, human-computer interaction

Research Background and Problem

Issues and Challenges

  • Virtual humans are widely used as persuasive agents, but the impact of their speech delivery (especially using text-to-speech technology) on social perceptions remains poorly understood.
  • Synthetic speech is often perceived as less trustworthy and natural compared to human-recorded speech. While neural text-to-speech (Neural TTS) technology is closer to human speech than existing technologies, its effects on social perception require further investigation.
  • It remains unclear whether the listener's gender influences social perceptions of virtual human speech.

Significance

  • Social perception is critical for the acceptance and user experience of virtual humans, playing a key role in trust and technology adoption. In persuasive scenarios such as medical consultations and energy conservation, these perceptions directly impact application effectiveness.

Motivation and Related Work

  • There is a lack of consensus in the analysis of the relationship between speech quality and social perception. Early studies suggest differences in social perception between synthetic and human speech, but the results are inconsistent.
  • Although Neural TTS is considered more natural and closer to human speech, its ability to enhance the social perception of virtual humans remains unverified.
  • Research on the influence of gender on technology acceptance is often superficial, and this study aims to address this gap.

Proposed Solution

Methods or Approach

  • A 3×2 experimental design was employed to evaluate the effects of three speech technologies (Standard TTS, Neural TTS, and human-recorded speech) on the social perception and persuasive effectiveness of virtual humans, considering the gender of the listener (male and female).
  • The study assessed social perception metrics for the virtual human, including speaker competence, trustworthiness, and courage, as well as perceptions of the speech content, arguments, and voice.

Innovations

  • This study is the first to directly compare Neural TTS with Standard TTS while exploring social perception differences from the perspective of listener gender.
  • It builds on earlier methodologies but incorporates modern TTS technologies, particularly Neural TTS, providing a foundation for improving virtual human design in the future.

Implementation Steps and Key Techniques

  1. Create three speech conditions (Standard TTS, Neural TTS, and Human Speech), synchronizing all speech with the virtual human.
  2. Use questionnaires and ratings to evaluate social perception and persuasive effectiveness.
  3. Generate Standard and Neural TTS speech using Amazon Polly tools.
  4. Analyze participant response data using Principal Component Analysis (PCA) and Analysis of Variance (ANOVA).

Research Findings

Specific Results

  • TTS (text-to-speech) conditions did not significantly improve the social perception of virtual humans. Virtual humans using Neural TTS were perceived as less trustworthy compared to those using Standard TTS.
  • Not all listeners responded equally to speech quality: male listeners rated trustworthiness significantly lower than female listeners.
  • Negative evaluations of synthetic speech may be linked to the "uncanny valley" effect, which describes discomfort caused by speech that approximates but fails to fully replicate human-like qualities.

Advantages Over Existing Solutions

  • Provides empirical data on whether modern TTS technologies can enhance user experience.
  • Explores the complex interplay between listener gender and social perception, offering guidance for designing gender-sensitive virtual humans.

Limitations and Future Directions

  • The experiment was limited to female virtual humans and predefined persuasive topics, excluding studies on virtual humans of different genders or ethnicities.
  • The listener sample primarily consisted of U.S. university students, which may not represent the general population.
  • Future research should expand to other cultural contexts, more diverse virtual human designs, and integrate mixed quantitative and qualitative methods for data collection.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/72043/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3517564
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Agent Personality & Anthropomorphism
work
Professions
article
Content Status
Full text indexed
hub
Related Papers
10 related papers