Designing for Speech Practice Systems: How Do User-Controlled Voice Manipulation and Model Speakers Impact Self-Perceptions of Voice?

Voice User Interface (VUI) DesignVoice Accessibility

Title of the Paper

Designing for Speech Practice Systems: How Do User-Controlled Voice Manipulation and Model Speakers Impact Self-Perceptions of Voice?

Paper Information

  • Domain: Human-Computer Interaction (HCI), public speaking training, self-perception, and voice modification technologies
  • Keywords: Public speaking, speech training, voice manipulation, model audio, self-perception, gender and voice, user experience research

Research Background and Questions

  • Background:

    • Speaking skills are a critical component of human communication, but identifying effective improvement methods can be challenging.
    • Current speech training systems often emphasize external evaluations and feedback, with limited focus on users' self-perception and self-assessment abilities.
    • Women and gender minority groups, in particular, face societal pressure to emulate the "authoritative" deep voices associated with men, which can negatively impact their self-perception.
  • Research Questions:

    • How can user-controlled voice manipulation and the use of model audio enhance users' perception of their own voice and speaking abilities?
    • Can a speech practice system be designed to better support users in improving self-perception?
  • Motivation:

    • Existing research indicates that self-perception can improve speaking skills, confidence, and deeper understanding of communication abilities, ultimately influencing social well-being.
    • By exploring the combination of voice manipulation and model audio, the authors aim to provide more personalized and effective solutions for gender minority groups, enhancing their awareness of their own voice and speaking abilities.

Solution

  • Methods or Solution:

    • Propose a speech practice model that combines user-controlled voice manipulation and model audio.
    • Design experiments to analyze the impact of these two methods (used individually or in combination) on speakers' self-perception.
  • Innovations:

    • Introduce user-controlled voice manipulation (e.g., pitch and duration adjustments via Praat software) with highly recognizable interactive designs.
    • Emphasize users' self-assessment rather than relying solely on external system evaluations.
    • Investigate the potential influence of social and gender factors on voice perception, specifically for women and gender minority groups.
  • Implementation Steps:

    • Use quantitative analysis and qualitative interviews to study the effects of voice manipulation and model audio on users' confidence, satisfaction, communication, and self-expression abilities.
    • Design four experimental conditions (combining or separately using voice manipulation and model audio) to compare their impact on enhancing self-perception.
    • Apply thematic analysis to process interview and survey data, extracting key recommendations for designing speech training systems.

Research Findings

  • Specific Findings:

    1. Quantitative Experiment Analysis:

      • Model audio significantly improved users' confidence in tone, while voice manipulation did not show a notable impact on self-perception in the quantitative analysis.
      • Voice manipulation demonstrated a positive trend in enhancing self-expression.
    2. Interviews and Thematic Analysis:

      • Voice manipulation software (Praat) was widely appreciated, helping users understand their voice characteristics through immediate auditory feedback and audio visualization.
      • Most users provided positive feedback on model audio, although some felt pressure when comparing their voice to the model audio.
      • Participants reported that social and gender identity significantly influenced their perception of voice and speaking.
    3. Changes in Self-Perception:

      • Users showed a significant improvement in speech satisfaction compared to pre-experiment levels, but confidence, communication skills, and self-expression abilities did not show notable enhancement.
  • Comparison with Existing Solutions:

    • Encouraged users to actively engage in the voice adjustment process, increasing their involvement and self-expression abilities.
    • Focused on personalized design in speech practice, supporting users in optimizing skills based on their unique characteristics.
  • Limitations and Future Directions:

    • Limitations:
      • Small sample size, consisting entirely of women or gender minority groups, limiting generalizability.
      • Practice content was based on provided template texts, without considering user-customized texts or live audience scenarios.
      • Focused solely on English-speaking environments, making results inapplicable to multilingual or global contexts.
    • Future Directions:
      • Expand sample size to include more gender groups and linguistic backgrounds.
      • Explore the potential impact of dynamic audience feedback on self-perception.
      • Further develop personalized audio recommendation systems to address current algorithmic limitations.

Design Implications

  • User-Customized Voice Manipulation:

    • Provide users with more control options, allowing them to adjust specific voice parameters such as pitch, frequency, and rhythm as needed.
    • Enhance user education to help them understand how voice manipulation can improve speaking skills.
  • Personalized Model Audio:

    • Recommend model audio with similar attributes to users' voice characteristics and speaking goals to improve relevance and practicality.
    • Ensure diversity in model audio collections, including speakers of different genders, social backgrounds, and accents to meet varied needs.
  • Annotations and Visualization Support:

    • Offer script annotation features to help users record and recall key points when modifying their voice.
    • Use real-time feedback and visualization tools to improve the usability and comprehensibility of voice manipulation.

Contribution of the Paper

  • Provides a new perspective on designing speech training systems, helping users improve speaking skills through self-assessment and self-perception.
  • Highlights the importance of gender and social factors in voice perception research, offering inspiration for future studies and designs.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68730/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3502093
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Voice User Interface (VUI) Design, Voice Accessibility
work
Professions
article
Content Status
Full text indexed
hub
Related Papers
4 related papers