Exploring the Design of Human Speech Indicators to Enhance Waiting Experience in Voice User Interface

Voice User Interface (VUI) DesignAgent Personality & Anthropomorphism

Research Background and Issues

  • Identified Challenges and Issues: The system loading wait experience in Voice User Interfaces (VUIs) often impacts user satisfaction, leading to dissatisfaction. Traditional visual feedback mechanisms (e.g., progress bars) cannot be directly applied to popular voice assistants (e.g., Siri, Alexa), which typically lack effective waiting feedback mechanisms. Therefore, this study focuses on optimizing the waiting experience in VUIs, particularly through the design of auditory indicators.
  • Significance: With the increasing prevalence of voice assistants, their interactions are becoming more akin to interpersonal communication. Designing auditory indicators that effectively convey system status while enhancing user satisfaction through human-like speech is particularly important.
  • Research Motivation and Related Work: Although previous studies have attempted to use non-verbal auditory indicators (e.g., beeps, tapping sounds) to improve the waiting experience, these indicators can only convey limited information. Prior research has shown that humorous or explanatory speech can enhance users' perception of the social capabilities of voice assistants. Additionally, human speech helps build trust and increases emotional engagement, but its application in system loading wait scenarios remains underexplored.

Solution

  • Proposed Solution: The authors propose designing human speech indicators by integrating elements of "explanation" and "humor" to optimize the waiting experience in VUIs.
  • Innovative Features:
    1. Combining "high explanatory" and "high humorous" speech indicators to enable users to understand the system's loading status while maintaining emotional engagement.
    2. Proposing differentiated speech designs based on varying loading durations (single-element designs for short loading times, multi-element designs for longer loading times).
    3. Providing practical design guidelines, such as personalized content, diverse feedback, and emotional enhancement.
  • Implementation Steps and Key Techniques:
    1. Focus Groups: Recruiting 35 users with experience using voice assistants to discuss their waiting experiences and suggestions.
    2. Experimental Testing: Designing four types of speech indicators (based on high/low humor and explanatory levels) and a control group (no speech indicators), tested in two scenarios: short loading (5 seconds) and long loading (15 seconds).
    3. Data Collection: Using questionnaires to evaluate user attention, perceived time, enjoyment, and overall satisfaction; followed by semi-structured interviews for in-depth feedback analysis.

Research Outcomes

  • Specific Findings:
    1. Experimental results show that in long loading scenarios, speech indicators combining "explanation" and "humor" performed best; in short loading scenarios, indicators with single high explanatory or humorous elements were more effective.
    2. Confirmed the positive impact of "humor" and "explanation" on users' waiting experience, particularly in enhancing attention and enjoyment.
    3. Provided design guidelines for voice assistant interactions based on experimental data.
  • Advantages Over Existing Solutions:
    1. Compared to traditional non-verbal feedback or abstract sound indicators, human speech combining humor and explanation offers greater social and emotional value, helping to build user trust.
    2. Scenario-specific strategies highlight design flexibility, such as selecting different elements based on duration to avoid cognitive overload for users.
  • Experimental or Evaluation Results:
    1. Users generally recognized the role of humorous content in enhancing the waiting experience, with "high humorous" speech indicators significantly improving enjoyment.
    2. "High explanatory" indicators provided transparency, reducing user confusion and enhancing understanding of system status.
    3. In slow loading scenarios, speech indicators combining explanation and humor performed best, while fast scenarios required less complex feedback.
  • Limitations and Future Directions:
    1. Questionnaire Design Limitations: Single-scale measurements were limited to exploratory goals; future studies should develop multi-scale metrics for broader evaluations.
    2. Lack of Personalization: The speech content in the experiments was predefined, failing to fully meet users' personalized needs. Future work could leverage large language models to generate real-time dynamic personalized content.
    3. Cultural and Linguistic Diversity: Participants were predominantly Mandarin speakers, which did not adequately account for the global linguistic and cultural diversity of voice assistant users.
    4. Unexplored Design Dimensions: Beyond humor and explanation, dimensions such as emotional tone and vocal characteristics also warrant further investigation in future studies.

This study provides significant theoretical support and practical guidance for the design of human speech indicators in VUIs' system loading wait scenarios, laying a solid foundation for subsequent research.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188429/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713090
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Voice User Interface (VUI) Design, Agent Personality & Anthropomorphism
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers