A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures

Honorable Mention
Intelligent Voice Assistants (Alexa, Siri, etc.)Voice AccessibilityHuman-LLM Collaboration

Document Title

A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures

Document Information

  • Subject Area: Human-Computer Interaction, Voice Assistant Trust Research
  • Keywords: Voice Assistant, Trust, Survey, Interview Study, Failure Cases, User Behavior, Natural Language Processing, Dataset, Reliability, User Experience

Research Background and Issues

  • Problems and Challenges:

    • Despite significant advancements in Natural Language Processing (NLP), voice assistants often fail to meet user expectations, particularly in accurately responding to complex tasks.
    • Users generally exhibit low levels of trust in voice assistants, which hinders their broader adoption.
  • Significance:

    • User trust is a critical factor for technology adoption, and voice assistants are increasingly being designed for complex tasks such as healthcare, mental health advice, and high-stakes decision-making.
    • Investigating the impact of voice assistant failures on user trust can help optimize technology design and enhance user experience.
  • Research Motivation and Related Work:

    • NLP models are typically evaluated using standardized datasets, which differ significantly from real-world application scenarios (e.g., mixed speech, incomplete inputs). This results in poor adaptability of the technology in actual user contexts.
    • Few studies have systematically examined user perceptions of voice assistant failures and the specific impact of these failures on trust.

Solution

  • Methods and Approach:

    • A mixed-methods approach is proposed to examine the impact of voice assistant failures on user trust.
    • Conduct interviews to gain in-depth insights into user experiences with voice assistant failures.
    • Build a dataset of user-reported failure cases and perform classification analysis.
    • Design and implement a survey to quantify the impact of different failure types on trust (across the dimensions of competence, benevolence, and integrity).
  • Innovations:

    • Extended existing NLP failure classification models to create a taxonomy of 12 failure sources specific to voice assistants.
    • Released a dataset containing 199 user-submitted failure cases, providing a foundation for future research.
    • Combined quantitative assessments of trust with real-world failure experiences, offering systematic insights.
  • Implementation Steps:

    • Interviews: Recruit 12 participants to collect detailed descriptions of their experiences with voice assistant failures.
    • Failure Data Collection: Use Amazon Mechanical Turk to crowdsource 199 voice assistant failure cases.
    • Survey: Develop a questionnaire to evaluate the impact of different failure scenarios on user trust and willingness to use the assistant for future tasks.
    • Data Analysis: Analyze differences in trust dimensions using quantitative statistical methods and qualitative coding frameworks.

Research Findings

  • Specific Results:

    • Classification results identified 12 failure sources, including MISSED TRIGGER (no response), SPURIOUS TRIGGER (false activation), etc. The most frequent failure type was MISUNDERSTANDING, while the least frequent was TRUNCATION.
    • Survey results indicated that certain failures (e.g., ambiguity-induced failures) had minimal negative impact on user trust, whereas failures like MISSED TRIGGER and CORRECT ACTION significantly undermined trust in the assistant's competence and benevolence.
  • Comparison with Existing Solutions:

    • This study is the first to systematically compare the multidimensional impact of various failure types on trust, deepening the understanding of user behavior patterns.
    • Unlike studies focusing solely on technical dimensions, this research provides critical insights into user perceptions and behaviors.
  • Experimental or Evaluation Results:

    • Quantitative analysis revealed that users were more likely to regain trust in voice assistants after failures in simple tasks, such as playing music.
    • The predictive effect of trust dimensions on retrying complex tasks (e.g., money transfers): Competence > Benevolence > Integrity.
  • Limitations and Future Directions:

    • Limitations:
      • Data collection relied on user recollections, which may involve recall bias.
      • The survey used simulated failure scenarios, which might not fully reflect reactions in real-world contexts.
      • The sample primarily consisted of frequent voice assistant users, which may not represent infrequent users or those more sensitive to failures.
    • Future Directions:
      • Use contextual experiments or log studies to capture real-time failure scenarios.
      • Expand the study to include a broader user demographic.
      • Investigate trust repair mechanisms and dynamic processes in greater depth.

Output Format

  • The research findings are clearly structured and align with practical needs for improving user experience and technology design.
  • This study provides direct guidance for developing strategies and directions for future voice assistant trust systems, offering robust data support for further exploration of user-technology interaction behaviors.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96119/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581152
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
Honorable Mention
group
Authors
6 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Voice Accessibility, Human-LLM Collaboration
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
2 related papers