A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures
Honorable MentionAuthors
Intelligent Voice Assistants (Alexa, Siri, etc.)Voice AccessibilityHuman-LLM Collaboration
Document Title
A Mixed-Methods Approach to Understanding User Trust after Voice Assistant Failures
Document Information
- Subject Area: Human-Computer Interaction, Voice Assistant Trust Research
- Keywords: Voice Assistant, Trust, Survey, Interview Study, Failure Cases, User Behavior, Natural Language Processing, Dataset, Reliability, User Experience
Research Background and Issues
-
Problems and Challenges:
- Despite significant advancements in Natural Language Processing (NLP), voice assistants often fail to meet user expectations, particularly in accurately responding to complex tasks.
- Users generally exhibit low levels of trust in voice assistants, which hinders their broader adoption.
-
Significance:
- User trust is a critical factor for technology adoption, and voice assistants are increasingly being designed for complex tasks such as healthcare, mental health advice, and high-stakes decision-making.
- Investigating the impact of voice assistant failures on user trust can help optimize technology design and enhance user experience.
-
Research Motivation and Related Work:
- NLP models are typically evaluated using standardized datasets, which differ significantly from real-world application scenarios (e.g., mixed speech, incomplete inputs). This results in poor adaptability of the technology in actual user contexts.
- Few studies have systematically examined user perceptions of voice assistant failures and the specific impact of these failures on trust.
Solution
-
Methods and Approach:
- A mixed-methods approach is proposed to examine the impact of voice assistant failures on user trust.
- Conduct interviews to gain in-depth insights into user experiences with voice assistant failures.
- Build a dataset of user-reported failure cases and perform classification analysis.
- Design and implement a survey to quantify the impact of different failure types on trust (across the dimensions of competence, benevolence, and integrity).
-
Innovations:
- Extended existing NLP failure classification models to create a taxonomy of 12 failure sources specific to voice assistants.
- Released a dataset containing 199 user-submitted failure cases, providing a foundation for future research.
- Combined quantitative assessments of trust with real-world failure experiences, offering systematic insights.
-
Implementation Steps:
- Interviews: Recruit 12 participants to collect detailed descriptions of their experiences with voice assistant failures.
- Failure Data Collection: Use Amazon Mechanical Turk to crowdsource 199 voice assistant failure cases.
- Survey: Develop a questionnaire to evaluate the impact of different failure scenarios on user trust and willingness to use the assistant for future tasks.
- Data Analysis: Analyze differences in trust dimensions using quantitative statistical methods and qualitative coding frameworks.
Research Findings
-
Specific Results:
- Classification results identified 12 failure sources, including MISSED TRIGGER (no response), SPURIOUS TRIGGER (false activation), etc. The most frequent failure type was MISUNDERSTANDING, while the least frequent was TRUNCATION.
- Survey results indicated that certain failures (e.g., ambiguity-induced failures) had minimal negative impact on user trust, whereas failures like MISSED TRIGGER and CORRECT ACTION significantly undermined trust in the assistant's competence and benevolence.
-
Comparison with Existing Solutions:
- This study is the first to systematically compare the multidimensional impact of various failure types on trust, deepening the understanding of user behavior patterns.
- Unlike studies focusing solely on technical dimensions, this research provides critical insights into user perceptions and behaviors.
-
Experimental or Evaluation Results:
- Quantitative analysis revealed that users were more likely to regain trust in voice assistants after failures in simple tasks, such as playing music.
- The predictive effect of trust dimensions on retrying complex tasks (e.g., money transfers): Competence > Benevolence > Integrity.
-
Limitations and Future Directions:
- Limitations:
- Data collection relied on user recollections, which may involve recall bias.
- The survey used simulated failure scenarios, which might not fully reflect reactions in real-world contexts.
- The sample primarily consisted of frequent voice assistant users, which may not represent infrequent users or those more sensitive to failures.
- Future Directions:
- Use contextual experiments or log studies to capture real-time failure scenarios.
- Expand the study to include a broader user demographic.
- Investigate trust repair mechanisms and dynamic processes in greater depth.
- Limitations:
Output Format
- The research findings are clearly structured and align with practical needs for improving user experience and technology design.
- This study provides direct guidance for developing strategies and directions for future voice assistant trust systems, offering robust data support for further exploration of user-technology interaction behaviors.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- Which failure types during complex task execution by voice assistants most affect user trust?Category: Dialogue Error Repair and Failure Recovery StrategiesSimilar questionsarrow_forward
- Does user trust in voice assistants vary across different dimensions of failure scenarios (e.g., competence, benevolence, integrity)?Category: Dialogue Error Repair and Failure Recovery StrategiesSimilar questionsarrow_forward
- After which failure types are users more likely to retry using voice assistants?Category: Dialogue Error Repair and Failure Recovery StrategiesSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Users lose trust due to frequent voice assistant failures and are unwilling to use complex task features.Category: Dialogue Error Repair and Failure Recovery StrategiesSimilar questionsarrow_forward
- 67%
From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech Recognition
CHI '23· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
- 67%
A Prompt Chaining Framework for Long-Term Recall in LLM-Powered Intelligent Assistant
IUI '25· Intelligent Voice Assistants (Alexa, Siri, etc.) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581152
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
Honorable Mention
group
Authors
6 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Voice Accessibility, Human-LLM Collaboration
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
2 related papers