Assessing the Impact of Automated Suggestions on Decision Making: Domain Experts Mediate Model Errors but Take Less Initiative
Authors
Title of the Paper
Assessing the Impact of Automated Suggestions on Decision Making: Domain Experts Mediate Model Errors but Take Less Initiative
Paper Information
- Subject Areas: Artificial Intelligence and Human Collaboration, Medical Informatics, Human-Computer Interaction
- Keywords: Clinical Text Annotation, Ontology, Cognitive Models, AI Teams, Label Recommendation, User Agency, Automation Trust
Research Background and Issues
-
Identified Problems or Challenges:
- In hybrid intelligence systems, human reliance on automation errors may lead to incorrect decisions, reduced accuracy, and even loss of task engagement.
- Particularly in clinical annotation tasks, it remains unclear how domain experts interact with automated tools, such as whether trust in the system might cause them to overlook critical issues that require human judgment.
-
Significance: Clinical text annotation demands high expertise and is costly, while automated tools have the potential to reduce workload and improve efficiency. However, this may risk accepting system errors and diminishing users' creativity or critical engagement, ultimately impacting the quality of subsequent machine learning model training.
-
Research Motivation and Related Work:
- Research Motivation: Using the complex task of clinical text annotation as an example, this study investigates whether domain experts can maintain critical engagement and correct machine errors under automated assistance.
- Related Work: Previous studies indicate that users may overtrust automation, but most focus on simple tasks or general user groups, with limited research on domain experts and complex annotation tasks.
Solution
-
Proposed Methods or Solutions:
- Developed a novel text annotation platform incorporating existing label recommendation and automated pre-annotation functionalities.
- The system features a multi-panel interface, including a text panel, label recommendation panel, and highlight panel, supporting label recommendations and automated annotations generated by machine learning models.
- Designed a two-phase user experiment to analyze the impact of automated suggestion accuracy on user behavior, efficiency, and accuracy.
-
Innovations:
- Utilized a large medical label space containing over 400,000 labels, allowing the study to reflect real-world model error patterns.
- Conducted long-term tracking of user behavior (up to 8 hours), revealing how automation influences experts' trust and initiative.
-
Implementation Steps and Key Techniques:
- Phase 1: Observing the impact of label recommendations under standard and weakened modes on user behavior and decision-making.
- Phase 2: Investigating the effect of pre-automated annotations on user initiative and accuracy.
- Data Recording: Changes in user-platform interactions recorded in JSON format.
- Data Evaluation: Comparing recall rates, label accuracy, and user behavior in manually creating new annotations.
Research Findings
-
Specific Findings:
- Dependence on Label Recommendations:
- Experts were able to perceive deficiencies in the recommendation model and take retrieval measures when automated results were incomplete (recall rate reached 80%).
- Users were prone to accept suboptimal suggestions when the model's recommendation logic was confusing or conflicted with expert intuition.
- Performance in Pre-filled Annotations:
- Users exhibited high trust in erroneous annotations, with limited ability to modify incorrect annotations (accepting only 16%-33% of erroneous suggestions).
- After using pre-annotations, user initiative significantly decreased, with the frequency of creating new annotations dropping by 12%.
- Dependence on Label Recommendations:
-
Advantages Over Existing Solutions:
- The platform introduces a large-scale label space and closely integrates user behavior analysis, extending the scope of existing small-scale experimental studies.
-
Experimental or Evaluation Results:
- Label recommendations improved efficiency (annotation time reduced to an average of 3 seconds), and most correct recommendations were easily accepted.
- While pre-filled annotations reduced user workload, users were more likely to lose initiative and failed to recognize changes in their behavioral patterns.
-
Limitations and Future Directions:
- Limitations: The experiments were conducted in offline, low-risk tasks, insufficient to validate user behavior in high-pressure environments.
- Improvement Directions:
- Introduce real-time error feedback mechanisms (e.g., prompts after users accept erroneous labels).
- Further optimize the user interface to highlight the importance of uncertain labels.
- Explore personalized automation assistance based on user behavior.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do domain experts balance correcting system errors with maintaining their own agency when using automated label recommendation systems?Category: Uncertainty in Active Learning and Machine TeachingSimilar questionsarrow_forward
- How does accuracy of automated label recommendation and pre-filled annotations affect user behavior, efficiency, and decision quality?Category: Uncertainty in Active Learning and Machine TeachingSimilar questionsarrow_forward
- Are existing UI designs suitable for surfacing uncertainty in automated recommendations to improve user engagement effectiveness?Category: Uncertainty in Active Learning and Machine TeachingSimilar questionsarrow_forward
Practical Problems
1- Medical domain experts struggle to both improve efficiency and avoid accepting erroneous recommendations in automated annotation.Category: Uncertainty in Active Learning and Machine TeachingSimilar questionsarrow_forward
- 80%
EXMOS: Explanatory Model Steering through Multifaceted Explanations and Data Configurations
CHI '24· Explainable AI (XAI) +1
- 80%
Human-Algorithmic Interaction Using a Large Language Model-Augmented Artificial Intelligence Clinical Decision Support System
CHI '24· Human-LLM Collaboration +2
- 75%
Ignore, Trust, or Negotiate: Understanding Clinician Acceptance of AI-Based Treatment Recommendations in Health Care
CHI '23· Explainable AI (XAI) +1
- 67%
Accurate Insights, Trustworthy Interactions: Designing a Collaborative AI-Human Multi-Agent System with Knowledge Graph for Diagnosis Prediction
CHI '25· Brain-Computer Interface (BCI) & Neurofeedback +2
- 60%
Designing Theory-Driven User-Centric Explainable AI
CHI '19· Explainable AI (XAI) +1
- 60%
Designing AI for Trust and Collaboration in Time-Constrained Medical Decisions: A Sociotechnical Lens
CHI '21· Explainable AI (XAI) +1
- 60%
Healthcare AI Treatment Decision Support: Design Principles to Enhance Clinician Adoption and Trust
CHI '23· Explainable AI (XAI) +1
- 60%
“If I Had All the Time in the World”: Ophthalmologists' Perceptions of Anchoring Bias Mitigation in Clinical AI Support
CHI '23· Explainable AI (XAI) +1
- 60%
Harnessing Biomedical Literature to Calibrate Clinicians' Trust in AI Decision Support Systems
CHI '23· Explainable AI (XAI) +1
- 60%
Rethinking the Role of AI with Physicians in Oncology: Revealing Perspectives from Clinical and Research Workflows
CHI '23· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)