Assessing the Impact of Automated Suggestions on Decision Making: Domain Experts Mediate Model Errors but Take Less Initiative

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationPhysicians, Nurses & CliniciansUniversity Professors & Researchers

Title of the Paper

Assessing the Impact of Automated Suggestions on Decision Making: Domain Experts Mediate Model Errors but Take Less Initiative

Paper Information

  • Subject Areas: Artificial Intelligence and Human Collaboration, Medical Informatics, Human-Computer Interaction
  • Keywords: Clinical Text Annotation, Ontology, Cognitive Models, AI Teams, Label Recommendation, User Agency, Automation Trust

Research Background and Issues

  • Identified Problems or Challenges:

    1. In hybrid intelligence systems, human reliance on automation errors may lead to incorrect decisions, reduced accuracy, and even loss of task engagement.
    2. Particularly in clinical annotation tasks, it remains unclear how domain experts interact with automated tools, such as whether trust in the system might cause them to overlook critical issues that require human judgment.
  • Significance: Clinical text annotation demands high expertise and is costly, while automated tools have the potential to reduce workload and improve efficiency. However, this may risk accepting system errors and diminishing users' creativity or critical engagement, ultimately impacting the quality of subsequent machine learning model training.

  • Research Motivation and Related Work:

    • Research Motivation: Using the complex task of clinical text annotation as an example, this study investigates whether domain experts can maintain critical engagement and correct machine errors under automated assistance.
    • Related Work: Previous studies indicate that users may overtrust automation, but most focus on simple tasks or general user groups, with limited research on domain experts and complex annotation tasks.

Solution

  • Proposed Methods or Solutions:

    1. Developed a novel text annotation platform incorporating existing label recommendation and automated pre-annotation functionalities.
    2. The system features a multi-panel interface, including a text panel, label recommendation panel, and highlight panel, supporting label recommendations and automated annotations generated by machine learning models.
    3. Designed a two-phase user experiment to analyze the impact of automated suggestion accuracy on user behavior, efficiency, and accuracy.
  • Innovations:

    1. Utilized a large medical label space containing over 400,000 labels, allowing the study to reflect real-world model error patterns.
    2. Conducted long-term tracking of user behavior (up to 8 hours), revealing how automation influences experts' trust and initiative.
  • Implementation Steps and Key Techniques:

    1. Phase 1: Observing the impact of label recommendations under standard and weakened modes on user behavior and decision-making.
    2. Phase 2: Investigating the effect of pre-automated annotations on user initiative and accuracy.
    3. Data Recording: Changes in user-platform interactions recorded in JSON format.
    4. Data Evaluation: Comparing recall rates, label accuracy, and user behavior in manually creating new annotations.

Research Findings

  • Specific Findings:

    1. Dependence on Label Recommendations:
      • Experts were able to perceive deficiencies in the recommendation model and take retrieval measures when automated results were incomplete (recall rate reached 80%).
      • Users were prone to accept suboptimal suggestions when the model's recommendation logic was confusing or conflicted with expert intuition.
    2. Performance in Pre-filled Annotations:
      • Users exhibited high trust in erroneous annotations, with limited ability to modify incorrect annotations (accepting only 16%-33% of erroneous suggestions).
      • After using pre-annotations, user initiative significantly decreased, with the frequency of creating new annotations dropping by 12%.
  • Advantages Over Existing Solutions:

    • The platform introduces a large-scale label space and closely integrates user behavior analysis, extending the scope of existing small-scale experimental studies.
  • Experimental or Evaluation Results:

    • Label recommendations improved efficiency (annotation time reduced to an average of 3 seconds), and most correct recommendations were easily accepted.
    • While pre-filled annotations reduced user workload, users were more likely to lose initiative and failed to recognize changes in their behavioral patterns.
  • Limitations and Future Directions:

    1. Limitations: The experiments were conducted in offline, low-risk tasks, insufficient to validate user behavior in high-pressure environments.
    2. Improvement Directions:
      • Introduce real-time error feedback mechanisms (e.g., prompts after users accept erroneous labels).
      • Further optimize the user interface to highlight the importance of uncertain labels.
      • Explore personalized automation assistance based on user behavior.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47633/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445522
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
Physicians, Nurses & Clinicians, University Professors & Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers