Do People Appropriately Rely on AI-Advice? An Analytical Review of HCI Research on Human-AI Decision-Making
Honorable MentionAuthors
Muhammad Raees
Rochester Institute of TechnologyVassilis-Javed Khan
independentPaper Title
Do People Appropriately Rely on AI-Advice? An Analytical Review of HCI Research on Human-AI Decision-Making
Publication Info
- Topic area: Human-AI interaction and decision-making
- Keywords: Human-AI reliance, appropriate reliance, over-reliance, under-reliance, decision-making, XAI, cognitive biases, user engagement, interaction design, HCI
Background and Problem
- Problem / challenge: Existing research on human-AI reliance often lacks realistic decision-making scenarios, leading to gaps in understanding how users appropriately rely on AI advice. There is also no consensus on definitions and metrics for "appropriate reliance."
- Significance: Understanding and fostering appropriate reliance on AI advice is crucial for improving decision-making in high-stakes domains such as healthcare, finance, and justice, where accountability and ethical concerns are paramount.
- Motivation and related work: Prior studies have explored trust in automation, explainable AI (XAI), and user biases, but they often fail to address how users discriminate between correct and incorrect AI advice or adapt their decisions accordingly. This paper builds on existing work by focusing on recent studies and objective metrics for human-AI reliance.
Solution
- Proposed approach: Analytical review of 56 studies (2018–2025) on human-AI reliance, focusing on users, AI systems, and interaction methods to consolidate findings and provide recommendations for achieving appropriate reliance.
- Novelty:
- Comprehensive analysis of recent empirical studies on human-AI reliance, covering diverse domains and methods.
- Discussion and advocacy for a unified definition of "appropriate reliance."
- Identification of gaps and recommendations for improving reliance through user engagement, task design, and interaction strategies.
- Procedure and key techniques:
- Systematic review using SCOPUS and ACM Digital Library with PRISMA framework.
- Categorization of studies by reliance metrics, experimental design, participant expertise, and AI system fidelity.
- Analysis of interventions and their impact on reliance, including cognitive forcing, XAI methods, and user engagement strategies.
Results
- Concrete findings:
- 71% of studies use multi-step decision-making protocols, allowing users to revise decisions after seeing AI advice.
- Common metrics include decision accuracy (35 studies), agreement fraction (21 studies), and switch fraction (15 studies).
- Most studies involve novice participants (48/56), often recruited via crowdsourcing platforms.
- 33% of studies use simulated AI systems, raising concerns about the realism of findings.
- Advantage over baselines:
- Identifies tailored metrics like Relative AI Reliance (RAIR) and Relative Self-Reliance (RSR) to better capture appropriate reliance.
- Highlights the effectiveness of cognitive forcing and frictional designs in reducing over-reliance.
- Experiments / evaluation:
- Studies span domains like healthcare, business, education, and leisure, with tasks ranging from medical diagnostics to sentiment analysis.
- Most studies rely on crowd workers, with limited involvement of domain experts or realistic tasks.
- Limitations and future work:
- Over-reliance on novice participants and simulated AI systems limits generalizability.
- Lack of standardized metrics and definitions for appropriate reliance.
- Future work should focus on realistic tasks, domain-specific expertise, and layered interaction designs.
Summary
This paper reviews 56 studies on human-AI reliance, highlighting gaps in realistic decision-making scenarios and the lack of consensus on appropriate reliance metrics. It identifies key factors influencing reliance, including user biases, task complexity, and interaction design. The findings emphasize the importance of multi-step decision protocols, tailored metrics like RAIR and RSR, and interventions such as cognitive forcing to improve reliance. However, the reliance on novice participants and simulated AI systems limits the applicability of current research. Future work should prioritize realistic tasks, domain expertise, and reflective interaction designs to better understand and foster appropriate reliance on AI advice.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Farsight: Fostering Responsible AI Awareness During AI Application Prototyping
CHI '24· Explainable AI (XAI) +2
- 83%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 83%
Trust in AI-assisted Decision Making: Perspectives from Those Behind the System and Those for Whom the Decision is Made
CHI '24· Explainable AI (XAI) +2
- 83%
"AI enhances our performance, I have no doubt this one will do the same": The Placebo effect is robust to negative descriptions of AI
CHI '24· Explainable AI (XAI) +2
- 83%
Trusting Autonomous Teammates in Human-AI Teams - A Literature Review
CHI '25· Explainable AI (XAI) +2
- 71%
Evaluating the Interpretability of Generative Models by Interactive Reconstruction
CHI '21· Explainable AI (XAI) +2
- 71%
Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI Systems
CHI '23· Explainable AI (XAI) +2
- 71%
PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule Synthesis
CHI '23· Explainable AI (XAI) +2
- 71%
Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
CHI '23· Explainable AI (XAI) +2
- 71%
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
CHI '24· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)