To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models
Honorable MentionAuthors
Research Background and Issues
-
Issues or Challenges:
This study investigates user reliance behaviors on large language models (LLMs). Users may exhibit over-reliance due to the model's confident outputs or under-reliance due to distrust in the model. A core challenge lies in enabling users to trust and utilize LLM recommendations appropriately, achieving "reasonable reliance." -
Significance:
As LLMs are increasingly integrated into various tasks (e.g., writing, education, decision support), their errors can significantly impact human decision-making. Over-reliance may lead users to accept incorrect information, while under-reliance may result in missed opportunities to benefit from accurate LLM recommendations. -
Research Motivation and Related Work:
Previous research has primarily focused on reducing over-reliance, with less attention given to maintaining reasonable reliance while minimizing erroneous reliance. Unlike prior studies on human-computer collaboration, this research aims to systematically and empirically evaluate several reliance intervention measures.
Solution
-
Proposed Methods or Solutions:
The authors designed and tested three LLM reliance intervention measures:- Reliance Disclaimer: Adding static reminders for users to verify information.
- Uncertainty Highlighting: Visualizing and marking uncertain parts of the LLM output.
- Implicit Answer: Presenting answers in a hidden format to compel users to engage more deeply with the suggestions.
These methods aim to calibrate user reliance by increasing cognitive intervention and reducing blind trust in model outputs.
-
Innovative Contributions:
This study not only explores methods to reduce over-reliance but also evaluates whether they promote reasonable reliance. It examines the effects on human-computer collaboration using multidimensional metrics such as "changes in user confidence" and "time engagement." -
Implementation Steps and Key Techniques:
- Conducted a randomized controlled experiment with 400 participants.
- Provided two task types: LSAT logical reasoning (multiple-choice) and image quantity estimation (numerical problems).
- Designed the experiment in two stages: participants first completed tasks independently, then reviewed modified LLM output suggestions and re-answered.
- Collected data metrics including reliance rationality (correct/incorrect reliance), confidence calibration (confidence changes), and time engagement.
Research Findings
-
Specific Results:
- Reliance Disclaimer effectively improved reasonable reliance in LSAT tasks with minimal impact on time engagement.
- Uncertainty Highlighting reduced over-reliance but increased user hesitation toward model suggestions, leading to declines in subjective perceptions (e.g., accuracy and usefulness).
- Implicit Answer encouraged users to make more independent decisions but increased task completion time and reduced reliance due to the additional cognitive effort required.
-
Comparison with Existing Solutions and Advantages:
Reliance Disclaimer proved to be the simplest and most effective method, with minimal added complexity to user interactions. Other interventions, while reducing over-reliance, introduced greater time costs or diminished user trust in LLMs. -
Experimental or Evaluation Results:
- In LSAT tasks, Reliance Disclaimer significantly improved the Appropriate Reliance Ratio (ARR).
- Uncertainty Highlighting had limited effectiveness and, in some cases, exacerbated erroneous judgments of suggestions.
- Implicit Answer increased cognitive engagement in numerical estimation tasks but led to higher under-reliance.
- Confidence calibration was poor, with users displaying higher confidence when over-relying or erroneously relying on the model.
-
Limitations and Future Directions:
- The study did not fully simulate user behaviors in long-term and real-world interactions with LLMs (e.g., dynamic tasks).
- Experiments were conducted exclusively with GPT-4o, so results may not generalize to all LLMs.
- Intervention measures may require further refinement, such as alternative methods for highlighting uncertainty (e.g., marking high-confidence parts in green).
- Future research should test interventions in more open-ended and generative tasks (e.g., writing or code generation) and incorporate qualitative studies to explore users' actual perceptions and interaction patterns with model suggestions.
Conclusion
This study provides an important exploration of user reliance behaviors on LLMs, systematically evaluating how different interventions influence reasonable reliance and confidence calibration. The results indicate that simple interventions (e.g., reliance disclaimers) perform better in certain tasks, while more complex interventions may yield unintended effects. Future research should expand the diversity of study tasks and model types while further improving user behavior modeling and intervention design. This research offers valuable insights for building safer and more efficient human-computer collaboration systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can interventions be designed to reduce users' over-reliance or under-reliance on large language models (LLM)?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How do different interventions (e.g., reliance disclaimers, uncertainty labels, and implicit answers) affect user trust and appropriate reliance behavior?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- What factors influence users' trust calibration toward LLM advice across different task types?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
Practical Problems
1- Users' inability to appropriately trust LLMs leads to wrong decisions or missed opportunities.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 100%
Rethinking Interaction: From Instrumental Interaction to Human-Computer Partnerships
CHI '18· Human-LLM Collaboration +1
- 100%
Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
CHI '25· Human-LLM Collaboration +1
- 100%
Enhancing AI-Assisted Group Decision Making through LLM-Powered Devil's Advocate
IUI '24· Human-LLM Collaboration +1
- 100%
C-PAK: Correcting and Completing Variable-length Prefix-based Abbreviated Keystrokes
UIST '23· Human-LLM Collaboration +1
- 67%
Automating Clinical Documentation with Digital Scribes: Understanding the Impact on Physicians
CHI '21· Human-LLM Collaboration +1
- 67%
Unified Conversational Models with System-Initiated Transitions between Chit-Chat and Task-Oriented Dialogues
CUI '23· Conversational Chatbots +2
- 67%
TiiS: A Review of User Interface Design for Interactive Machine Learning
IUI '19· Human-LLM Collaboration +2
- 67%
Induction of an active attitude by short speech reaction time toward interaction for decision-making with multiple agents
IUI '19· Agent Personality & Anthropomorphism +2
- 67%
Exploring the Effects of Machine Learning Literacy Interventions on Laypeople's Reliance on Machine Learning Models
IUI '22· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)