Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
Honorable MentionAuthors
Human-LLM CollaborationExplainable AI (XAI)Software Engineers & DevelopersHCI ResearchersCognitive Scientists
Research Background and Issues
- Issues and Challenges: The output of large language models (LLMs) may contain convincing yet factually incorrect information, leading to problems of user over-reliance. Such misinformation can prompt users to act based on incorrect data, which can be particularly dangerous in high-risk scenarios.
- Significance: As LLMs rapidly gain popularity in fields such as search and information retrieval, the issue of over-reliance is considered a critical research topic in human-AI interaction. Addressing this issue directly impacts whether users can correctly utilize this technology.
- Research Motivation and Related Work: While much of the prior research on AI reliance has focused on traditional AI models, studies on LLMs are still in their infancy. Existing work indicates that explanations and other features can influence user trust in systems, but it also reveals that these explanations may foster unreasonable reliance. This study aims to provide insights into mitigating over-reliance by examining system-generated explanations, source citations, and inconsistencies.
Solution
- Proposed Approach: The authors conducted two studies: 1) Using the "think aloud" method to understand how users perceive ChatGPT responses and identify the features influencing their reliance decisions; 2) Designing a large-scale, pre-registered experiment with tightly controlled variables to analyze key factors affecting user reliance.
- Innovations: The research not only investigates the impact of explanations, sources, and inconsistencies but also analyzes the interactions between these features. Additionally, the study uses real LLM outputs and creates simulated environments, making the experimental results more applicable to real-world scenarios.
- Implementation Steps:
- Study Design: In the initial exploratory study, 16 participants engaged in multi-turn interactions with ChatGPT to answer questions, while their observations on explanations, inconsistencies, and sources were recorded.
- Experimental Manipulation: In the second experiment, three variables (correct/incorrect answers, presence of explanations, presence of sources) were strictly controlled to study their impact on user accuracy and reliance behaviors.
- Technical Methods: ChatGPT and Perplexity AI (a model focused on providing sources) were used to generate the required dialogues for the experiment, and the logic of explanations was encoded. Source-click behaviors were meticulously tracked during the experiment.
Research Findings
-
Key Results:
- Providing explanations significantly increased user reliance but also amplified over-reliance on incorrect responses. Providing sources effectively reduced over-reliance and improved appropriate reliance on correct answers. Inconsistent explanations reduced user reliance on incorrect answers.
- Users expressed lower trust and engagement in the absence of explanations or sources, indicating the critical role of these features in building trust.
-
Advantages Over Existing Solutions:
- The study introduced emphasizing inconsistencies in LLM outputs as an effective strategy to mitigate over-reliance, offering a new direction beyond the existing focus on "explanation visualization."
- It highlighted the importance and impact of source quality, proposing design improvements based on source credibility and usability.
-
Experimental and Evaluation Results:
- The first study revealed that users value explanations and sources and consider logical inconsistencies as important indicators of reliability.
- In the second experiment, user answer accuracy was influenced by different variable combinations. Providing sources significantly promoted accurate decision-making, while the absence of explanations or inconsistent explanations reduced reliance on incorrect answers.
- The experiment demonstrated that users tend to rely on explanations as a measure of credibility; however, lengthy or incorrect explanations may "confuse" users, leading to over-reliance.
-
Limitations and Future Directions:
- The experiment was limited to binary question-answering tasks and cannot be directly generalized to other LLM application scenarios (e.g., writing or dialogue generation).
- The error rate of the LLM was set relatively low in the experimental design, and the study did not explore how higher accuracy rates might affect reliance.
- Due to the simplified interaction process used to control research variables, the study did not simulate dynamic changes in multi-turn interaction scenarios.
Suggested Future Research Directions:
- Expand research on the impact of inconsistencies on user behavior in other task contexts (e.g., creative or complex reasoning tasks).
- Investigate how to optimize the presentation of explanations while reducing the negative effects of over-reliance on inaccurate outputs.
- Explore ways to integrate user background knowledge and model visualization techniques to enhance the efficiency of LLM usage.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How do users determine trust in large language models (LLMs) based on features such as explanations, citations, and consistency?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How do system explanations, cited sources, and output consistency affect reliance on incorrect versus correct information?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- Can highlighting inconsistency in output quality effectively reduce users' over-reliance on incorrect LLM outputs?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Users easily over-trust incorrect information from LLMs, leading to decision errors.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 80%
EvAlignUX: Advancing UX Evaluation through LLM-Supported Metrics Exploration
CHI '25· Human-LLM Collaboration +1
- 67%
"Why is 'Chicago' deceptive?" Towards Building Model-Driven Tutorials for Humans
CHI '20· Human-LLM Collaboration +2
- 67%
Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking
IUI '24· Human-LLM Collaboration +1
- 67%
ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts
UIST '25· Human-LLM Collaboration +2
- 60%
COGAM: Measuring and Moderating Cognitive Load in Machine Learning Model Explanations
CHI '20· Explainable AI (XAI)
- 60%
Continual Human-in-the-Loop Optimization
CHI '25· Human-LLM Collaboration
- 60%
MAPLE: Mobile App Prediction Leveraging Large Language Model Embeddings
UbiComp '24· Human-LLM Collaboration
- 60%
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
UIST '24· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714020
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI)
work
Professions
Software Engineers & Developers, HCI Researchers, Cognitive Scientists
article
Content Status
Full text indexed
hub
Related Papers
8 related papers