Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies

Honorable Mention
Human-LLM CollaborationExplainable AI (XAI)Software Engineers & DevelopersHCI ResearchersCognitive Scientists

Research Background and Issues

  • Issues and Challenges: The output of large language models (LLMs) may contain convincing yet factually incorrect information, leading to problems of user over-reliance. Such misinformation can prompt users to act based on incorrect data, which can be particularly dangerous in high-risk scenarios.
  • Significance: As LLMs rapidly gain popularity in fields such as search and information retrieval, the issue of over-reliance is considered a critical research topic in human-AI interaction. Addressing this issue directly impacts whether users can correctly utilize this technology.
  • Research Motivation and Related Work: While much of the prior research on AI reliance has focused on traditional AI models, studies on LLMs are still in their infancy. Existing work indicates that explanations and other features can influence user trust in systems, but it also reveals that these explanations may foster unreasonable reliance. This study aims to provide insights into mitigating over-reliance by examining system-generated explanations, source citations, and inconsistencies.

Solution

  • Proposed Approach: The authors conducted two studies: 1) Using the "think aloud" method to understand how users perceive ChatGPT responses and identify the features influencing their reliance decisions; 2) Designing a large-scale, pre-registered experiment with tightly controlled variables to analyze key factors affecting user reliance.
  • Innovations: The research not only investigates the impact of explanations, sources, and inconsistencies but also analyzes the interactions between these features. Additionally, the study uses real LLM outputs and creates simulated environments, making the experimental results more applicable to real-world scenarios.
  • Implementation Steps:
    1. Study Design: In the initial exploratory study, 16 participants engaged in multi-turn interactions with ChatGPT to answer questions, while their observations on explanations, inconsistencies, and sources were recorded.
    2. Experimental Manipulation: In the second experiment, three variables (correct/incorrect answers, presence of explanations, presence of sources) were strictly controlled to study their impact on user accuracy and reliance behaviors.
    3. Technical Methods: ChatGPT and Perplexity AI (a model focused on providing sources) were used to generate the required dialogues for the experiment, and the logic of explanations was encoded. Source-click behaviors were meticulously tracked during the experiment.

Research Findings

  • Key Results:

    1. Providing explanations significantly increased user reliance but also amplified over-reliance on incorrect responses. Providing sources effectively reduced over-reliance and improved appropriate reliance on correct answers. Inconsistent explanations reduced user reliance on incorrect answers.
    2. Users expressed lower trust and engagement in the absence of explanations or sources, indicating the critical role of these features in building trust.
  • Advantages Over Existing Solutions:

    1. The study introduced emphasizing inconsistencies in LLM outputs as an effective strategy to mitigate over-reliance, offering a new direction beyond the existing focus on "explanation visualization."
    2. It highlighted the importance and impact of source quality, proposing design improvements based on source credibility and usability.
  • Experimental and Evaluation Results:

    • The first study revealed that users value explanations and sources and consider logical inconsistencies as important indicators of reliability.
    • In the second experiment, user answer accuracy was influenced by different variable combinations. Providing sources significantly promoted accurate decision-making, while the absence of explanations or inconsistent explanations reduced reliance on incorrect answers.
    • The experiment demonstrated that users tend to rely on explanations as a measure of credibility; however, lengthy or incorrect explanations may "confuse" users, leading to over-reliance.
  • Limitations and Future Directions:

    1. The experiment was limited to binary question-answering tasks and cannot be directly generalized to other LLM application scenarios (e.g., writing or dialogue generation).
    2. The error rate of the LLM was set relatively low in the experimental design, and the study did not explore how higher accuracy rates might affect reliance.
    3. Due to the simplified interaction process used to control research variables, the study did not simulate dynamic changes in multi-turn interaction scenarios.

Suggested Future Research Directions:

  • Expand research on the impact of inconsistencies on user behavior in other task contexts (e.g., creative or complex reasoning tasks).
  • Investigate how to optimize the presentation of explanations while reducing the negative effects of over-reliance on inaccurate outputs.
  • Explore ways to integrate user background knowledge and model visualization techniques to enhance the efficiency of LLM usage.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188664/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714020
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI)
work
Professions
Software Engineers & Developers, HCI Researchers, Cognitive Scientists
article
Content Status
Full text indexed
hub
Related Papers
8 related papers