To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models

Honorable Mention
Human-LLM CollaborationAI-Assisted Decision-Making & Automation

Research Background and Issues

  • Issues or Challenges:
    This study investigates user reliance behaviors on large language models (LLMs). Users may exhibit over-reliance due to the model's confident outputs or under-reliance due to distrust in the model. A core challenge lies in enabling users to trust and utilize LLM recommendations appropriately, achieving "reasonable reliance."

  • Significance:
    As LLMs are increasingly integrated into various tasks (e.g., writing, education, decision support), their errors can significantly impact human decision-making. Over-reliance may lead users to accept incorrect information, while under-reliance may result in missed opportunities to benefit from accurate LLM recommendations.

  • Research Motivation and Related Work:
    Previous research has primarily focused on reducing over-reliance, with less attention given to maintaining reasonable reliance while minimizing erroneous reliance. Unlike prior studies on human-computer collaboration, this research aims to systematically and empirically evaluate several reliance intervention measures.

Solution

  • Proposed Methods or Solutions:
    The authors designed and tested three LLM reliance intervention measures:

    1. Reliance Disclaimer: Adding static reminders for users to verify information.
    2. Uncertainty Highlighting: Visualizing and marking uncertain parts of the LLM output.
    3. Implicit Answer: Presenting answers in a hidden format to compel users to engage more deeply with the suggestions.

    These methods aim to calibrate user reliance by increasing cognitive intervention and reducing blind trust in model outputs.

  • Innovative Contributions:
    This study not only explores methods to reduce over-reliance but also evaluates whether they promote reasonable reliance. It examines the effects on human-computer collaboration using multidimensional metrics such as "changes in user confidence" and "time engagement."

  • Implementation Steps and Key Techniques:

    1. Conducted a randomized controlled experiment with 400 participants.
    2. Provided two task types: LSAT logical reasoning (multiple-choice) and image quantity estimation (numerical problems).
    3. Designed the experiment in two stages: participants first completed tasks independently, then reviewed modified LLM output suggestions and re-answered.
    4. Collected data metrics including reliance rationality (correct/incorrect reliance), confidence calibration (confidence changes), and time engagement.

Research Findings

  • Specific Results:

    1. Reliance Disclaimer effectively improved reasonable reliance in LSAT tasks with minimal impact on time engagement.
    2. Uncertainty Highlighting reduced over-reliance but increased user hesitation toward model suggestions, leading to declines in subjective perceptions (e.g., accuracy and usefulness).
    3. Implicit Answer encouraged users to make more independent decisions but increased task completion time and reduced reliance due to the additional cognitive effort required.
  • Comparison with Existing Solutions and Advantages:
    Reliance Disclaimer proved to be the simplest and most effective method, with minimal added complexity to user interactions. Other interventions, while reducing over-reliance, introduced greater time costs or diminished user trust in LLMs.

  • Experimental or Evaluation Results:

    • In LSAT tasks, Reliance Disclaimer significantly improved the Appropriate Reliance Ratio (ARR).
    • Uncertainty Highlighting had limited effectiveness and, in some cases, exacerbated erroneous judgments of suggestions.
    • Implicit Answer increased cognitive engagement in numerical estimation tasks but led to higher under-reliance.
    • Confidence calibration was poor, with users displaying higher confidence when over-relying or erroneously relying on the model.
  • Limitations and Future Directions:

    1. The study did not fully simulate user behaviors in long-term and real-world interactions with LLMs (e.g., dynamic tasks).
    2. Experiments were conducted exclusively with GPT-4o, so results may not generalize to all LLMs.
    3. Intervention measures may require further refinement, such as alternative methods for highlighting uncertainty (e.g., marking high-confidence parts in green).
    4. Future research should test interventions in more open-ended and generative tasks (e.g., writing or code generation) and incorporate qualitative studies to explore users' actual perceptions and interaction patterns with model suggestions.

Conclusion

This study provides an important exploration of user reliance behaviors on LLMs, systematically evaluating how different interventions influence reasonable reliance and confidence calibration. The results indicate that simple interventions (e.g., reliance disclaimers) perform better in certain tasks, while more complex interventions may yield unintended effects. Future research should expand the diversity of study tasks and model types while further improving user behavior modeling and intervention design. This research offers valuable insights for building safer and more efficient human-computer collaboration systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188670/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714097
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
9 related papers