Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance

Human-LLM CollaborationExplainable AI (XAI)AI-Assisted Decision-Making & AutomationUI/UX DesignersAI/ML Researchers & EngineersHCI Researchers

Research Background and Issues

  • Issues and Challenges:
    This paper investigates how search tools based on large language models (LLMs) influence decision-making processes, particularly in terms of speed, accuracy, and over-reliance. Traditional search tools limit user efficiency when handling complex queries, whereas LLM-based search tools significantly enhance user experience but may lead to erroneous decisions due to the phenomenon of "hallucination" (i.e., generating information without factual basis).

  • Significance:
    With the widespread application of LLM technology in online tools, these tools are transforming the search and decision-making behaviors of billions of users. As a critical gateway for information retrieval, the impact of search tools on decision quality needs to be thoroughly understood, especially in high-stakes scenarios such as purchasing a car.

  • Research Motivation and Related Work:
    The authors review the potential of LLMs to enhance productivity (e.g., writing, programming) while highlighting the risks of these tools generating false content and users' over-reliance on them. Additionally, the study explores design interventions to mitigate over-reliance, discussing how visual signals (e.g., color coding) can reduce this phenomenon.

Solution

  • Method or Solution:
    The authors conducted two online experiments comparing the performance of traditional search tools and LLM-based search tools. In the second experiment, they tested a color-coding mechanism (indicating the confidence level of model-generated information) to reduce users' reliance on low-confidence outputs.

  • Innovations:

    1. Investigated the direct impact of LLMs on users' efficiency in handling complex queries and decision quality.
    2. Proposed and validated a confidence-based visual feedback mechanism using simple color coding to alert users to potential errors in LLM-generated information.
    3. Highly controlled research variables, including configuring the LLM with zero "temperature" to ensure output consistency.
  • Implementation Steps and Key Techniques:

    1. Experiment 1: Compared the efficiency, query quantity and complexity, and decision accuracy between LLM-based search tools and traditional search tools.
      • Participants completed a series of car comparison tasks (selecting based on the ratio of cargo space to total length).
      • Recorded participants' query time, quantity, complexity, and final decisions.
    2. Experiment 2: Explored the impact of color-coded visual cues on users' ability to identify low-confidence information from LLMs.
      • Divided participants into three groups: no cues, low-confidence cues only, and both high- and low-confidence cues.
      • Used GPT-3 to generate confidence probabilities (>50% as high confidence, ≤50% as low confidence).

Research Findings

  • Specific Findings:

    1. Using LLM-based search tools significantly improved task completion speed, with users completing tasks using fewer but more complex queries.
    2. Compared to traditional search tools, LLMs performed similarly in normal tasks but led to significantly higher error rates when generating incorrect information, demonstrating users' over-reliance on erroneous outputs.
    3. The color-coding mechanism effectively improved users' ability to identify incorrect information, increasing task accuracy by over 70%.
  • Advantages:
    Compared to traditional search tools, LLM-based search tools significantly improved query efficiency and user experience. After implementing the cue mechanism, users demonstrated higher proactivity in detecting incorrect information, mitigating the "hallucination" effect of the tools.

  • Experimental or Evaluation Results:

    1. Experiment 1: LLM-based search tools saved approximately 50% of task completion time compared to traditional search tools, with higher query complexity but no significant reduction in error rates. In error scenarios, traditional search tools maintained a 93% accuracy rate, whereas LLMs achieved only 47% (Task 5).
    2. Experiment 2: Color coding improved accuracy in error-prone scenarios, with the two cue groups achieving 58% and 53% accuracy, respectively, compared to 26% in the no-cue group.
    3. In terms of user satisfaction, while the reliability ratings of the two search tools showed little difference, users of LLM-based tools reported greater satisfaction with the overall search experience.
  • Limitations and Future Directions:

    • Limitations:

      1. The experiments were conducted solely in the car-purchasing scenario, excluding more complex or open-ended decision tasks.
      2. The tools developed did not utilize conversational LLM models (e.g., Bing Chat or Google Gemini).
      3. The study employed specific confidence cues (color coding) and did not compare other cue mechanisms (e.g., confidence scores or source links).
    • Future Directions:

      1. Conduct further research on open-ended tasks, such as interactive experiences involving complex search goals.
      2. Apply new LLM models (e.g., retrieval-augmented generation methods combining traditional search and LLMs).
      3. Explore richer cue designs, such as expressing uncertainty in natural language or dynamically generating confidence scores.
      4. Investigate whether different user groups have differentiated needs for confidence cues, such as tailoring cue designs based on users' education levels or technical proficiency.

Through this study, the authors demonstrate the immense potential of LLM-based search tools in the field of information retrieval while providing critical guidance for designing intelligent search tools with greater reliability and user adaptability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188754/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714082
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers