Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
Authors
Research Background and Issues
-
Issues and Challenges:
This paper investigates how search tools based on large language models (LLMs) influence decision-making processes, particularly in terms of speed, accuracy, and over-reliance. Traditional search tools limit user efficiency when handling complex queries, whereas LLM-based search tools significantly enhance user experience but may lead to erroneous decisions due to the phenomenon of "hallucination" (i.e., generating information without factual basis). -
Significance:
With the widespread application of LLM technology in online tools, these tools are transforming the search and decision-making behaviors of billions of users. As a critical gateway for information retrieval, the impact of search tools on decision quality needs to be thoroughly understood, especially in high-stakes scenarios such as purchasing a car. -
Research Motivation and Related Work:
The authors review the potential of LLMs to enhance productivity (e.g., writing, programming) while highlighting the risks of these tools generating false content and users' over-reliance on them. Additionally, the study explores design interventions to mitigate over-reliance, discussing how visual signals (e.g., color coding) can reduce this phenomenon.
Solution
-
Method or Solution:
The authors conducted two online experiments comparing the performance of traditional search tools and LLM-based search tools. In the second experiment, they tested a color-coding mechanism (indicating the confidence level of model-generated information) to reduce users' reliance on low-confidence outputs. -
Innovations:
- Investigated the direct impact of LLMs on users' efficiency in handling complex queries and decision quality.
- Proposed and validated a confidence-based visual feedback mechanism using simple color coding to alert users to potential errors in LLM-generated information.
- Highly controlled research variables, including configuring the LLM with zero "temperature" to ensure output consistency.
-
Implementation Steps and Key Techniques:
- Experiment 1: Compared the efficiency, query quantity and complexity, and decision accuracy between LLM-based search tools and traditional search tools.
- Participants completed a series of car comparison tasks (selecting based on the ratio of cargo space to total length).
- Recorded participants' query time, quantity, complexity, and final decisions.
- Experiment 2: Explored the impact of color-coded visual cues on users' ability to identify low-confidence information from LLMs.
- Divided participants into three groups: no cues, low-confidence cues only, and both high- and low-confidence cues.
- Used GPT-3 to generate confidence probabilities (>50% as high confidence, ≤50% as low confidence).
- Experiment 1: Compared the efficiency, query quantity and complexity, and decision accuracy between LLM-based search tools and traditional search tools.
Research Findings
-
Specific Findings:
- Using LLM-based search tools significantly improved task completion speed, with users completing tasks using fewer but more complex queries.
- Compared to traditional search tools, LLMs performed similarly in normal tasks but led to significantly higher error rates when generating incorrect information, demonstrating users' over-reliance on erroneous outputs.
- The color-coding mechanism effectively improved users' ability to identify incorrect information, increasing task accuracy by over 70%.
-
Advantages:
Compared to traditional search tools, LLM-based search tools significantly improved query efficiency and user experience. After implementing the cue mechanism, users demonstrated higher proactivity in detecting incorrect information, mitigating the "hallucination" effect of the tools. -
Experimental or Evaluation Results:
- Experiment 1: LLM-based search tools saved approximately 50% of task completion time compared to traditional search tools, with higher query complexity but no significant reduction in error rates. In error scenarios, traditional search tools maintained a 93% accuracy rate, whereas LLMs achieved only 47% (Task 5).
- Experiment 2: Color coding improved accuracy in error-prone scenarios, with the two cue groups achieving 58% and 53% accuracy, respectively, compared to 26% in the no-cue group.
- In terms of user satisfaction, while the reliability ratings of the two search tools showed little difference, users of LLM-based tools reported greater satisfaction with the overall search experience.
-
Limitations and Future Directions:
-
Limitations:
- The experiments were conducted solely in the car-purchasing scenario, excluding more complex or open-ended decision tasks.
- The tools developed did not utilize conversational LLM models (e.g., Bing Chat or Google Gemini).
- The study employed specific confidence cues (color coding) and did not compare other cue mechanisms (e.g., confidence scores or source links).
-
Future Directions:
- Conduct further research on open-ended tasks, such as interactive experiences involving complex search goals.
- Apply new LLM models (e.g., retrieval-augmented generation methods combining traditional search and LLMs).
- Explore richer cue designs, such as expressing uncertainty in natural language or dynamically generating confidence scores.
- Investigate whether different user groups have differentiated needs for confidence cues, such as tailoring cue designs based on users' education levels or technical proficiency.
-
Through this study, the authors demonstrate the immense potential of LLM-based search tools in the field of information retrieval while providing critical guidance for designing intelligent search tools with greater reliability and user adaptability.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do LLM search tools affect user efficiency, accuracy, and decision quality on complex queries?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- Can confidence visual feedback mechanisms such as color coding reduce over-reliance on low-confidence information?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- Under what conditions do LLM-generated errors significantly increase users' decision error rates?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
Practical Problems
1- Users relying on LLM tools may depend on incorrect information and make poor decisions.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 100%
How Users Perceive Mixed-Initiative AI: Attitudes Toward Assistance in Problem Solving
IUI '26· Human-LLM Collaboration +2
- 86%
Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 86%
An Exploration of Default Images in Text-to-Image Generation
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 86%
Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents
CHI '26· Human-LLM Collaboration +3
- 83%
Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User Experience
CHI '23· Human-LLM Collaboration +2
- 83%
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
CHI '25· Human-LLM Collaboration +2
- 83%
Co-Disclosing the Computer: LLM-Mediated Computing through Reflective Conversation
CHI '26· Human-LLM Collaboration +2
- 83%
What can AI do for me: Evaluating Machine Learning Interpretations in Cooperative Play
IUI '19· Human-LLM Collaboration +2
- 83%
CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent Agents
IUI '26· Human-LLM Collaboration +2
- 83%
User Reliance on AI Support for Collaborative Partner Selection
IUI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)