LLM-box vs. Thinking-box: Designing for Deliberate User Engagement with Distorted Information in Conversational Search
Authors
Paper Title
LLM-box vs. Thinking-box: Designing for Deliberate User Engagement with Distorted Information in Conversational Search
Publication Info
- Topic area: User-level interventions to mitigate distorted information in conversational search powered by Large Language Models (LLMs).
- Keywords: Conversational search, Large Language Models, hallucination, distorted information, user engagement, reflective nudges, deliberate thinking, information seeking, trust calibration, cognitive load.
Background and Problem
- Problem / challenge: LLM-generated responses in conversational search often contain distorted information, which is difficult for users to recognize due to the fluent and persuasive style of outputs. Technical mitigations have reduced hallucination rates but cannot fully eliminate distortions.
- Significance: Distorted information in conversational search impacts decision-making, learning, and exploration, necessitating user-level support to foster critical engagement and mitigate risks.
- Motivation and related work: Prior research has explored technical solutions like retrieval augmentation and confidence scores, as well as reflective nudges to encourage critical thinking. However, these approaches often focus on detection accuracy or trust calibration, neglecting how users perceive and interact with such features in real-world scenarios.
Solution
- Proposed approach: Two user-level interventions—LLM-box (confidence scores and descriptions) and Thinking-box (reflection-oriented checkpoints)—to support deliberate engagement with LLM responses.
- Novelty:
- Comparative evaluation of technical and user-level approaches to mitigate distorted information.
- Integration of hallucination taxonomies into reflective prompts for user-level guidance.
- Empirical evidence on how interventions influence user engagement, cognitive load, and detection accuracy.
- Procedure and key techniques:
- Developed three systems: baseline (standard LLM interface), LLM-box (confidence scores and descriptions), and Thinking-box (sentence-level reflective prompts).
- Conducted a within-subjects study with 16 participants using pre-generated LLM responses containing annotated distortions.
- Measured detection accuracy, task completion time, and user perceptions through surveys, think-aloud sessions, and interviews.
Results
- Concrete findings:
- Thinking-box achieved the highest detection accuracy (true positives ↑) but increased cognitive load (completion time ↑).
- LLM-box reduced cognitive effort but led to more false positives and narrowed user agency.
- Both probes improved awareness of distortions compared to the baseline.
- Advantage over baselines:
- Thinking-box reduced missed distortions significantly (false negatives ↓) and fostered deeper engagement.
- LLM-box raised distortion awareness but limited exploration beyond flagged content.
- Experiments / evaluation:
- Nine tasks across three systems, randomized and counterbalanced.
- Quantitative metrics: detection accuracy (TP, FP, FN), task completion time.
- Qualitative insights: user strategies, trust calibration, and preferences for guidance.
- Limitations and future work:
- Limited demographic diversity (young adults, LLM-literate participants).
- Pre-generated content lacks real-time conversational dynamics.
- Future research should explore broader user groups, discourse-level guidance, and longitudinal effects of reflective prompts.
Summary
This study compared two interventions—LLM-box and Thinking-box—to support users in engaging more critically with distorted information in conversational search. Thinking-box improved detection accuracy and fostered deliberate engagement but increased cognitive load, while LLM-box reduced effort but narrowed user agency. Both approaches enhanced distortion awareness compared to a baseline interface. Findings suggest that user-level interventions can complement technical solutions by scaffolding reflection without undermining autonomy. Design implications include framing guidance as suggestive, layering information with prioritization, and situating prompts as unobtrusive interruptions. Future systems should balance convenience with constructive effort to sustain critical engagement over time.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language Models
CHI '24· Human-LLM Collaboration +2
- 71%
DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking
CHI '26· Human-LLM Collaboration +2
- 71%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 71%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 71%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 71%
Exploring the Innovation Opportunities for Pre-trained Models
DIS '25· Generative AI (Text, Image, Music, Video) +2
- 67%
Validating AI-Generated Code with Live Programming
CHI '24· Human-LLM Collaboration +1
- 67%
Exploring the Design Space of Real-time LLM Knowledge Support Systems: A Case Study of Jargon Explanations
CHI '25· Human-LLM Collaboration +1
- 67%
User Modelling for Avoiding Overfitting in Interactive Knowledge Elicitation for Prediction
IUI '18· Human-LLM Collaboration +1
- 67%
Less or More: Towards Glanceable Explanations for LLM Recommendations Using Ultra-Small Devices
IUI '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)