LLM-box vs. Thinking-box: Designing for Deliberate User Engagement with Distorted Information in Conversational Search

Human-LLM CollaborationExplainable AI (XAI)Conversational Search & QA SystemsAI/ML Researchers & EngineersSoftware Engineers & DevelopersHCI Researchers

Paper Title

LLM-box vs. Thinking-box: Designing for Deliberate User Engagement with Distorted Information in Conversational Search

Publication Info

  • Topic area: User-level interventions to mitigate distorted information in conversational search powered by Large Language Models (LLMs).
  • Keywords: Conversational search, Large Language Models, hallucination, distorted information, user engagement, reflective nudges, deliberate thinking, information seeking, trust calibration, cognitive load.

Background and Problem

  • Problem / challenge: LLM-generated responses in conversational search often contain distorted information, which is difficult for users to recognize due to the fluent and persuasive style of outputs. Technical mitigations have reduced hallucination rates but cannot fully eliminate distortions.
  • Significance: Distorted information in conversational search impacts decision-making, learning, and exploration, necessitating user-level support to foster critical engagement and mitigate risks.
  • Motivation and related work: Prior research has explored technical solutions like retrieval augmentation and confidence scores, as well as reflective nudges to encourage critical thinking. However, these approaches often focus on detection accuracy or trust calibration, neglecting how users perceive and interact with such features in real-world scenarios.

Solution

  • Proposed approach: Two user-level interventions—LLM-box (confidence scores and descriptions) and Thinking-box (reflection-oriented checkpoints)—to support deliberate engagement with LLM responses.
  • Novelty:
    1. Comparative evaluation of technical and user-level approaches to mitigate distorted information.
    2. Integration of hallucination taxonomies into reflective prompts for user-level guidance.
    3. Empirical evidence on how interventions influence user engagement, cognitive load, and detection accuracy.
  • Procedure and key techniques:
    • Developed three systems: baseline (standard LLM interface), LLM-box (confidence scores and descriptions), and Thinking-box (sentence-level reflective prompts).
    • Conducted a within-subjects study with 16 participants using pre-generated LLM responses containing annotated distortions.
    • Measured detection accuracy, task completion time, and user perceptions through surveys, think-aloud sessions, and interviews.

Results

  • Concrete findings:
    • Thinking-box achieved the highest detection accuracy (true positives ↑) but increased cognitive load (completion time ↑).
    • LLM-box reduced cognitive effort but led to more false positives and narrowed user agency.
    • Both probes improved awareness of distortions compared to the baseline.
  • Advantage over baselines:
    • Thinking-box reduced missed distortions significantly (false negatives ↓) and fostered deeper engagement.
    • LLM-box raised distortion awareness but limited exploration beyond flagged content.
  • Experiments / evaluation:
    • Nine tasks across three systems, randomized and counterbalanced.
    • Quantitative metrics: detection accuracy (TP, FP, FN), task completion time.
    • Qualitative insights: user strategies, trust calibration, and preferences for guidance.
  • Limitations and future work:
    • Limited demographic diversity (young adults, LLM-literate participants).
    • Pre-generated content lacks real-time conversational dynamics.
    • Future research should explore broader user groups, discourse-level guidance, and longitudinal effects of reflective prompts.

Summary

This study compared two interventions—LLM-box and Thinking-box—to support users in engaging more critically with distorted information in conversational search. Thinking-box improved detection accuracy and fostered deliberate engagement but increased cognitive load, while LLM-box reduced effort but narrowed user agency. Both approaches enhanced distortion awareness compared to a baseline interface. Findings suggest that user-level interventions can complement technical solutions by scaffolding reflection without undermining autonomy. Design implications include framing guidance as suggestive, layering information with prioritization, and situating prompts as unobtrusive interruptions. Future systems should balance convenience with constructive effort to sustain critical engagement over time.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222113/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790271
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), Conversational Search & QA Systems
work
Professions
AI/ML Researchers & Engineers, Software Engineers & Developers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers