Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
Authors
Paper Title
Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
Publication Info
- Topic area: Investigating user interaction behaviors that correlate with overreliance on conversational LLMs.
- Keywords: Overreliance, conversational LLMs, user behavior, interaction patterns, misinformation, cognitive processes, adaptive mitigation, real-time detection, behavioral clustering, human-AI collaboration.
Background and Problem
- Problem / challenge: Overreliance on conversational LLMs occurs when users accept incorrect AI recommendations, often due to the fluent and authoritative nature of LLM outputs. Current methods for detecting overreliance focus on task outcomes, which fail to capture the interaction process and intermediate errors.
- Significance: Overreliance can lead to serious errors, undermining the effectiveness of human-AI collaboration in critical tasks like decision-making, content creation, and problem-solving.
- Motivation and related work: Prior research has explored overreliance using outcome-oriented metrics and isolated behavioral signals but lacks a systematic framework to analyze interaction behaviors. This paper addresses this gap by linking user behaviors to overreliance and enabling real-time detection and adaptive mitigation.
Solution
- Proposed approach: A cluster-based analytical framework that identifies behavioral patterns correlated with overreliance on conversational LLMs by analyzing user interaction logs.
- Novelty:
- Creation of a dataset linking user interaction behaviors with overreliance metrics.
- Development of a clustering framework to quantify the relationship between behaviors and overreliance.
- Identification of five distinct behavioral patterns associated with overreliance, with cognitive interpretations and design implications.
- Procedure and key techniques:
- Conducted a controlled experiment with 77 participants completing three tasks (quiz solving, article summarization, trip planning) using an LLM injected with misinformation.
- Collected detailed interaction logs (e.g., mouse movements, clicks, keypresses) and processed them into feature vectors.
- Used a transformer-based autoencoder to embed interaction sequences into low-dimensional representations.
- Applied DBSCAN clustering to identify recurring behavioral patterns and validated clusters based on predictive capability and consistency.
Results
- Concrete findings:
- Five behavioral patterns were identified:
- High-frequency copying-pasting: Users with high overreliance frequently copied and pasted LLM outputs without editing.
- Focused task comprehension: Low-overreliance users spent more time understanding the task before engaging with the LLM.
- Frequency of referring to LLM responses: High-overreliance users repeatedly consulted the LLM, while low-overreliance users worked more independently.
- Coarse- vs. fine-grained locating and editing: High-overreliance users navigated and edited content roughly, while low-overreliance users made precise edits.
- Pausing and hesitation before prompting: High-overreliance users hesitated before crafting follow-up prompts but ultimately adopted LLM suggestions.
- Five behavioral patterns were identified:
- Advantage over baselines: The process-oriented approach captures intermediate behaviors and cognitive strategies, which outcome-based methods overlook.
- Experiments / evaluation:
- Tasks included quiz solving, article summarization, and trip planning, with injected misinformation to simulate real-world errors.
- Behavioral data were analyzed using clustering and validated through participant self-reports and cognitive interpretations.
- Limitations and future work:
- Artificially injected misinformation may not fully replicate natural LLM hallucinations.
- Findings are limited to the specific tasks and low-stakes scenarios studied.
- Future work should explore high-stakes tasks, real-time detection methods, and incorporate physiological measures for cognitive validation.
Summary
This study investigates overreliance on conversational LLMs by analyzing user interaction behaviors. Through a controlled experiment with 77 participants, five behavioral patterns were identified, highlighting differences in task comprehension, editing precision, and reliance on LLM outputs. The proposed clustering framework enables real-time detection of overreliance and informs adaptive mitigation strategies. While the study focuses on specific tasks, its findings provide a foundation for improving human-AI collaboration in diverse applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Plurals: A System for Guiding LLMs via Simulated Social Ensembles
CHI '25· Human-LLM Collaboration +2
- 83%
Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges
CHI '23· Human-LLM Collaboration +2
- 83%
Vulnerability of LLM Outputs to Heuristics-Inducing Prompt Structures
IUI '26· Human-LLM Collaboration +2
- 80%
User Modelling for Avoiding Overfitting in Interactive Knowledge Elicitation for Prediction
IUI '18· Human-LLM Collaboration +1
- 80%
DxHF: Providing High-Quality Human Feedback for LLM Alignment with Interactive Decomposition
UIST '25· Human-LLM Collaboration +1
- 71%
Who Controls the Conversation? User Perspectives On Generative AI (LLM) System Prompts
CHI '26· Human-LLM Collaboration +3
- 71%
All Accept, No Reject: Evaluating LLMs as “Peer” Reviewers
CHI '26· Human-LLM Collaboration +3
- 71%
Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational Agents
CHI '26· Human-LLM Collaboration +3
- 71%
Characterizing User-Reported Risks across LLM Chatbots
CHI '26· Human-LLM Collaboration +3
- 71%
AI and My Values: User Perceptions of LLMs’ Ability to Extract, Embody, and Explain Human Values from Casual Conversations
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)