The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
Authors
Paper Title
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
Publication Info
- Topic area: Human-computer interaction focusing on latency effects in LLM-based knowledge work.
- Keywords: Response latency, human-LLM interaction, task type, perceived quality, interaction design, cognitive workload, AI trust, user behavior, positive friction, knowledge work.
Background and Problem
- Problem / challenge: While latency is a key factor in user experience with LLMs, its nuanced effects on user behavior and perceptions across different task types remain underexplored.
- Significance: Understanding latency's impact can inform the design of LLM systems to optimize user satisfaction, trust, and task efficiency, especially as LLMs become integral to knowledge work.
- Motivation and related work: Prior HCI research has treated latency as a cost to minimize, focusing on thresholds for usability. However, conversational and probabilistic dynamics of LLMs introduce new interpretive dimensions to waiting, such as anthropomorphizing delays as "thinking time." This study addresses gaps in understanding how latency interacts with task type to shape user behavior and perceptions.
Solution
- Proposed approach: A controlled 2 × 3 experiment examining the effects of latency (2, 9, 20 seconds) and task type (Creation, Advice) on human-LLM interaction and perception.
- Novelty:
- Empirical evidence showing that task type, more than latency, drives interaction behaviors, while latency influences perceptions of response quality.
- Design implications for using latency as a tunable interaction variable rather than a cost to minimize.
- Development of a reusable experimental platform for studying latency effects in LLM interactions.
- Procedure and key techniques:
- Participants (N=240) completed three tasks in either Creation or Advice categories, interacting with a GPT-4-based assistant under fixed latency conditions (2, 9, or 20 seconds).
- Behavioral data (e.g., prompt submissions, copy-pasting) and subjective ratings (e.g., clarity, usefulness) were collected.
- Statistical analyses (e.g., ANOVA) and thematic coding of qualitative feedback were used to evaluate latency and task type effects.
Results
- Concrete findings:
- Creation tasks elicited more prompt submissions (M=6.20) than Advice tasks (M=5.14), regardless of latency.
- Short latencies (2 s) were rated less thoughtful and useful than moderate (9 s) or long (20 s) latencies.
- Moderate latencies (9 s) yielded the highest usefulness ratings, while very long latencies (20 s) sometimes raised reliability concerns.
- Participants often interpreted delays as AI "deliberation," enhancing perceived thoughtfulness.
- Advantage over baselines:
- Demonstrated that latency effects are not linear; moderate delays can enhance perceived quality compared to very short waits.
- Highlighted task-specific interaction patterns, with Creation tasks driving more iterative prompting than Advice tasks.
- Experiments / evaluation:
- Design: 2 × 3 between-subjects experiment.
- Datasets: Creation and Advice tasks adapted from prior HCI/AI studies.
- Metrics: Behavioral logs (e.g., prompt counts), subjective ratings (e.g., clarity, usefulness), and NASA-TLX workload scores.
- Limitations and future work:
- Focused only on time-to-first-token latency; future studies should explore other timing factors (e.g., streaming pace, total response time).
- Conducted in a controlled setting; real-world, time-sensitive contexts may yield different results.
- Limited to U.S.-based participants with high digital familiarity; cultural and demographic diversity should be examined in future studies.
Summary
This study investigated how response latency and task type influence user interaction with LLMs. Interaction behaviors were robust to latency but varied significantly by task type, with Creation tasks eliciting more prompts than Advice tasks. Perceptions of response quality were shaped by latency, with moderate delays (9 s) enhancing perceived usefulness and thoughtfulness, while very short (2 s) and very long (20 s) delays were less favorable. These findings suggest that latency can be a design variable rather than a cost to minimize, with implications for task-specific tuning and ethical considerations in LLM-based systems. Future research should explore latency effects in more complex, real-world settings.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 100%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 86%
DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
CHI '26· Human-LLM Collaboration +3
- 86%
When Help Hurts: Verification Load and Fatigue with AI Coding Assistants
CHI '26· Human-LLM Collaboration +3
- 83%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 83%
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
CHI '23· Human-LLM Collaboration +1
- 83%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 83%
Automatic Macro Mining from Interaction Traces at Scale
CHI '24· Human-LLM Collaboration +1
- 83%
Interactive Debugging and Steering of Multi-Agent AI Systems
CHI '25· Human-LLM Collaboration +2
- 83%
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
CHI '25· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)