When Help Hurts: Verification Load and Fatigue with AI Coding Assistants
Honorable MentionAuthors
Paper Title
When Help Hurts: Verification Load and Fatigue with AI Coding Assistants
Publication Info
- Topic area: Human-computer interaction in AI-assisted programming.
- Keywords: AI coding assistants, verification load, cognitive load, interaction design, programming interfaces, trust calibration, fatigue, usability, adaptive orchestration, software development.
Background and Problem
- Problem / challenge: While AI coding assistants improve productivity, developers face significant verification burdens to check and correct AI-generated code. Current research often conflates interface design with model capability, lacks cross-mode verification metrics, and does not clarify where interaction styles excel or fail.
- Significance: Understanding and mitigating verification burdens is crucial for improving developer productivity, reducing fatigue, and ensuring the correctness and security of AI-assisted code.
- Motivation and related work: Prior studies highlight speed gains from AI assistants but report mixed correctness and high validation costs. Research on cognitive load and trust calibration shows that poor interface design can exacerbate verification burdens, yet evaluations often fail to isolate interface effects from backend model capabilities.
Solution
- Proposed approach: A controlled study comparing three interaction modes—Inline Suggestions, Chat-based Interaction, and Structured Prompts—while holding the LLM backend constant. Introduces a mode-agnostic verification-load index to quantify process costs.
- Novelty:
- Interface attribution under backend parity: Isolates interface effects by using a fixed LLM backend across modes.
- Verification-load composite: Develops a behavioral index capturing compile/test failures, time-to-first-compile, code churn, pauses, and context switches.
- Complexity thresholds and expertise guidance: Identifies task complexity thresholds where modes excel and provides expertise-specific recommendations.
- Design and evaluation guidance: Proposes adaptive mode orchestration, transparency on demand, and verification-aware packaging.
- Procedure and key techniques:
- Conducted a mixed-methods study with 60 participants (36 AI-assisted, 24 no-AI control).
- Tasks involved realistic Python programming challenges with controlled complexity.
- Measured workload, completion time, correctness, stress, and fatigue, alongside behavioral logs.
- Applied linear mixed-effects models, Johnson-Neyman analysis for complexity thresholds, and mediation analysis for verification-load effects.
Results
- Concrete findings:
- AI assistance reduced workload by 18.2 TLX points, improved correctness (OR = 1.71), and shortened task time by 22% compared to no-AI.
- Inline was fastest and lowest-load for simple tasks; Chat improved correctness at high complexity without time penalties; Structured benefited novices at mid complexity.
- Verification-load increased fatigue and stress across tasks, partially mediating these effects (26% for fatigue, 24% for stress).
- Advantage over baselines:
- AI outperformed the no-AI control in workload, time, and correctness.
- Chat surpassed Inline in correctness for complex tasks (beyond complexity z ≈ +0.41).
- Structured scaffolds reduced ambiguity for novices, improving correctness at mid complexity.
- Experiments / evaluation:
- Counterbalanced within-subjects design for AI modes; between-subjects design for AI vs. no-AI.
- Tasks included parsing, debugging, and REST client implementation with controlled oracles.
- Metrics included TLX workload, task time, correctness (unit tests passed), and a verification-load composite.
- Limitations and future work:
- Focused on Python tasks with a single-session exposure; generalization to other languages, frameworks, or long-term use is untested.
- Excluded repository-aware or tool-augmented backends to isolate interface effects.
- Future work should explore longitudinal adoption, richer integrations, and team-level dynamics.
Summary
This study demonstrates that interface design significantly impacts the trade-off between productivity and verification burden in AI-assisted programming. Inline Suggestions are optimal for simple tasks, Chat excels at high complexity, and Structured Prompts support novices at mid complexity. A mode-agnostic verification-load index quantifies the process costs of checking AI output and partially explains rising fatigue and stress under repeated use. These findings highlight the importance of adaptive mode orchestration, transparency on demand, and verification-aware packaging to optimize AI coding assistants.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 86%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 86%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 86%
Never-ending Learning of User Interfaces
UIST '23· Human-LLM Collaboration +2
- 75%
DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
CHI '26· Human-LLM Collaboration +3
- 71%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 71%
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
CHI '23· Human-LLM Collaboration +1
- 71%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 71%
Automatic Macro Mining from Interaction Traces at Scale
CHI '24· Human-LLM Collaboration +1
- 71%
Interactive Debugging and Steering of Multi-Agent AI Systems
CHI '25· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)