When Help Hurts: Verification Load and Fatigue with AI Coding Assistants

Honorable Mention
Human-LLM CollaborationAI-Assisted Decision-Making & AutomationExplainable AI (XAI)AutoML InterfacesSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Paper Title

When Help Hurts: Verification Load and Fatigue with AI Coding Assistants

Publication Info

  • Topic area: Human-computer interaction in AI-assisted programming.
  • Keywords: AI coding assistants, verification load, cognitive load, interaction design, programming interfaces, trust calibration, fatigue, usability, adaptive orchestration, software development.

Background and Problem

  • Problem / challenge: While AI coding assistants improve productivity, developers face significant verification burdens to check and correct AI-generated code. Current research often conflates interface design with model capability, lacks cross-mode verification metrics, and does not clarify where interaction styles excel or fail.
  • Significance: Understanding and mitigating verification burdens is crucial for improving developer productivity, reducing fatigue, and ensuring the correctness and security of AI-assisted code.
  • Motivation and related work: Prior studies highlight speed gains from AI assistants but report mixed correctness and high validation costs. Research on cognitive load and trust calibration shows that poor interface design can exacerbate verification burdens, yet evaluations often fail to isolate interface effects from backend model capabilities.

Solution

  • Proposed approach: A controlled study comparing three interaction modes—Inline Suggestions, Chat-based Interaction, and Structured Prompts—while holding the LLM backend constant. Introduces a mode-agnostic verification-load index to quantify process costs.
  • Novelty:
    1. Interface attribution under backend parity: Isolates interface effects by using a fixed LLM backend across modes.
    2. Verification-load composite: Develops a behavioral index capturing compile/test failures, time-to-first-compile, code churn, pauses, and context switches.
    3. Complexity thresholds and expertise guidance: Identifies task complexity thresholds where modes excel and provides expertise-specific recommendations.
    4. Design and evaluation guidance: Proposes adaptive mode orchestration, transparency on demand, and verification-aware packaging.
  • Procedure and key techniques:
    • Conducted a mixed-methods study with 60 participants (36 AI-assisted, 24 no-AI control).
    • Tasks involved realistic Python programming challenges with controlled complexity.
    • Measured workload, completion time, correctness, stress, and fatigue, alongside behavioral logs.
    • Applied linear mixed-effects models, Johnson-Neyman analysis for complexity thresholds, and mediation analysis for verification-load effects.

Results

  • Concrete findings:
    • AI assistance reduced workload by 18.2 TLX points, improved correctness (OR = 1.71), and shortened task time by 22% compared to no-AI.
    • Inline was fastest and lowest-load for simple tasks; Chat improved correctness at high complexity without time penalties; Structured benefited novices at mid complexity.
    • Verification-load increased fatigue and stress across tasks, partially mediating these effects (26% for fatigue, 24% for stress).
  • Advantage over baselines:
    • AI outperformed the no-AI control in workload, time, and correctness.
    • Chat surpassed Inline in correctness for complex tasks (beyond complexity z ≈ +0.41).
    • Structured scaffolds reduced ambiguity for novices, improving correctness at mid complexity.
  • Experiments / evaluation:
    • Counterbalanced within-subjects design for AI modes; between-subjects design for AI vs. no-AI.
    • Tasks included parsing, debugging, and REST client implementation with controlled oracles.
    • Metrics included TLX workload, task time, correctness (unit tests passed), and a verification-load composite.
  • Limitations and future work:
    • Focused on Python tasks with a single-session exposure; generalization to other languages, frameworks, or long-term use is untested.
    • Excluded repository-aware or tool-augmented backends to isolate interface effects.
    • Future work should explore longitudinal adoption, richer integrations, and team-level dynamics.

Summary

This study demonstrates that interface design significantly impacts the trade-off between productivity and verification burden in AI-assisted programming. Inline Suggestions are optimal for simple tasks, Chat excels at high complexity, and Structured Prompts support novices at mid complexity. A mode-agnostic verification-load index quantifies the process costs of checking AI output and partially explains rising fatigue and stress under repeated use. These findings highlight the importance of adaptive mode orchestration, transparency on demand, and verification-aware packaging to optimize AI coding assistants.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223408/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791176
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Explainable AI (XAI), AutoML Interfaces
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers