"Here, Let Me Help": An Empirical Study of User Interventions in Human–Web Agent Collaboration
Honorable MentionAuthors
Paper Title
'Here, Let Me Help': An Empirical Study of User Interventions in Human–Web Agent Collaboration
Publication Info
- Topic area: User interventions in human–web agent collaboration.
- Keywords: Web agents, human–agent collaboration, user intervention, mixed-initiative systems, usability, task completion, user behavior, large language models, interface automation, empirical study.
Background and Problem
- Problem / challenge: Current web agents struggle with accuracy, latency, and task complexity, necessitating human intervention. Existing research focuses on outcome-based metrics, neglecting the process-level dynamics of human–agent collaboration.
- Significance: Understanding user interventions is critical for improving web agent usability, reliability, and collaboration, especially as agents remain imperfect and prone to errors.
- Motivation and related work: Prior work has explored web agents, benchmarks, and mixed-initiative systems but lacks a systematic understanding of why and how users intervene during agent execution. This study addresses this gap by analyzing intervention behaviors and proposing a taxonomy.
Solution
- Proposed approach: A systematic empirical study of user interventions during web agent execution, focusing on reasons, types, and patterns of intervention.
- Novelty:
- A taxonomy of intervention reasons and types, distinguishing explicit and implicit interventions.
- Analysis of usability perceptions and their relationship with intervention behaviors.
- Actionable design implications for improving mixed-initiative systems.
- Insights into task and domain-specific factors influencing interventions.
- Procedure and key techniques:
- Conducted an in-lab study with 30 participants performing 12 tasks across shopping, travel, and information domains using live websites.
- Collected interaction logs, user inputs, and screen recordings to analyze intervention behaviors.
- Developed a taxonomy of intervention reasons (e.g., instruction–action misalignment, incomplete task termination) and types (e.g., prompt restructuring, direct substitution).
- Quantitatively analyzed subjective usability ratings and behavioral metrics.
Results
- Concrete findings:
- 44.2% of tasks were completed without intervention, 29.7% with intervention, and 14.7% failed despite intervention.
- UI misinterpretation (41%) and instruction misunderstandings (23%) were the most common reasons for intervention.
- Explicit interventions (e.g., prompt restructuring, direct substitution) were more frequent than implicit ones (e.g., tracking, cross-checking).
- Higher intervention frequency correlated with lower perceived effectiveness, satisfaction, and trust.
- Advantage over baselines: Provides process-level insights into user interventions, which are absent in outcome-centric evaluations of web agents.
- Experiments / evaluation:
- Tasks were derived from benchmarks and structured around four activity types (navigating, finding, gathering, transacting).
- Used a generalist web agent with GPT-4o for task execution.
- Measured usability (effectiveness, ease, satisfaction, trust) and intervention behaviors.
- Limitations and future work:
- Limited to novice users and three domains (shopping, travel, information).
- Does not capture long-term behavioral evolution or higher-stakes contexts.
- Future work should explore broader domains, experienced users, and richer behavioral signals (e.g., eye-tracking).
Summary
This study investigates user interventions in human–web agent collaboration through a systematic in-lab experiment with 30 participants performing 12 tasks across three domains. It introduces a taxonomy of intervention reasons (e.g., instruction–action misalignment, incomplete task termination) and types (e.g., prompt restructuring, direct substitution), revealing systematic patterns and their impact on usability. Findings highlight the trade-off between intervention frequency and user satisfaction, emphasizing the need for collaborative, mixed-initiative systems. The study provides actionable design implications, such as improving prompt scaffolds, enhancing progress visibility, and addressing latency-driven interventions, to support seamless human–agent collaboration.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 88%
Interview-Informed Generative Agents for Product Discovery: A Validation Study
CHI '26· Human-LLM Collaboration +3
- 88%
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
CHI '26· Human-LLM Collaboration +3
- 88%
ImpReSS: Designing and Evaluating a Lightweight Implicit Recommender System in Conversational Support Agents
IUI '26· Human-LLM Collaboration +4
- 86%
Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support
CHI '25· Human-LLM Collaboration +2
- 86%
PointAloud: An Interaction Suite for AI-Supported Pointer-Centric Think-Aloud Computing
CHI '26· Human-LLM Collaboration +2
- 75%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 75%
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
CHI '26· Human-LLM Collaboration +3
- 75%
The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
CHI '26· Human-LLM Collaboration +3
- 75%
Live in the Loop: Rapid Run-time Feedback for Prompts
CHI '26· Human-LLM Collaboration +3
- 71%
'It Is Not Always Discovery Time': Four Pragmatic Approaches in Designing AI Systems
CHI '22· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)