"Here, Let Me Help": An Empirical Study of User Interventions in Human–Web Agent Collaboration

Honorable Mention
Human-LLM CollaborationAI-Assisted Decision-Making & AutomationUser Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingSoftware Engineers & DevelopersUI/UX DesignersAI/ML Researchers & Engineers

Paper Title

'Here, Let Me Help': An Empirical Study of User Interventions in Human–Web Agent Collaboration

Publication Info

  • Topic area: User interventions in human–web agent collaboration.
  • Keywords: Web agents, human–agent collaboration, user intervention, mixed-initiative systems, usability, task completion, user behavior, large language models, interface automation, empirical study.

Background and Problem

  • Problem / challenge: Current web agents struggle with accuracy, latency, and task complexity, necessitating human intervention. Existing research focuses on outcome-based metrics, neglecting the process-level dynamics of human–agent collaboration.
  • Significance: Understanding user interventions is critical for improving web agent usability, reliability, and collaboration, especially as agents remain imperfect and prone to errors.
  • Motivation and related work: Prior work has explored web agents, benchmarks, and mixed-initiative systems but lacks a systematic understanding of why and how users intervene during agent execution. This study addresses this gap by analyzing intervention behaviors and proposing a taxonomy.

Solution

  • Proposed approach: A systematic empirical study of user interventions during web agent execution, focusing on reasons, types, and patterns of intervention.
  • Novelty:
    1. A taxonomy of intervention reasons and types, distinguishing explicit and implicit interventions.
    2. Analysis of usability perceptions and their relationship with intervention behaviors.
    3. Actionable design implications for improving mixed-initiative systems.
    4. Insights into task and domain-specific factors influencing interventions.
  • Procedure and key techniques:
    • Conducted an in-lab study with 30 participants performing 12 tasks across shopping, travel, and information domains using live websites.
    • Collected interaction logs, user inputs, and screen recordings to analyze intervention behaviors.
    • Developed a taxonomy of intervention reasons (e.g., instruction–action misalignment, incomplete task termination) and types (e.g., prompt restructuring, direct substitution).
    • Quantitatively analyzed subjective usability ratings and behavioral metrics.

Results

  • Concrete findings:
    • 44.2% of tasks were completed without intervention, 29.7% with intervention, and 14.7% failed despite intervention.
    • UI misinterpretation (41%) and instruction misunderstandings (23%) were the most common reasons for intervention.
    • Explicit interventions (e.g., prompt restructuring, direct substitution) were more frequent than implicit ones (e.g., tracking, cross-checking).
    • Higher intervention frequency correlated with lower perceived effectiveness, satisfaction, and trust.
  • Advantage over baselines: Provides process-level insights into user interventions, which are absent in outcome-centric evaluations of web agents.
  • Experiments / evaluation:
    • Tasks were derived from benchmarks and structured around four activity types (navigating, finding, gathering, transacting).
    • Used a generalist web agent with GPT-4o for task execution.
    • Measured usability (effectiveness, ease, satisfaction, trust) and intervention behaviors.
  • Limitations and future work:
    • Limited to novice users and three domains (shopping, travel, information).
    • Does not capture long-term behavioral evolution or higher-stakes contexts.
    • Future work should explore broader domains, experienced users, and richer behavioral signals (e.g., eye-tracking).

Summary

This study investigates user interventions in human–web agent collaboration through a systematic in-lab experiment with 30 participants performing 12 tasks across three domains. It introduces a taxonomy of intervention reasons (e.g., instruction–action misalignment, incomplete task termination) and types (e.g., prompt restructuring, direct substitution), revealing systematic patterns and their impact on usability. Findings highlight the trade-off between intervention frequency and user satisfaction, emphasizing the need for collaborative, mixed-initiative systems. The study provides actionable design implications, such as improving prompt scaffolds, enhancing progress visibility, and addressing latency-driven interventions, to support seamless human–agent collaboration.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222579/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791536
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
Software Engineers & Developers, UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers