Characterizing Unintended Consequences of GUI Agents For Web Browsing
Authors
Paper Title
Characterizing Unintended Consequences of GUI Agents For Web Browsing
Publication Info
- Topic area: User-reported issues and unintended consequences of GUI agents in web browsing automation.
- Keywords: GUI agents, web automation, unintended consequences, user experience, security risks, mitigation strategies, LLMs, human-agent interaction, operational failures, design implications.
Background and Problem
- Problem / challenge: GUI agents for web browsing, powered by LLMs, face significant challenges in grounding abstract user intent, adapting to dynamic interfaces, and avoiding erroneous actions. These issues lead to operational failures, security risks, and user dissatisfaction, which are underexplored in the literature.
- Significance: Addressing these challenges is critical as GUI agents are increasingly integrated into workflows for tasks like e-commerce, social media management, and information retrieval. Failures in these systems can result in financial losses, security breaches, and erosion of trust.
- Motivation and related work: While prior research has focused on technical benchmarks and AI-induced harms in other domains (e.g., IoT vulnerabilities), there is a lack of user-centric studies on GUI agents. This paper aims to fill this gap by systematically analyzing user complaints and their consequences.
Solution
- Proposed approach: A two-phase mixed-methods study combining social media analysis (N=221 posts) and semi-structured interviews (N=21) to characterize user-reported complaints, their influences, and mitigation strategies for GUI agents in web browsing.
- Novelty:
- Development of a user-centric taxonomy of GUI agent complaints, focusing on failures in sense-making, adaptation, and operational burdens.
- Analysis of the multi-level impacts of these failures, including operational, security, and societal consequences.
- Synthesis of user-initiated mitigation strategies and design implications for safer, more reliable GUI agents.
- Procedure and key techniques:
- Phase 1: Social media analysis to identify naturally occurring complaints about GUI agents.
- Phase 2: Semi-structured interviews to contextualize and deepen understanding of user experiences.
- Thematic analysis to categorize complaints and derive insights into their consequences and mitigation strategies.
Results
- Concrete findings:
- Identified three categories of failures: (1) sense-making and intent alignment (e.g., task decomposition errors, instruction misinterpretation), (2) adaptation and execution (e.g., faulty GUI actions, poor UI adaptability), and (3) frictions and burdens (e.g., high operational costs, slow system response).
- Highlighted multi-level influences, including task abandonment, security vulnerabilities (e.g., uncontrolled file system access), and societal concerns like erosion of trust.
- Documented user mitigation strategies such as ecological sandboxing, cursor shadowing, and manual oversight.
- Advantage over baselines: Unlike prior studies that focus on technical benchmarks, this work provides a user-centric perspective, emphasizing real-world interaction failures and their broader implications.
- Experiments / evaluation:
- Social media analysis: Reviewed 1,850 posts, filtered to 221 relevant posts with 25,366 likes and 9,370 comments.
- Interviews: Conducted with 21 participants across diverse tasks (e.g., e-commerce, social media management) to validate and expand on findings.
- Limitations and future work:
- Scope limited to web browsing tasks; findings may not generalize to other domains.
- Participant pool skewed toward tech-savvy users in China; future work should include less technical users and other cultural contexts.
- Need for comparative studies across different agent platforms and task complexities.
Summary
This paper systematically investigates the unintended consequences of GUI agents in web browsing, identifying failures in sense-making, adaptation, and operational burdens. These failures lead to significant operational, security, and societal impacts, including task abandonment, security breaches, and erosion of trust. Users currently rely on ad-hoc mitigation strategies like ecological sandboxing and cursor shadowing. The study proposes design implications for consequence-aware agents, emphasizing interaction-aligned intervention, risk-aware operations, and ecological sandboxes. These findings offer actionable insights for improving GUI agent reliability and safety in real-world applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Sensemaking in Multi-Agent LLM Interfaces: How Users Interpret Transparency and Trustworthiness Cues
CHI '26· Human-LLM Collaboration +2
- 83%
The AI Memory Gap: Users Misremember What They Created With AI or Without
CHI '26· Human-LLM Collaboration +2
- 80%
UICrit: Enhancing Automated Design Evaluation with a UI Critique Dataset
UIST '24· Human-LLM Collaboration +1
- 71%
Privacy Control in Conversational LLM Platforms: A Walkthrough Study
CHI '26· Explainable AI (XAI) +3
- 71%
Characterizing User-Reported Risks across LLM Chatbots
CHI '26· Human-LLM Collaboration +3
- 71%
AI and My Values: User Perceptions of LLMs’ Ability to Extract, Embody, and Explain Human Values from Casual Conversations
CHI '26· Human-LLM Collaboration +3
- 67%
Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User Experience
CHI '23· Human-LLM Collaboration +2
- 67%
ONYX: Assisting Users in Teaching Natural Language Interfaces Through Multi-Modal Interactive Task Learning
CHI '23· Voice User Interface (VUI) Design +2
- 67%
ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
CHI '24· Human-LLM Collaboration +2
- 67%
Conversation Progress Guide : UI System for Enhancing Self-Efficacy in Conversational AI
CHI '25· Conversational Chatbots +2
Based on Jaccard similarity of research subtopics & professions (≥60%)