A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
Authors
Paper Title
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
Publication Info
- Topic area: Accessibility and human-computer interaction in computer use agents (CUAs).
- Keywords: Accessibility, computer use agents, blind and low-vision users, assistive technology, screen readers, magnifiers, interaction datasets, task success, human-computer interaction, accessibility benchmarking.
Background and Problem
- Problem / challenge: Current CUAs are designed to mimic sighted users' interactions, neglecting the needs and practices of blind and low-vision users (BLVUs). There is no existing dataset capturing real-world BLVU interactions to evaluate CUAs under assistive technology (AT) conditions.
- Significance: Addressing this gap is crucial for ensuring CUAs are accessible to all users, enabling equitable collaboration, error prevention, and privacy protection for BLVUs.
- Motivation and related work: Prior datasets and benchmarks focus on sighted users (SUs) and web-based interactions, leaving a gap in understanding CUAs' performance in AT-mediated environments. Existing studies on BLVUs often isolate specific interaction methods or tasks, without providing comprehensive datasets for benchmarking CUAs.
Solution
- Proposed approach: Introduction of the A11y-CUA dataset, which captures multimodal interaction traces of BLVUs and SUs performing 60 everyday tasks across desktop and web applications.
- Novelty:
- Creation of a dataset with 40.4 hours of interaction data and 158,325 events from 16 participants (8 SUs, 8 BLVUs).
- Development of a computer use recorder to capture synchronized, replayable traces of real-world interactions.
- Comparative analysis of interaction styles between SUs and BLVUs, highlighting between-group and within-group variability.
- Evaluation of state-of-the-art CUAs under default and AT conditions, revealing significant accessibility gaps.
- Procedure and key techniques:
- Data collection involved participants completing 60 tasks in five categories (Browsing & Web, System Operations, Document Editing, Workflow, Media) using a controlled Windows environment.
- Tasks were designed to reflect real-world scenarios, with verifiable end states and multimodal logging of interactions (e.g., keystrokes, mouse actions, accessibility settings).
- CUAs (Claude Sonnet 4.5 and Qwen3-VL-32B-Instruct) were evaluated under three conditions: default, screen-reader (keyboard-only), and magnifier (150% viewport scaling).
Results
- Concrete findings:
- SUs completed tasks with a 99.1% success rate in an average of 92.3 seconds, while BLVUs achieved 84.6% success in 211.1 seconds.
- Default-CUA achieved a 78.33% success rate, while SR-CUA and Magnifier-CUA succeeded in only 41.67% and 28.33% of tasks, respectively.
- CUAs under AT conditions exhibited slower task completion times and higher failure rates compared to human participants.
- Advantage over baselines:
- The dataset provides a unique benchmark for evaluating CUAs under AT constraints, highlighting perceptual, cognitive, and action gaps in current systems.
- It enables analysis of diverse interaction strategies within and across user groups, offering insights for improving CUA design.
- Experiments / evaluation:
- Tasks were evaluated for success rates, completion times, and interaction methods across participants and CUAs.
- Metrics included mouse and keyboard actions, hotkey usage, and navigation patterns.
- CUAs were tested under default, keyboard-only, and magnified viewport conditions, with manual review of task recordings for error analysis.
- Limitations and future work:
- The dataset focuses on Windows OS and closed-source applications, limiting generalizability.
- Future work could expand to other operating systems, open-source applications, and more complex or open-ended tasks.
- Enhancements to CUAs could include AT-synchronized perception feeds, improved task state tracking, and robust error recovery mechanisms.
Summary
The A11y-CUA dataset addresses the accessibility gap in CUAs by capturing real-world interaction traces of SUs and BLVUs across 60 tasks. Analysis reveals distinct interaction styles, with SUs favoring mouse-dominant strategies and BLVUs relying on keyboard navigation and screen-reader feedback. CUAs under AT conditions perform poorly, with significant perceptual, cognitive, and action gaps. This dataset provides a foundation for benchmarking and improving accessibility-aware CUAs, promoting equitable computer use for all users. Future directions include expanding the dataset's scope and enhancing CUA capabilities for better AT integration.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)