AXNav: Replaying Accessibility Tests from Natural Language
Authors
Document Title
AXNav: Replaying Accessibility Tests from Natural Language
Document Information
- Topic Area: Automation of accessibility testing based on natural language
- Keywords: Accessibility testing, UI automation, large language models, natural language processing, testing tools, human-computer interaction, test automation, multimodal planning, screen reader, dynamic font
Research Background and Problem
-
What problems or challenges did the authors identify?
- Most applications provide inadequate support for accessibility features (e.g., screen readers, dynamic fonts) due to a lack of experience, awareness, or organizational support.
- Although many testing and inspection tools exist, the industry still relies heavily on time-consuming and difficult-to-scale manual testing.
- Automated testing tools (e.g., GUI testing and accessibility checkers) are fragile, unable to adapt effectively to UI changes, and their scanning results may contain numerous false positives.
- It is challenging to achieve complete coverage of all scenarios and features through manual testing, especially in development environments requiring frequent updates.
-
Why is this problem important?
- Accessibility is crucial for ensuring applications are suitable and equitable for individuals with disabilities.
- Insufficient automated testing not only increases labor costs but may also lead to undetected issues, ultimately harming user experience.
- Efficient accessibility testing tools can significantly improve development quality and help achieve broader user coverage.
-
Research Motivation and Related Work
- Although existing research has attempted to use large language models (LLMs) for automating tasks such as UI interaction and debugging report reproduction, no studies have focused on leveraging LLMs for accessibility testing tasks.
- Current manual testing tools often rely on static step records, which are prone to failure during UI updates. This research aims to improve the situation through dynamic responses driven by natural language and models.
- Practitioners have expressed concerns about the high repetition and update costs of large-scale manual testing, underscoring the need for a method to reduce manual steps while retaining the flexibility of human inspection.
Solution
-
What methods or solutions did the authors propose?
- Developed a new system named AXNav, which understands manual testing instructions from natural language descriptions and reproduces tests in real applications.
- AXNav enables and configures various accessibility features (e.g., VoiceOver, dynamic fonts, bold fonts) and generates actionable step plans through model inference for execution.
- After testing, it produces annotated videos with chapters to help testers quickly locate potential accessibility issues.
-
What are the innovative aspects of this solution?
- LLM-driven multi-agent architecture: AXNav employs a multi-agent architecture (Planner, Action, and Evaluator) to dynamically generate and revise plans, allowing for adjustments in unexpected scenarios (e.g., UI changes, permission requests).
- Pixel-level UI recognition combined with natural language interface: Utilizes pixel-level UI detection models to provide screen structure and interaction suggestions to the LLM, enabling adaptation to unknown applications.
- Multi-level accessibility performance analysis: Developed a set of custom heuristic algorithms to detect and annotate specific accessibility issues (e.g., navigation loops, inconsistent font scaling).
-
What are the implementation steps and key technologies used?
- Device allocation and control: Remote cloud-based iOS devices are used to install target applications and configure accessibility features.
- Test planning and execution:
- Plan generation: Generates actionable step JSON structures based on natural language instructions.
- UI element interaction: Combines pixel-level UI recognition with VoiceOver gesture commands for operation.
- Dynamic re-planning: Adjustments are triggered if errors occur or plans are infeasible.
- Heuristic detection: Detects issues such as:
- Inconsistent font scaling for Dynamic Type.
- Click element identification problems in Button Shapes.
- Navigation loops or missing elements in VoiceOver.
- Output results: Automatically generates interactive videos with annotated key issue chapters to improve review efficiency.
Research Outcomes
-
What specific results were achieved?
- Replay accuracy on real test sets and self-constructed open datasets reached 85.5% (complex task accuracy 61.1%) and 70% (complex task accuracy 64.3%), respectively.
- The system was evaluated as "very useful" by 10 professional testers, especially in standardizing and efficiently executing multiple testing steps.
- The system's automatically generated interactive videos helped testers quickly locate potential issues and served as supplementary materials for debugging reports.
-
What advantages does it have compared to existing solutions?
- Higher flexibility and dynamic adjustment capability: Does not rely on hardcoded interaction paths and adapts to frequent UI changes.
- Semantic natural language interface: Users can initiate complex tasks through concise descriptions without needing in-depth knowledge of automation frameworks.
- Higher coverage: Integrates support for four different accessibility features (VoiceOver, Dynamic Type, etc.) and analyzes specific issues based on heuristic rules.
-
What were the experimental or evaluation results?
- Technical evaluation: AXNav performed reasonably well in navigating complex tasks but still has room for improvement, such as more effectively handling collection item selection.
- User evaluation: Video chapters and LLM-guided logical outputs were deemed significant supplements to manual testing workflows.
- User preference: Participants believed that in future integrations, the tool could significantly reduce manual labor through parallel testing.
-
Limitations and Future Directions:
- Limitations:
- Currently supports only the iOS platform; future expansion to Android and other operating systems is needed.
- UI navigation plans may fail in specific scenarios, such as tasks requiring scrolling or multi-layer nesting.
- Video content is not user-friendly for screen reader users.
- Future Directions:
- Enhance LLM navigation capabilities, such as integrating specific contextual modeling.
- Support additional accessibility features (e.g., color inversion, voice control) and broader heuristic checks.
- Provide interaction displays suitable for non-visual users and improve transparency (e.g., error cause explanations).
- Introduce more efficient UI summary mechanisms, such as dashboard-style problem summaries.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can large language models (LLMs) convert natural language descriptions into executable accessibility testing plans?Category: ML/AI Model Visualization and Explainable AnalysisSimilar questionsarrow_forward
- How does the AXNav system dynamically adjust testing plans to improve flexibility when UIs change frequently?Category: ML/AI Model Visualization and Explainable AnalysisSimilar questionsarrow_forward
- Can visualized outputs (such as annotated interactive videos) effectively improve the efficiency and accuracy of accessibility testing?Category: ML/AI Model Visualization and Explainable AnalysisSimilar questionsarrow_forward
Practical Problems
1- Developers struggle to efficiently test and optimize accessibility support in applications.Category: ML/AI Model Visualization and Explainable AnalysisSimilar questionsarrow_forward
- 80%
Persona-L has Entered the Chat: Leveraging LLMs and Ability-based Framework for Personas of People with Complex Needs
CHI '25· Voice Accessibility +2
- 67%
IdeaBot: Investigating Social Facilitation in Human-Machine Team Creativity
CHI '21· Conversational Chatbots +3
- 60%
Mapping Machine Learning Advances from HCI Research to Reveal Starting Places for Design Innovation
CHI '18· Human-LLM Collaboration
- 60%
Typing Efficiency and Suggestion Accuracy Influence the Benefits and Adoption of Word Suggestions
CHI '21· Human-LLM Collaboration +1
- 60%
OPTIMISM: Enabling Collaborative Implementation of Domain-Specific Metaheuristic Optimization
CHI '23· Generative AI (Text, Image, Music, Video) +1
- 60%
Generating Automatic Feedback on UI Mockups with Large Language Models
CHI '24· Human-LLM Collaboration +1
- 60%
Canvil: Designerly Adaptation for LLM-Powered User Experiences
CHI '25· 360° Video & Panoramic Content +1
- 60%
No Evidence for LLMs Being Useful in Problem Reframing
CHI '25· Human-LLM Collaboration +1
- 60%
PlanTogether: Facilitating AI Application Planning Using Information Graphs and Large Language Models
CHI '25· Human-LLM Collaboration +1
- 60%
The Bot on Speaking Terms: The Effects of Conversation Architecture on Perceptions of Conversational Agents
CUI '23· Agent Personality & Anthropomorphism +1
Based on Jaccard similarity of research subtopics & professions (≥60%)