fAIlureNotes: Supporting Designers in Understanding the Limits of AI Models for Computer Vision Tasks
Authors
Title of the Paper
fAIlureNotes: Supporting Designers in Understanding the Limits of AI Models for Computer Vision Tasks
Paper Information
- Subject Area: Human-Computer Interaction, Design Tools, AI Error Analysis
- Keywords: AI Design, Human-Computer Interaction, Computer Vision Models, Failure Analysis, User Experience Design, Tool Development, Model Exploration, Error Classification
Research Background and Problem
-
What problems or challenges did the authors identify?
- When designing AI-driven products, UX designers need to evaluate the alignment between the model and user needs. However, designers currently lack methods to explore AI model behavior and its limitations.
- There is a gap between user research and AI model behavior, which is particularly pronounced when designers lack technical knowledge.
- Existing exploration tools (e.g., interactive model cards) are inefficient, leading to model exploration being primarily conducted through resource-intensive experiments or post-launch user feedback.
-
Why is this problem important?
- If the limitations of AI models are not identified during the design phase, it can result in costly rework and user experience issues after launch.
- AI model failures are often tied to specific user data and usage scenarios; failing to identify these issues in advance can have adverse effects on users.
- Effective tools and practices can help designers better understand AI models, ultimately creating more human-centered and high-quality product designs.
-
Research Motivation and Related Work
- Previous research has shown that designers without tool support often feel frustrated and helpless when working with machine learning (ML) models.
- Current model behavior analysis tools are primarily designed for AI engineers, not UX designers.
- This study aims to design a tool focused on helping UX practitioners understand AI model limitations early in the design process to avoid costly fixes and iterations later.
Solution
-
What methods or solutions did the authors propose?
- The authors proposed and implemented a tool called "fAIlureNotes," which supports designers in exploring AI model behavior and failure patterns early in the design process through user research-driven workflows.
- By introducing an automated error classification engine combined with text-to-image generation models, fAIlureNotes enables comprehensive AI failure analysis from a user perspective.
-
What are the innovative aspects of this solution?
- Integration of User Scenarios: The tool allows designers to build user scenarios based on user research, extending model exploration to specific user contexts.
- Iterative Failure Discovery: It provides data augmentation and generation features (e.g., image generation and editing) to guide designers in testing complex model failure cases.
- Intuitive Design Overview: It offers a canvas to visually present failure patterns and supports designers in formulating specific recovery strategies for errors.
- Automated Failure Classification Engine: It categorizes and labels error types (e.g., False Detection, Unnecessary Detection, Out-of-Distribution) based on failure patterns.
-
What are the implementation steps and key technologies used?
- Importing User Research Data: Designers create user personas and task scenarios in the tool, upload relevant images, or use built-in generation models to create input data.
- Model Exploration and Error Classification: The tool runs pre-trained AI models (e.g., DETR), compares model outputs with user expectations, and automatically generates classification labels (e.g., error types).
- Iterative Exploration Expansion: Using features like generation prompts and image augmentation, the system guides designers in testing different hypotheses.
- Design Synthesis and Failure Review: The tool provides a canvas summarizing failure cards, enabling designers to develop a systematic understanding of errors and formulate design recommendations.
Research Outcomes
-
What specific outcomes were achieved?
- fAIlureNotes significantly improved the depth and quality of designers' identification of model failure patterns.
- The study showed that compared to existing interactive model cards, fAIlureNotes offered significant advantages in supporting analogy to user scenarios, error grouping, and summarization.
- Designers were able to use the tool to propose and document detailed design intervention strategies, such as implementing user feedback mechanisms and adding local and global model explanations.
-
What advantages does it have compared to existing solutions?
- Compared to interactive model card exploration tools (e.g., HuggingFace):
- fAIlureNotes is more intuitive and advanced in integrating user context characteristics and failure information.
- It supports a complete workflow, from exploration to design synthesis, rather than isolated model testing.
- It automates error classification and provides recovery suggestions, reducing cognitive load for designers.
- Compared to interactive model card exploration tools (e.g., HuggingFace):
-
What were the experimental or evaluation results?
- User research and evaluation with 10 UX designers revealed that fAIlureNotes significantly improved error detection efficiency and the quality of design intervention suggestions.
- During actual use by designers, the tool effectively reduced the need to switch between tools, increasing focus on design tasks.
- fAIlureNotes enabled designers to create an average of 1.6 failure groups and 1.4 recovery strategies per user.
-
Limitations and Future Directions
- Limitations:
- The tool currently supports only a single task (object detection) and does not yet cover a broader range of computer vision tasks or other domains.
- It has not yet integrated cross-team collaboration features, and interaction between engineers and designers needs further optimization.
- Error understanding is still limited to the subjective analysis of designers, making it challenging to address more complex socio-technical issues (e.g., fairness and ethics).
- Future Directions:
- Expand tool support to other AI tasks, such as text classification and large language models.
- Support dataset management, subset analysis, and multi-model performance comparison.
- Evaluate tool performance and multi-stakeholder usage in real-world industrial scenarios.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can designers more efficiently understand the limitations of computer vision AI models?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- Which tools can help users establish connections between their research and AI model behavior?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- How can automated error classification provide designers with intuitive model failure analysis?Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
Practical Problems
1- Designers struggle to identify AI model limitations during product design, leading to experience issues and high costs.Category: AI/LLM as Design Collaborators and Creative ToolsSimilar questionsarrow_forward
- 80%
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
CHI '22· Explainable AI (XAI) +1
- 67%
Questioning the AI: Informing Design Practices for Explainable AI User Experiences
CHI '20· Explainable AI (XAI) +1
- 67%
How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?
CHI '22· Explainable AI (XAI) +1
- 67%
Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User Experience
CHI '23· Human-LLM Collaboration +2
- 67%
Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making
CHI '23· Explainable AI (XAI) +1
- 67%
Zeno: An Interactive Framework for Behavioral Evaluation of Machine Learning
CHI '23· Explainable AI (XAI) +1
- 67%
VIME: Visual Interactive Model Explorer for Identifying Capabilities and Limitations of Machine Learning Models for Sequential Decision-Making
UIST '24· Eye Tracking & Gaze Interaction +2
Based on Jaccard similarity of research subtopics & professions (≥60%)