Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
Authors
Title of the Paper
Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
Paper Information
- Field of Study: Research on methods and practices for handling failures in deep-learning models in computer vision
- Keywords: Practices, machine learning testing, debugging, interpretability, model deployment, error handling, failure management workflows
Research Background and Problem
- Problem or Challenge: Deep-learning models may encounter various issues during development and deployment, such as overfitting, data bias, vulnerability, and opaque handling of input data. These issues can lead to erroneous outputs and even harmful impacts in real-world applications. For instance, a model designed to distinguish between benign and malignant moles may perform inaccurately on individuals with darker skin tones.
- Significance of the Problem: As deep learning becomes increasingly prevalent in computer vision, ensuring the safety and effectiveness of these models has become a critical issue. Properly addressing failures before deployment can significantly reduce potential risks.
- Motivation and Related Work: Current research on models focuses more on improving accuracy and less on the effectiveness and applicability of failure handling in real-world applications. Additionally, there is a lack of systematic investigation into the actual needs and processes of practitioners, which may lead to a disconnect between research and practical applications.
Solution
- Research Methodology: The authors conducted semi-structured interviews with 18 machine learning practitioners to explore their goals, workflows, tools, and challenges in handling failures.
- Innovative Contributions: Through real-world interviews and contextual design tasks, the study reveals gaps in current practices and proposes research opportunities, such as establishing a best practices repository, developing failure management tools, and providing specialized training tailored to practitioners' needs.
- Implementation Steps and Techniques:
- Formulate two research questions: What are practitioners' goals before deploying models? What specific actions do they take?
- Design a hypothetical scenario: Conduct failure checking and handling for an existing scene classification deep-learning model.
- Encode and thematically analyze the interview content to summarize the framework and key challenges in practice.
Research Findings
-
Specific Findings:
- Proposed a structured framework for failure management practices in computer vision models, including workflow steps and a set of related failure-handling issues.
- Analyzed the relationship between existing methods (e.g., interpretability tools) and practitioners' actual approaches to handling failures.
- Identified practitioners' needs and design opportunities for using failure management tools and defining the final satisfaction point of a model.
-
Strengths:
- Addresses the real needs of practitioners, making it closer to practical applications than purely theoretical approaches.
- Suggests underexplored design directions, such as leveraging textual and global explanations and designing appropriate interactive failure management tools.
-
Experimental or Evaluation Results:
- Most practitioners rarely use existing research methods and tools, relying more on their own experience.
- Practitioners have subjective judgments regarding when a model is "ready," often prioritizing accuracy while lacking detailed attention to the model's "completeness."
- Interpretability methods are limited in application, with a focus on local visual explanations (e.g., saliency maps) rather than global or interactive approaches.
-
Limitations and Future Directions:
- Limitations: The study used a single task, potentially overlooking the specificities of other application scenarios; it did not deeply explore post-deployment failure management.
- Future Directions:
- Develop comprehensive failure management tools that integrate interactive interfaces with data/feature exploration functionalities.
- Improve knowledge exchange between research and practice by establishing an open, collaborative best practices repository.
- Design more specific evaluation metrics for model performance and failure management to help developers clarify deployment standards.
- Incorporate failure management education into computer vision curricula.
Conclusion
Through interviews, this paper provides an in-depth analysis of the current state and pain points of failure management in computer vision models from development to deployment. It proposes a series of research recommendations to enhance practitioners' work efficiency and the effectiveness of failure handling. These recommendations span from methodological tools to educational resources and research opportunities aligned with practical needs, offering valuable insights for strengthening the connection between academic research and real-world applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What are the specific needs and practices for handling computer vision deep learning model failures before deployment?Category: Coding Assistants and Multi-Turn Code SupportSimilar questionsarrow_forward
- What is the gap between existing methods (e.g., explainability tools) and practitioners' actual approaches in failure management?Category: Coding Assistants and Multi-Turn Code SupportSimilar questionsarrow_forward
- How can new tools and training meet practitioners' needs in model failure handling?Category: Coding Assistants and Multi-Turn Code SupportSimilar questionsarrow_forward
Practical Problems
1- ML developers lack efficient tools and workflows for handling model errors, hindering model deployment.Category: Coding Assistants and Multi-Turn Code SupportSimilar questionsarrow_forward
- 100%
PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule Synthesis
CHI '23· Explainable AI (XAI) +2
- 100%
Unakite: Scaffolding Developers’ Decision-Making Using the Web
UIST '19· Explainable AI (XAI) +2
- 86%
DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
CHI '26· Human-LLM Collaboration +3
- 83%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 83%
Rules or Weights? Comparing User Understanding of Explainable AI Techniques with the Cognitive XAI-Adaptive Model
IUI '26· Explainable AI (XAI) +2
- 75%
Towards Guidelines for Designing Human-in-the-Loop Machine Training Interfaces
IUI '21· Human-LLM Collaboration +3
- 71%
Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine Learning
CHI '20· Explainable AI (XAI) +2
- 71%
Evaluating the Interpretability of Generative Models by Interactive Reconstruction
CHI '21· Explainable AI (XAI) +2
- 71%
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
CHI '25· Explainable AI (XAI) +2
- 71%
"Should I Rely on You or the AI?" Leaders' Trust and Perceptions in Mixed Human-AI Teams
CHI '26· Human-Robot Collaboration (HRC) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)