Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationComputational Methods in HCISoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs

Paper Information

  • Field of Study: Research on methods and practices for handling failures in deep-learning models in computer vision
  • Keywords: Practices, machine learning testing, debugging, interpretability, model deployment, error handling, failure management workflows

Research Background and Problem

  • Problem or Challenge: Deep-learning models may encounter various issues during development and deployment, such as overfitting, data bias, vulnerability, and opaque handling of input data. These issues can lead to erroneous outputs and even harmful impacts in real-world applications. For instance, a model designed to distinguish between benign and malignant moles may perform inaccurately on individuals with darker skin tones.
  • Significance of the Problem: As deep learning becomes increasingly prevalent in computer vision, ensuring the safety and effectiveness of these models has become a critical issue. Properly addressing failures before deployment can significantly reduce potential risks.
  • Motivation and Related Work: Current research on models focuses more on improving accuracy and less on the effectiveness and applicability of failure handling in real-world applications. Additionally, there is a lack of systematic investigation into the actual needs and processes of practitioners, which may lead to a disconnect between research and practical applications.

Solution

  • Research Methodology: The authors conducted semi-structured interviews with 18 machine learning practitioners to explore their goals, workflows, tools, and challenges in handling failures.
  • Innovative Contributions: Through real-world interviews and contextual design tasks, the study reveals gaps in current practices and proposes research opportunities, such as establishing a best practices repository, developing failure management tools, and providing specialized training tailored to practitioners' needs.
  • Implementation Steps and Techniques:
    1. Formulate two research questions: What are practitioners' goals before deploying models? What specific actions do they take?
    2. Design a hypothetical scenario: Conduct failure checking and handling for an existing scene classification deep-learning model.
    3. Encode and thematically analyze the interview content to summarize the framework and key challenges in practice.

Research Findings

  1. Specific Findings:

    • Proposed a structured framework for failure management practices in computer vision models, including workflow steps and a set of related failure-handling issues.
    • Analyzed the relationship between existing methods (e.g., interpretability tools) and practitioners' actual approaches to handling failures.
    • Identified practitioners' needs and design opportunities for using failure management tools and defining the final satisfaction point of a model.
  2. Strengths:

    • Addresses the real needs of practitioners, making it closer to practical applications than purely theoretical approaches.
    • Suggests underexplored design directions, such as leveraging textual and global explanations and designing appropriate interactive failure management tools.
  3. Experimental or Evaluation Results:

    • Most practitioners rarely use existing research methods and tools, relying more on their own experience.
    • Practitioners have subjective judgments regarding when a model is "ready," often prioritizing accuracy while lacking detailed attention to the model's "completeness."
    • Interpretability methods are limited in application, with a focus on local visual explanations (e.g., saliency maps) rather than global or interactive approaches.
  4. Limitations and Future Directions:

    • Limitations: The study used a single task, potentially overlooking the specificities of other application scenarios; it did not deeply explore post-deployment failure management.
    • Future Directions:
      • Develop comprehensive failure management tools that integrate interactive interfaces with data/feature exploration functionalities.
      • Improve knowledge exchange between research and practice by establishing an open, collaborative best practices repository.
      • Design more specific evaluation metrics for model performance and failure management to help developers clarify deployment standards.
      • Incorporate failure management education into computer vision curricula.

Conclusion

Through interviews, this paper provides an in-depth analysis of the current state and pain points of failure management in computer vision models from development to deployment. It proposes a series of research recommendations to enhance practitioners' work efficiency and the effectiveness of failure handling. These recommendations span from methodological tools to educational resources and research opportunities aligned with practical needs, offering valuable insights for strengthening the connection between academic research and real-world applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96542/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581555
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation, Computational Methods in HCI
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers