How Accurate Does It Feel? - Human Perception of Different Types of Classification Mistakes

Explainable AI (XAI)Visualization Perception & CognitionHCI ResearchersCognitive Scientists

Title of the Paper

How Accurate Does It Feel? – Human Perception of Different Types of Classification Mistakes

Paper Information

  • Research Area: Human-Computer Interaction (HCI) and User Perception of Machine Learning Systems
  • Keywords: Accuracy, User Perception, Label Noise, Classification Errors, Data Quality, User Study

Research Background and Problem

  • Identified Issues or Challenges:

    • Supervised learning systems are typically trained on datasets with "ground truth" labels, but these labels may contain errors or subjective noise.
    • Classification errors in the data may not align with users' perceived experience, making traditional accuracy metrics insufficient to reflect user experience accurately.
    • Label errors or ambiguities in the data can significantly impact human trust and acceptance of algorithms.
  • Significance:

    • Accurately evaluating machine learning system performance is critical for improving user trust, optimizing user experience, and avoiding societal issues such as bias.
    • The impact of classification difficulty and prediction errors on user perception has not been thoroughly studied, which could help refine model evaluation standards.
  • Research Motivation and Related Work:

    • Previous studies have shown that users may lose trust in a system after observing classification errors, potentially disregarding its recommendations.
    • Traditional metrics such as accuracy, precision, and recall fail to capture users' subjective perception of system performance.
    • This study focuses on how users perceive errors in classifications of varying difficulty and proposes a human-centered approach to system evaluation.

Solution

  • Proposed Approach or Methodology:

    • Design artificial classifiers to simulate different types of classification errors (e.g., easy-to-classify, hard-to-classify, and impossible-to-classify cases).
    • Collect users' subjective evaluations of classifier performance and compare these with traditional metrics (e.g., accuracy, F1 score).
  • Innovative Contributions:

    • Introduces the concept of "perceived accuracy" by incorporating user-perceived uncertainty into system evaluation metrics, studying the impact of classification difficulty on user perception.
    • Highlights the gap between traditional machine learning performance evaluation and user perception, suggesting the importance of classification difficulty in human-computer interaction tasks.
  • Implementation Steps and Key Techniques:

    • Data Construction:
      • Use crowdsourcing tasks to label sentences in a binary classification dataset as easy-to-classify, hard-to-classify, or impossible-to-classify.
      • Create two datasets for experiments: one with easy-to-classify data and one with mixed classification difficulty.
    • Experimental Design:
      • Simulate multiple "classifiers" with varying error distributions (on easy, hard, and impossible-to-classify sentences).
      • Have participants interact with the classifiers and provide subjective evaluations of their performance ("perceived accuracy").
    • Evaluation Methods:
      • Assess the differences between users' "calculated accuracy" (traditional metrics) and "perceived accuracy."
      • Use statistical tests to examine the significance of the impact of classification errors on user perception under different experimental conditions.

Research Findings

  • Key Findings:

    • Errors on easy-to-classify data lead users to significantly underestimate system performance (lower perceived accuracy).
    • Errors on hard-to-classify and impossible-to-classify data have less impact on user perception, and users may even overestimate system performance in such cases.
    • The inherent difficulty of the dataset itself can reduce users' perceived accuracy of a classifier, even if it achieves 100% accuracy.
  • Relative Advantages:

    • Compared to traditional accuracy metrics, analyzing perceived accuracy provides deeper insights into the interaction mechanisms between users and machine learning models.
    • Experimental results support the introduction of new evaluation methods that better reflect system performance from the user's perspective.
  • Experimental or Evaluation Results:

    • User subjective evaluations of classifiers showed significant differences across conditions with different types of classification errors, confirming the relationship between perceived accuracy and error type.
    • Tests of traditional metrics (e.g., precision, F1 score) revealed systematic biases in their ability to reflect user perception.
  • Limitations and Future Directions:

    • Limitations:
      • The experiment required participants to make binary judgments for each classification task, which may limit its applicability to real-world scenarios.
      • The study's focus on a single task type (text classification) may restrict the generalizability of the findings.
    • Future Directions:
      • Extend the research to machine learning tasks involving more complex data types (e.g., images and audio).
      • Explore non-binary user choices to investigate more granular perception evaluations.
      • Examine the impact of different interaction scenarios and task settings on users' perceived accuracy.

The study demonstrates that traditional accuracy metrics alone are insufficient to reflect system performance from the user's perspective. It emphasizes the necessity of developing new evaluation standards from a human perception standpoint, particularly in scenarios involving classification difficulty.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68760/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501915
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Explainable AI (XAI), Visualization Perception & Cognition
work
Professions
HCI Researchers, Cognitive Scientists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers