How Accurate Does It Feel? - Human Perception of Different Types of Classification Mistakes
Authors
Title of the Paper
How Accurate Does It Feel? – Human Perception of Different Types of Classification Mistakes
Paper Information
- Research Area: Human-Computer Interaction (HCI) and User Perception of Machine Learning Systems
- Keywords: Accuracy, User Perception, Label Noise, Classification Errors, Data Quality, User Study
Research Background and Problem
-
Identified Issues or Challenges:
- Supervised learning systems are typically trained on datasets with "ground truth" labels, but these labels may contain errors or subjective noise.
- Classification errors in the data may not align with users' perceived experience, making traditional accuracy metrics insufficient to reflect user experience accurately.
- Label errors or ambiguities in the data can significantly impact human trust and acceptance of algorithms.
-
Significance:
- Accurately evaluating machine learning system performance is critical for improving user trust, optimizing user experience, and avoiding societal issues such as bias.
- The impact of classification difficulty and prediction errors on user perception has not been thoroughly studied, which could help refine model evaluation standards.
-
Research Motivation and Related Work:
- Previous studies have shown that users may lose trust in a system after observing classification errors, potentially disregarding its recommendations.
- Traditional metrics such as accuracy, precision, and recall fail to capture users' subjective perception of system performance.
- This study focuses on how users perceive errors in classifications of varying difficulty and proposes a human-centered approach to system evaluation.
Solution
-
Proposed Approach or Methodology:
- Design artificial classifiers to simulate different types of classification errors (e.g., easy-to-classify, hard-to-classify, and impossible-to-classify cases).
- Collect users' subjective evaluations of classifier performance and compare these with traditional metrics (e.g., accuracy, F1 score).
-
Innovative Contributions:
- Introduces the concept of "perceived accuracy" by incorporating user-perceived uncertainty into system evaluation metrics, studying the impact of classification difficulty on user perception.
- Highlights the gap between traditional machine learning performance evaluation and user perception, suggesting the importance of classification difficulty in human-computer interaction tasks.
-
Implementation Steps and Key Techniques:
- Data Construction:
- Use crowdsourcing tasks to label sentences in a binary classification dataset as easy-to-classify, hard-to-classify, or impossible-to-classify.
- Create two datasets for experiments: one with easy-to-classify data and one with mixed classification difficulty.
- Experimental Design:
- Simulate multiple "classifiers" with varying error distributions (on easy, hard, and impossible-to-classify sentences).
- Have participants interact with the classifiers and provide subjective evaluations of their performance ("perceived accuracy").
- Evaluation Methods:
- Assess the differences between users' "calculated accuracy" (traditional metrics) and "perceived accuracy."
- Use statistical tests to examine the significance of the impact of classification errors on user perception under different experimental conditions.
- Data Construction:
Research Findings
-
Key Findings:
- Errors on easy-to-classify data lead users to significantly underestimate system performance (lower perceived accuracy).
- Errors on hard-to-classify and impossible-to-classify data have less impact on user perception, and users may even overestimate system performance in such cases.
- The inherent difficulty of the dataset itself can reduce users' perceived accuracy of a classifier, even if it achieves 100% accuracy.
-
Relative Advantages:
- Compared to traditional accuracy metrics, analyzing perceived accuracy provides deeper insights into the interaction mechanisms between users and machine learning models.
- Experimental results support the introduction of new evaluation methods that better reflect system performance from the user's perspective.
-
Experimental or Evaluation Results:
- User subjective evaluations of classifiers showed significant differences across conditions with different types of classification errors, confirming the relationship between perceived accuracy and error type.
- Tests of traditional metrics (e.g., precision, F1 score) revealed systematic biases in their ability to reflect user perception.
-
Limitations and Future Directions:
- Limitations:
- The experiment required participants to make binary judgments for each classification task, which may limit its applicability to real-world scenarios.
- The study's focus on a single task type (text classification) may restrict the generalizability of the findings.
- Future Directions:
- Extend the research to machine learning tasks involving more complex data types (e.g., images and audio).
- Explore non-binary user choices to investigate more granular perception evaluations.
- Examine the impact of different interaction scenarios and task settings on users' perceived accuracy.
- Limitations:
The study demonstrates that traditional accuracy metrics alone are insufficient to reflect system performance from the user's perspective. It emphasizes the necessity of developing new evaluation standards from a human perception standpoint, particularly in scenarios involving classification difficulty.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do users perceive the impact of different types of classification errors on system accuracy?Category: Metric Comprehension and Analytical Explanation SupportSimilar questionsarrow_forward
- How does classification task difficulty affect users' subjective evaluation of system performance?Category: Metric Comprehension and Analytical Explanation SupportSimilar questionsarrow_forward
- Why cannot traditional accuracy metrics adequately reflect user-perceived classification performance?Category: Metric Comprehension and Analytical Explanation SupportSimilar questionsarrow_forward
Practical Problems
1- Users often lose trust in machine learning systems upon seeing classification errors.Category: Algorithm Aversion, User Control, and Trust CalibrationSimilar questionsarrow_forward
- 80%
Toward Foraging for Understanding of StarCraft Agents: An Empirical Study
IUI '18· Explainable AI (XAI) +2
- 75%
COGAM: Measuring and Moderating Cognitive Load in Machine Learning Model Explanations
CHI '20· Explainable AI (XAI)
- 75%
Reading Between the Pixels: Investigating the Barriers to Visualization Literacy
CHI '24· Visualization Perception & Cognition
- 75%
Did You Misclick? Reversing 5-Point Satisfaction Scales Causes Unintended Responses
CHI '24· Visualization Perception & Cognition
- 60%
TopoText: Context-Preserving Text Data Exploration Across Multiple Spatial Scales
CHI '18· Interactive Data Visualization +1
- 60%
How Do We Measure That?! Quick Scale Development
CHI '18· Visualization Perception & Cognition +1
- 60%
BubbleView: An Interface for Crowdsourcing Image Importance Maps and Tracking Visual Attention
CHI '18· Eye Tracking & Gaze Interaction +1
- 60%
Understanding Digitally-Mediated Empathy: An Exploration of Visual, Narrative, and Biosensory Informational Cues
CHI '19· Visualization Perception & Cognition +1
- 60%
The Low/High Index of Pupillary Activity
CHI '20· Eye Tracking & Gaze Interaction +1
- 60%
How Relevant is Hick's Law for HCI?
CHI '20· Visualization Perception & Cognition +1
Based on Jaccard similarity of research subtopics & professions (≥60%)