iSEA : An Interactive Pipeline for Semantic Error Analysis of NLP Models
Error analysis in NLP models is essential to successful model development and deployment. One common approach for diagnosing errors is to identify subpopulations in the dataset where the model produces the most errors. However, existing approaches typically define subpopulations based on pre-defined features, which requires users to form hypotheses of errors in advance. To complement these approaches, we propose iSEA, an Interactive Pipeline for Semantic Error Analysis in NLP Models, which automatically discovers semantically-grounded subpopulations with high error rates in the context of a human-in-the-loop interactive system. iSEA enables model developers to learn more about their model errors through discovered subpopulations, validate the sources of errors through interactive analysis on the discovered subpopulations, and test hypotheses about model errors by defining custom subpopulations. The tool supports semantic descriptions of error-prone subpopulations at the token and concept level, as well as pre-defined higher-level features. Through use cases and expert interviews, we demonstrate how iSEA can assist error understanding and analysis.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can NLP model error analysis be improved by automatically discovering semantically related error subgroups?Category: NLP Model Error Analysis and Visualization SupportSimilar questionsarrow_forward
- How can multi-level semantic features (e.g., lexical, conceptual, high-level features) explain NLP model error distributions and behavior?Category: NLP Model Error Analysis and Visualization SupportSimilar questionsarrow_forward
- Are graphical user interfaces effective in supporting non-technical users in semantic error analysis?Category: NLP Model Error Analysis and Visualization SupportSimilar questionsarrow_forward
Practical Problems
1- NLP model error sources are difficult to understand, and existing tools offer limited support for non-technical users.Category: NLP Model Error Analysis and Visualization SupportSimilar questionsarrow_forward
- 100%
Angler: Helping Machine Translation Practitioners Prioritize Model Improvements
CHI '23· Explainable AI (XAI) +2
- 80%
UMLAUT: Debugging Deep Learning Programs using Program Structure and Model Behavior
CHI '21· Explainable AI (XAI) +1
- 80%
Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive Theories
CHI '25· Explainable AI (XAI) +1
- 71%
StepMIND: A Visual Framework for Stepwise, Multimodal, and Bidirectional Explanations of AI-Generated Data Analysis Pipeline
IUI '26· Explainable AI (XAI) +3
- 67%
Questioning the AI: Informing Design Practices for Explainable AI User Experiences
CHI '20· Explainable AI (XAI) +1
- 67%
AI-Moderated Decision-Making: Capturing and Balancing Anchoring Bias in Sequential Decision Tasks
CHI '22· Explainable AI (XAI) +2
- 67%
How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?
CHI '22· Explainable AI (XAI) +1
- 67%
Zeno: An Interactive Framework for Behavioral Evaluation of Machine Learning
CHI '23· Explainable AI (XAI) +1
- 67%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 67%
"AI enhances our performance, I have no doubt this one will do the same": The Placebo effect is robust to negative descriptions of AI
CHI '24· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)