Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine Learning
Honorable MentionAuthors
Machine learning (ML) models are now routinely deployed in domains ranging from criminal justice to healthcare. With this newfound ubiquity, ML has moved beyond academia and grown into an engineering discipline. To that end, interpretability tools have been designed to help data scientists and machine learning practitioners better understand how ML models work. However, there has been little evaluation of the extent to which these tools achieve this goal. We study data scientists' use of two existing interpretability tools, the InterpretML implementation of GAMs and the SHAP Python package. We conduct a contextual inquiry (N=11) and a survey (N=197) of data scientists to observe how they use interpretability tools to uncover common issues that arise when building and evaluating ML models. Our results indicate that data scientists over-trust and misuse interpretability tools. Furthermore, few of our participants were able to accurately describe the visualizations output by these tools. We highlight qualitative themes for data scientists' mental models of interpretability tools. We conclude with implications for researchers and tool designers, and contextualize our findings in the social science literature.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Rules or Weights? Comparing User Understanding of Explainable AI Techniques with the Cognitive XAI-Adaptive Model
IUI '26· Explainable AI (XAI) +2
- 71%
Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models
CHI '19· Explainable AI (XAI) +2
- 71%
No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive ML
CHI '20· Explainable AI (XAI) +2
- 71%
PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule Synthesis
CHI '23· Explainable AI (XAI) +2
- 71%
ESCAPE: Countering Systematic Errors from Machine's Blind Spots via Interactive Visual Analysis
CHI '23· Explainable AI (XAI) +2
- 71%
Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
CHI '23· Explainable AI (XAI) +2
- 71%
RELIC: Investigating Large Language Model Responses using Self-Consistency
CHI '24· Explainable AI (XAI) +2
- 71%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
- 71%
“I Don’t Think RAI Applies to My Model” – Engaging Non-champions with Sticky Stories for Responsible AI Work
CHI '26· AI Ethics, Fairness & Accountability +2
- 71%
Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
CHI '26· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)