iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries
Authors
The recent explosion in popularity of large language models (LLMs) has inspired learning engineers to incorporate them into adaptive educational tools that automatically score summary writing. Understanding and evaluating LLMs is vital before deploying them in critical learning environments, yet their unprecedented size and expanding number of parameters inhibits transparency and impedes trust when they underperform. Through a collaborative user-centered design process with several learning engineers building and deploying summary scoring LLMs, we characterized fundamental design challenges and goals around interpreting their models, including aggregating large text inputs, tracking score provenance, and scaling LLM interpretability methods. To address their concerns, we developed iScore, an interactive visual analytics tool for learning engineers to upload, score, and compare multiple summaries simultaneously. Tightly integrated views allow users to iteratively revise the language in summaries, track changes in the resulting LLM scores, and visualize model weights at multiple levels of abstraction. To validate our approach, we deployed iScore with three learning engineers over the course of a month. We present a case study where interacting with iScore led a learning engineer to improve their LLM's score accuracy by three percentage points. Finally, we conducted qualitative interviews with the learning engineers that revealed how iScore enabled them to understand, evaluate, and build trust in their LLMs during deployment.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can visual analytics tools explain how large language models work in automated essay scoring?Category: Machine Learning Model Visualization, Debugging, and Explainability SupportSimilar questionsarrow_forward
- How can input perturbation and attention visualization reveal potential bias and logical issues in model scoring?Category: Machine Learning Model Visualization, Debugging, and Explainability SupportSimilar questionsarrow_forward
- How can multi-view integration methods improve model explainability and user-friendliness?Category: Machine Learning Model Visualization, Debugging, and Explainability SupportSimilar questionsarrow_forward
Practical Problems
1- In educational technology, learning engineers struggle to understand language model essay scoring mechanisms.Category: Machine Learning Model Visualization, Debugging, and Explainability SupportSimilar questionsarrow_forward
- 83%
How Do Analysts Understand and Verify AI-Assisted Data Analyses?
CHI '24· Human-LLM Collaboration +2
- 83%
Interactive Explainable Ranking
CHI '26· Explainable AI (XAI) +2
- 80%
FDHelper: Assist Unsupervised Fraud Detection Experts with Interactive Feature Selection and Evaluation
CHI '20· Explainable AI (XAI) +1
- 80%
Trustworthy by Design: The Viewer's Perspective on Trust in Data Visualization
CHI '25· Explainable AI (XAI) +1
- 80%
Inferencing Underspecified Natural Language Utterances in Visual Analysis
IUI '19· Explainable AI (XAI) +1
- 71%
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
CHI '25· Human-LLM Collaboration +2
- 67%
Considering Agency and Data Granularity in the Design of Visualization Tools
CHI '18· Explainable AI (XAI) +2
- 67%
Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making
CHI '23· Explainable AI (XAI) +1
- 67%
AnchorViz: Facilitating Classifier Error Discovery through Interactive Semantic Data Exploration
IUI '18· Explainable AI (XAI) +2
- 67%
The Effects of Customisation on the Usability of Visual Analytics Dashboards: the Good, the Bad, and the Ugly
IUI '25· Interactive Data Visualization +1
Based on Jaccard similarity of research subtopics & professions (≥60%)