EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences
Authors
With the broad availability of large language models and their ability to generate vast outputs using varied prompts and configurations, determining the best output for a given task requires an intensive evaluation process, one where machine learning practitioners must decide how to assess the outputs and then carefully carry out the evaluation. This process is both time-consuming and costly. As practitioners work with an increasing number of models, they must now evaluate outputs to determine which model performs best for a given task. LLMs are increasingly used as evaluators to filter training data, evaluate model performance or assist human evaluators with detailed assessments. Our application, EvalAssist, supports this process by aiding users in interactively refining evaluation criteria. In our study with machine learning practitioners (n=15), each completing 6 tasks yielding 131 evaluations, we explore how task-related factors and judgment strategies influence criteria refinement and user perceptions. Findings show that users performed more evaluations with direct assessment by making criteria task-specific, modifying judgments, and changing the AI evaluator model. We conclude with recommendations for how systems can better support practitioners with AI-assisted evaluations.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
- 83%
Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging System
CHI '26· Interactive Data Visualization +2
- 83%
The Bots of Persuasion: Examining How Conversational Agents' Linguistic Expressions of Personality Affect User Perceptions and Decisions
CHI '26· Agent Personality & Anthropomorphism +2
- 83%
Belief Updating and Delegation in Multi-Task Human–AI Interaction: Evidence from Controlled Simulations
CHI '26· Human-LLM Collaboration +2
- 83%
Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
IUI '26· Human-LLM Collaboration +2
- 83%
Making Absence Visible in Intelligent Summarization Interfaces
IUI '26· Human-LLM Collaboration +2
- 83%
"Un-default" Behavior Tuning: Specifying Model Behavior outside the Norm with LLM Self-Playing and Self-Improving
IUI '26· Human-LLM Collaboration +2
- 80%
Effects of Communication Directionality and AI Agent Differences in Human-AI Interaction
CHI '21· Human-LLM Collaboration +1
- 80%
AI Knowledge: Improving AI Delegation through Human Enablement
CHI '23· Human-LLM Collaboration +1
- 80%
How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz Study
CHI '24· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)