Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
Authors
Paper Title
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
Publication Info
- Topic area: Evaluation practices for LLM-based products in production settings.
- Keywords: LLM evaluation, results-actionability gap, production systems, interpretive practices, vibe checks, organizational meta-work, qualitative methods, systematization, HCI, practitioner challenges.
Background and Problem
- Problem / challenge: Practitioners face significant challenges in evaluating LLM-based products due to the unpredictable and context-dependent nature of these systems. Existing evaluation frameworks and metrics often fail to provide actionable insights, leaving practitioners unable to translate evaluation results into system improvements.
- Significance: As LLMs are integrated into critical domains like healthcare, education, and enterprise software, inadequate evaluation poses risks such as business failures and societal harm. Effective evaluation is essential for ensuring reliability, safety, and user satisfaction.
- Motivation and related work: Previous studies have documented challenges in LLM evaluation, such as reliance on manual testing and the inadequacy of traditional metrics. However, these studies have focused on well-resourced organizations or academic settings, leaving a gap in understanding how less-resourced teams navigate evaluation. This paper addresses this gap by studying diverse practitioners and identifying a novel challenge: the results-actionability gap.
Solution
- Proposed approach: The study investigates how practitioners evaluate LLM-based products, identifies challenges, and proposes strategies to bridge the results-actionability gap. It emphasizes supporting interpretive practices and systematizing evaluation methods.
- Novelty:
- Empirical account of evaluation practices across diverse organizational contexts, extending prior work focused on single organizations or academic settings.
- Introduction and conceptualization of the results-actionability gap, a novel challenge where evaluation data fails to lead to actionable improvements.
- Actionable strategies for bridging the results-actionability gap through organizational adaptations rather than new metrics.
- Procedure and key techniques:
- Conducted semi-structured interviews with 19 practitioners from diverse sectors.
- Thematic analysis of evaluation practices, challenges, and organizational meta-work.
- Identified ten evaluation practices and five key challenges, including the results-actionability gap.
Results
- Concrete findings:
- Identified ten evaluation practices, including informal vibe checks, user feedback collection, expert collaboration, and attempts at automated testing.
- Documented five challenges: aligning evaluation objectives, defining meaningful constructs, selecting viable methods, overcoming technical barriers, and the results-actionability gap.
- Found that 17 out of 19 participants experienced the results-actionability gap, where evaluation results did not translate into actionable system improvements.
- Advantage over baselines: Unlike prior studies that frame interpretive practices as transitional, this study argues they are necessary adaptations to LLM characteristics. It provides strategies to systematize these practices rather than replace them.
- Experiments / evaluation:
- Interviews spanned diverse sectors (healthcare, education, enterprise software) and roles (data scientists, designers, engineers).
- Analysis focused on evaluation execution, design, and organizational meta-work.
- Limitations and future work:
- Sample size (N = 19) may not capture teams that have abandoned evaluation entirely.
- Findings are based on self-reported practices; ethnographic studies could provide deeper insights.
- Future work could explore longitudinal patterns and organizational factors influencing evaluation.
Summary
This study investigates how practitioners evaluate LLM-based products in production settings, identifying ten evaluation practices and five challenges. The results-actionability gap, where evaluation data fails to lead to actionable improvements, emerged as a central issue affecting 17 of 19 participants. The study reframes interpretive practices like vibe checks as necessary adaptations to LLM characteristics rather than methodological failures. It proposes actionable strategies, including evaluation-by-design, continuous sense-making, and incremental testing, to bridge the results-actionability gap. These findings highlight opportunities for HCI research to support practitioners in systematizing their evaluation methods.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows
CHI '26· Human-LLM Collaboration +2
- 100%
RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
UIST '25· Human-LLM Collaboration +2
- 83%
Interactive Debugging and Steering of Multi-Agent AI Systems
CHI '25· Human-LLM Collaboration +2
- 83%
DIY: Helping People Assess the Correctness of Natural Language to SQL Systems
IUI '21· Human-LLM Collaboration +2
- 83%
CoPrompter: User-Centric Evaluation of LM Instruction Alignment for Improved Prompt Engineering
IUI '25· Human-LLM Collaboration +2
- 75%
DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models
CHI '24· Human-LLM Collaboration +3
- 71%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 71%
Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
CHI '24· Human-LLM Collaboration +2
- 71%
Dango: A Mixed-Initiative Data Wrangling System using Large Language Model
CHI '25· Human-LLM Collaboration +2
- 71%
OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question Answering
CHI '25· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)