Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationExplainable AI (XAI)AI/ML Researchers & EngineersSoftware Engineers & DevelopersData Scientists & Analysts

Paper Title

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild

Publication Info

  • Topic area: Evaluation practices for LLM-based products in production settings.
  • Keywords: LLM evaluation, results-actionability gap, production systems, interpretive practices, vibe checks, organizational meta-work, qualitative methods, systematization, HCI, practitioner challenges.

Background and Problem

  • Problem / challenge: Practitioners face significant challenges in evaluating LLM-based products due to the unpredictable and context-dependent nature of these systems. Existing evaluation frameworks and metrics often fail to provide actionable insights, leaving practitioners unable to translate evaluation results into system improvements.
  • Significance: As LLMs are integrated into critical domains like healthcare, education, and enterprise software, inadequate evaluation poses risks such as business failures and societal harm. Effective evaluation is essential for ensuring reliability, safety, and user satisfaction.
  • Motivation and related work: Previous studies have documented challenges in LLM evaluation, such as reliance on manual testing and the inadequacy of traditional metrics. However, these studies have focused on well-resourced organizations or academic settings, leaving a gap in understanding how less-resourced teams navigate evaluation. This paper addresses this gap by studying diverse practitioners and identifying a novel challenge: the results-actionability gap.

Solution

  • Proposed approach: The study investigates how practitioners evaluate LLM-based products, identifies challenges, and proposes strategies to bridge the results-actionability gap. It emphasizes supporting interpretive practices and systematizing evaluation methods.
  • Novelty:
    1. Empirical account of evaluation practices across diverse organizational contexts, extending prior work focused on single organizations or academic settings.
    2. Introduction and conceptualization of the results-actionability gap, a novel challenge where evaluation data fails to lead to actionable improvements.
    3. Actionable strategies for bridging the results-actionability gap through organizational adaptations rather than new metrics.
  • Procedure and key techniques:
    • Conducted semi-structured interviews with 19 practitioners from diverse sectors.
    • Thematic analysis of evaluation practices, challenges, and organizational meta-work.
    • Identified ten evaluation practices and five key challenges, including the results-actionability gap.

Results

  • Concrete findings:
    • Identified ten evaluation practices, including informal vibe checks, user feedback collection, expert collaboration, and attempts at automated testing.
    • Documented five challenges: aligning evaluation objectives, defining meaningful constructs, selecting viable methods, overcoming technical barriers, and the results-actionability gap.
    • Found that 17 out of 19 participants experienced the results-actionability gap, where evaluation results did not translate into actionable system improvements.
  • Advantage over baselines: Unlike prior studies that frame interpretive practices as transitional, this study argues they are necessary adaptations to LLM characteristics. It provides strategies to systematize these practices rather than replace them.
  • Experiments / evaluation:
    • Interviews spanned diverse sectors (healthcare, education, enterprise software) and roles (data scientists, designers, engineers).
    • Analysis focused on evaluation execution, design, and organizational meta-work.
  • Limitations and future work:
    • Sample size (N = 19) may not capture teams that have abandoned evaluation entirely.
    • Findings are based on self-reported practices; ethnographic studies could provide deeper insights.
    • Future work could explore longitudinal patterns and organizational factors influencing evaluation.

Summary

This study investigates how practitioners evaluate LLM-based products in production settings, identifying ten evaluation practices and five challenges. The results-actionability gap, where evaluation data fails to lead to actionable improvements, emerged as a central issue affecting 17 of 19 participants. The study reframes interpretive practices like vibe checks as necessary adaptations to LLM characteristics rather than methodological failures. It proposes actionable strategies, including evaluation-by-design, continuous sense-making, and incremental testing, to bridge the results-actionability gap. These findings highlight opportunities for HCI research to support practitioners in systematizing their evaluation methods.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222094/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791069
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Explainable AI (XAI)
work
Professions
AI/ML Researchers & Engineers, Software Engineers & Developers, Data Scientists & Analysts
article
Content Status
Full text indexed
hub
Related Papers
10 related papers