Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationUser Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingAI/ML Researchers & EngineersHCI ResearchersData Scientists & Analysts

Paper Title

Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria

Publication Info

  • Topic area: Evaluation workflows for Large Language Models (LLMs) in domain-specific tasks.
  • Keywords: Large Language Models, evaluation criteria, domain experts, lay users, human-AI collaboration, staged workflows, criteria drift, usability testing, natural language processing, LLM-as-a-Judge.

Background and Problem

  • Problem / challenge: Existing evaluation methods for LLMs often fail to account for the unique contributions of diverse sources (domain experts, lay users, and LLMs) and how evaluation criteria evolve across different stages of the evaluation process. Current approaches are limited by cost, scalability, and alignment with domain-specific standards.
  • Significance: Effective evaluation of LLMs is essential to ensure accuracy, reliability, and alignment with end-user needs, particularly in high-stakes domains like healthcare and education, where errors can have significant consequences.
  • Motivation and related work: Prior research has focused on tools and systems for LLM evaluation (e.g., EvalLM, EvalAssist) and the role of human-AI collaboration. However, gaps remain in understanding how different sources contribute to evaluation criteria and how to design workflows that integrate these perspectives effectively.

Solution

  • Proposed approach: A staged evaluation workflow that integrates domain experts, lay users, and LLMs to leverage their complementary strengths for creating and refining evaluation criteria.
  • Novelty:
    1. Empirical analysis of evaluation criteria created by domain experts, lay users, and LLMs across two phases (a priori and a posteriori).
    2. Identification of criteria drift and convergence patterns in the evaluation process.
    3. Design guidelines for a staged evaluation workflow that balances quality, cost, and scalability.
  • Procedure and key techniques:
    • Participants (domain experts and lay users) and LLMs generated evaluation criteria for prompts in nutrition and math pedagogy domains.
    • Criteria were created in two phases: a priori (based on the prompt) and a posteriori (after reviewing outputs).
    • Outputs were generated using three LLMs (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Flash).
    • Criteria were analyzed for similarities, differences, and evolution across phases using qualitative coding and thematic analysis.

Results

  • Concrete findings:
    • Domain experts created detailed, knowledge-informed criteria aligned with professional standards, focusing on long-term outcomes.
    • Lay users emphasized usability, clarity, and presentation, reflecting immediate practical needs.
    • LLMs generated generic criteria closely tied to the prompt or output, lacking depth and contextual awareness.
    • Criteria drift was observed, with domain experts and lay users introducing new criteria in response to gaps or errors in outputs, while LLMs showed limited adaptability.
    • Convergence occurred in the a posteriori phase, with all sources focusing on observable features in the outputs.
  • Advantage over baselines: The staged workflow leverages the complementary strengths of domain experts, lay users, and LLMs, addressing limitations of existing evaluation methods that rely on a single source or phase.
  • Experiments / evaluation:
    • Case study in nutrition and math pedagogy domains with five domain experts and five lay users per domain.
    • Outputs generated by three LLMs for three scenarios per domain.
    • Criteria analyzed across a priori and a posteriori phases using qualitative coding and thematic analysis.
  • Limitations and future work:
    • Small sample size limits generalizability; future studies should include larger participant pools and additional domains.
    • Practical implementation of the proposed workflow remains untested.
    • Need to evaluate the reliability of criteria in LLM-as-a-Judge systems and explore generalization across related prompts.

Summary

This study proposes a staged evaluation workflow for LLMs that integrates domain experts, lay users, and LLMs to create and refine evaluation criteria. Domain experts contribute knowledge-based criteria aligned with professional standards, lay users emphasize usability and clarity, and LLMs provide generic baseline criteria. The study identifies criteria drift and convergence patterns, highlighting the importance of involving human judgment early in the process to ensure quality and reduce risks of reinforcing flawed outputs. Design guidelines are provided to balance cost, scalability, and quality in evaluation workflows, with implications for improving LLM evaluation systems in real-world applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223093/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790897
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Data Scientists & Analysts
article
Content Status
Full text indexed
hub
Related Papers
10 related papers