Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
Authors
Paper Title
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
Publication Info
- Topic area: Evaluation workflows for Large Language Models (LLMs) in domain-specific tasks.
- Keywords: Large Language Models, evaluation criteria, domain experts, lay users, human-AI collaboration, staged workflows, criteria drift, usability testing, natural language processing, LLM-as-a-Judge.
Background and Problem
- Problem / challenge: Existing evaluation methods for LLMs often fail to account for the unique contributions of diverse sources (domain experts, lay users, and LLMs) and how evaluation criteria evolve across different stages of the evaluation process. Current approaches are limited by cost, scalability, and alignment with domain-specific standards.
- Significance: Effective evaluation of LLMs is essential to ensure accuracy, reliability, and alignment with end-user needs, particularly in high-stakes domains like healthcare and education, where errors can have significant consequences.
- Motivation and related work: Prior research has focused on tools and systems for LLM evaluation (e.g., EvalLM, EvalAssist) and the role of human-AI collaboration. However, gaps remain in understanding how different sources contribute to evaluation criteria and how to design workflows that integrate these perspectives effectively.
Solution
- Proposed approach: A staged evaluation workflow that integrates domain experts, lay users, and LLMs to leverage their complementary strengths for creating and refining evaluation criteria.
- Novelty:
- Empirical analysis of evaluation criteria created by domain experts, lay users, and LLMs across two phases (a priori and a posteriori).
- Identification of criteria drift and convergence patterns in the evaluation process.
- Design guidelines for a staged evaluation workflow that balances quality, cost, and scalability.
- Procedure and key techniques:
- Participants (domain experts and lay users) and LLMs generated evaluation criteria for prompts in nutrition and math pedagogy domains.
- Criteria were created in two phases: a priori (based on the prompt) and a posteriori (after reviewing outputs).
- Outputs were generated using three LLMs (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Flash).
- Criteria were analyzed for similarities, differences, and evolution across phases using qualitative coding and thematic analysis.
Results
- Concrete findings:
- Domain experts created detailed, knowledge-informed criteria aligned with professional standards, focusing on long-term outcomes.
- Lay users emphasized usability, clarity, and presentation, reflecting immediate practical needs.
- LLMs generated generic criteria closely tied to the prompt or output, lacking depth and contextual awareness.
- Criteria drift was observed, with domain experts and lay users introducing new criteria in response to gaps or errors in outputs, while LLMs showed limited adaptability.
- Convergence occurred in the a posteriori phase, with all sources focusing on observable features in the outputs.
- Advantage over baselines: The staged workflow leverages the complementary strengths of domain experts, lay users, and LLMs, addressing limitations of existing evaluation methods that rely on a single source or phase.
- Experiments / evaluation:
- Case study in nutrition and math pedagogy domains with five domain experts and five lay users per domain.
- Outputs generated by three LLMs for three scenarios per domain.
- Criteria analyzed across a priori and a posteriori phases using qualitative coding and thematic analysis.
- Limitations and future work:
- Small sample size limits generalizability; future studies should include larger participant pools and additional domains.
- Practical implementation of the proposed workflow remains untested.
- Need to evaluate the reliability of criteria in LLM-as-a-Judge systems and explore generalization across related prompts.
Summary
This study proposes a staged evaluation workflow for LLMs that integrates domain experts, lay users, and LLMs to create and refine evaluation criteria. Domain experts contribute knowledge-based criteria aligned with professional standards, lay users emphasize usability and clarity, and LLMs provide generic baseline criteria. The study identifies criteria drift and convergence patterns, highlighting the importance of involving human judgment early in the process to ensure quality and reduce risks of reinforcing flawed outputs. Design guidelines are provided to balance cost, scalability, and quality in evaluation workflows, with implications for improving LLM evaluation systems in real-world applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Integrating Complementary Feature Sets for Human-AI Decision-Making
IUI '26· Human-LLM Collaboration +3
- 86%
From Narrative to Numbers: Evaluating Survey Questionnaires with Large Language Models
IUI '26· Human-LLM Collaboration +2
- 86%
Rationalizer: Leveraging LLM to Support User Providing the Rationales Behind the Rating of Likert Scale Questionnaires
IUI '26· Human-LLM Collaboration +2
- 75%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 75%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 75%
Evalet: Evaluating Large Language Models through Functional Fragmentation
CHI '26· Human-LLM Collaboration +3
- 75%
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
CHI '26· Human-LLM Collaboration +3
- 75%
Live in the Loop: Rapid Run-time Feedback for Prompts
CHI '26· Human-LLM Collaboration +3
- 71%
AI of Oz: Enhancing Wizard of Oz Studies in HCI with AI Assistance for Human Moderation
CHI '26· Human-LLM Collaboration +2
- 71%
EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences
UIST '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)