Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
Authors
The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts. Previous research indicates that LLM-as-a-judge exhibits a strong correlation with human judges in the context of general instruction following. However, for instructions that require specialized knowledge, the validity of using LLMs as judges remains uncertain. In our study, we applied a mixed-methods approach, conducting pairwise comparisons in which both subject matter experts (SMEs) and LLMs evaluated outputs from domain-specific tasks. We focused on two distinct fields: dietetics, with registered dietitian experts, and mental health, with clinical psychologist experts. Our results showed that SMEs agreed with LLM judges 68% of the time in the dietetics domain and 64% in mental health when evaluating overall preference. Additionally, the results indicated variations in SME-LLM agreement across domain-specific aspect questions. Our findings emphasize the importance of keeping human experts in the evaluation process, as LLMs alone may not provide the depth of understanding required for complex, knowledge specific tasks. We also explore the implications of LLM evaluations across different domains and discuss how these insights can inform the design of evaluation workflows that ensure better alignment between human experts and LLMs in interactive systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How consistent are LLM evaluation results with domain expert judgments on complex tasks?Category: Medical Simulation Training and Communication SupportSimilar questionsarrow_forward
- What key factors cause discrepancies between LLM and domain expert evaluation results?Category: Medical Simulation Training and Communication SupportSimilar questionsarrow_forward
- Can simulating an "expert persona" improve LLM consistency in domain-specific expert evaluation?Category: Medical Simulation Training and Communication SupportSimilar questionsarrow_forward
Practical Problems
1- Expert evaluation in fields such as healthcare and psychological support is costly and difficult to maintain consistently.Category: Medical Simulation Training and Communication SupportSimilar questionsarrow_forward
- 67%
Exploring Customizable Interactive Tools for Therapeutic Homework Support in Mental Health Counseling
CHI '26· Human-LLM Collaboration +2
- 67%
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
CHI '26· Human-LLM Collaboration +2
- 67%
Digitizing the Pre-consultation Experience: Impacts and Design Recommendations
CHI '26· Human-LLM Collaboration +2
- 67%
Towards Better Health Conversations: The Benefits of Context-seeking
CHI '26· Human-LLM Collaboration +2
- 67%
Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-Making
CHI '26· Human-LLM Collaboration +2
- 60%
Effects of Communication Directionality and AI Agent Differences in Human-AI Interaction
CHI '21· Human-LLM Collaboration +1
- 60%
AI Knowledge: Improving AI Delegation through Human Enablement
CHI '23· Human-LLM Collaboration +1
- 60%
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
CHI '23· Human-LLM Collaboration +1
- 60%
How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz Study
CHI '24· Human-LLM Collaboration +1
- 60%
VAL: Interactive Task Learning with GPT Dialog Parsing
CHI '24· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)