An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
Authors
Paper Title
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
Publication Info
- Topic area: Evaluation of LLM errors in scholarly question-answering tasks.
- Keywords: Large Language Models, scholarly QA, error taxonomy, expert evaluation, hallucinations, synthesis issues, domain-specific evaluation, retrieval-augmented generation, contextual inquiry, human-AI collaboration.
Background and Problem
- Problem / challenge: Current automated evaluation metrics for LLMs lack the contextual nuance required for scholarly QA tasks and fail to reflect the assessment strategies used by domain experts.
- Significance: Reliable scholarly QA systems are critical for managing the growing volume of scientific literature and supporting researchers in navigating complex, domain-specific information.
- Motivation and related work: Prior work has focused on automated metrics and benchmarks, which often fail to capture domain-specific errors, synthesis failures, and hallucinations. Human evaluation remains essential, but structured frameworks for expert-driven assessments are lacking. This paper addresses this gap by developing a schema for evaluating LLM errors based on domain expert feedback.
Solution
- Proposed approach: Development and validation of an expert-derived schema for evaluating LLM errors in scholarly QA, encompassing 20 error patterns across seven categories.
- Novelty:
- Creation of a fine-grained error schema tailored to scholarly QA tasks.
- Use of a two-phase methodology involving domain experts to identify and validate error patterns.
- Analysis of expert evaluation strategies and their implications for hybrid evaluation approaches.
- Mapping of error patterns to question types, highlighting specific challenges in scholarly QA.
- Procedure and key techniques:
- Phase 1: Schema development with three domain experts and researchers, analyzing 68 QA pairs from a retrieval-augmented generation (RAG) system.
- Phase 2: Schema validation with 10 additional domain experts using contextual inquiries and structured inventories to evaluate 120 QA pairs.
- Identification of seven top-level error categories: Incorrect Answers, Hallucinations, Incomplete Answers, Question Interpretation, Synthesis Issues, Formatting Issues, and System Failures.
- Mapping 11 question types (e.g., methodological inquiry, critical evaluation) to error categories to understand error distribution.
Results
- Concrete findings:
- Identified 20 error patterns across seven categories, including hallucinations (e.g., fabricated citations), synthesis failures, and incomplete answers.
- Experts naturally identified errors in correctness and completeness but overlooked subtle hallucinations and citation errors without structured guidance.
- Higher-order questions (e.g., methodological inquiry) surfaced more errors than factual verification questions.
- Advantage over baselines:
- The schema enables systematic identification of errors that automated metrics and benchmarks fail to capture, such as domain-specific synthesis failures and nuanced hallucinations.
- Structured evaluation tools based on the schema improved experts’ ability to detect subtle errors.
- Experiments / evaluation:
- Two-phase study with 13 domain experts evaluating 188 QA pairs.
- Use of a retrieval-augmented generation system to generate answers based on scholarly papers.
- Analysis of expert feedback and question characteristics to refine the schema and identify patterns in error detection.
- Limitations and future work:
- Small sample size (N=10 for validation) and focus on STEM fields; schema may require adaptation for humanities and social sciences.
- Technical limitations of the RAG system used, which may not generalize to larger proprietary models.
- Future work includes testing schema generalizability across disciplines, automating error detection, and developing hybrid evaluation workflows.
Summary
This paper introduces a schema for evaluating LLM errors in scholarly QA, developed through collaboration with domain experts. The schema identifies 20 error patterns across seven categories, addressing limitations of automated metrics by capturing domain-specific issues such as hallucinations and synthesis failures. Validation with 10 experts showed that structured evaluation tools based on the schema improved error detection, especially for subtle issues. The findings highlight the need for hybrid evaluation approaches combining automated pre-screening with expert oversight, as well as personalized tools to support domain-specific error identification. The schema provides a foundation for improving LLM reliability in scholarly contexts and adapting evaluation frameworks to diverse academic disciplines.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Criticality: Scaffolding Decision-Making with Interactive Critical Thinking and Evidence-Based Reasoning Traces
IUI '26· Human-LLM Collaboration +3
- 71%
Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
CHI '26· Human-LLM Collaboration +2
- 71%
Investigating the Effects of LLM Use on Critical Thinking Under Time Constraints: Access Timing and Time Availability
CHI '26· Human-LLM Collaboration +2
- 71%
LLM or Human? Perceptions of Trust and Quality in Research Summaries
CHI '26· Human-LLM Collaboration +2
- 71%
Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
IUI '26· Human-LLM Collaboration +2
- 71%
A Multimodal Investigation of Controllability and Cognitive Load in Interactive Machine Learning
IUI '26· Human-LLM Collaboration +2
- 71%
Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
IUI '26· Human-LLM Collaboration +2
- 67%
Cody: An AI-Based System to Semi-Automate Coding for Qualitative Research
CHI '21· Human-LLM Collaboration +1
- 67%
User Modelling for Avoiding Overfitting in Interactive Knowledge Elicitation for Prediction
IUI '18· Human-LLM Collaboration +1
- 67%
ReviewFlow: Intelligent Scaffolding to Support Academic Peer Reviewing
IUI '24· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)