Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Authors
Due to the cumbersome nature of human evaluation and limitations of code-based evaluation, Large Language Models (LLMs) are increasingly being used to assist humans in evaluating LLM outputs. Yet LLM-generated evaluators simply inherit all the problems of the LLMs they evaluate, requiring further human validation. We present a mixed-initiative approach to “validate the validators”— aligning LLM-generated evaluation functions (be it prompts or code) with human requirements. Our interface, EvalGen, provides automated assistance to users in generating evaluation criteria and implementing assertions. While generating candidate implementations (Python functions, LLM grader prompts), EvalGen asks humans to grade a subset of LLM outputs; this feedback is used to select implementations that better align with user grades. A qualitative study finds overall support for EvalGen but underscores the subjectivity and iterative nature of alignment. In particular, we identify a phenomenon we dub criteria drift: users need criteria to grade outputs, but grading outputs helps users define criteria. What is more, some criteria appear dependent on the specific LLM outputs observed (rather than independent and definable a priori), raising serious questions for approaches that assume the independence of evaluation from observation of model outputs. We present our interface and implementation details, a comparison of our algorithm with a baseline approach, and implications for the design of future LLM evaluation assistants.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
AILA: Attentive Interactive Labeling Assistant for Document Classification through Attention-Based Deep Neural Networks
CHI '19· Human-LLM Collaboration +1
- 80%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 80%
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
CHI '23· Human-LLM Collaboration +1
- 80%
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
CHI '24· Human-LLM Collaboration +1
- 80%
Automatic Macro Mining from Interaction Traces at Scale
CHI '24· Human-LLM Collaboration +1
- 80%
RAG Without the Lag: Enabling "What-If" Analysis for Retrieval-Augmented Generation Pipelines
CHI '26· Human-LLM Collaboration +1
- 75%
MAPLE: Mobile App Prediction Leveraging Large Language Model Embeddings
UbiComp '24· Human-LLM Collaboration
- 75%
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
UIST '24· Human-LLM Collaboration
- 67%
Crystalline: Lowering the Cost for Developers to Collect and Organize Information for Decision Making
CHI '22· Human-LLM Collaboration +2
- 67%
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
CHI '24· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)