Evalet: Evaluating Large Language Models through Functional Fragmentation
Honorable MentionAuthors
Paper Title
Evalet: Evaluating Large Language Models through Functional Fragmentation
Publication Info
- Topic area: Evaluation methods for Large Language Models (LLMs)
- Keywords: Large Language Models, evaluation, functional fragmentation, LLM-as-a-Judge, model behavior, qualitative analysis, fragment-level functions, interactive systems, user study, visualization
Background and Problem
- Problem / challenge: Current LLM evaluation methods rely on holistic scores that obscure specific elements influencing assessments, making it difficult to identify systemic issues or validate evaluations at scale.
- Significance: Addressing this limitation is critical for ensuring safe deployment of LLMs, aligning outputs with user goals, and enabling practitioners to trust and act on evaluations effectively.
- Motivation and related work: Prior approaches provide numeric scores and justifications but require manual review to map evaluations to specific fragments, limiting scalability. Existing fine-grained methods focus on predefined error categories or text spans but lack emergent annotation capabilities.
Solution
- Proposed approach: Functional fragmentation, a method that decomposes LLM outputs into key fragments, interprets their rhetorical functions relative to evaluation criteria, and visualizes these functions for inspection, rating, and comparison.
- Novelty:
- Disentangling outputs into fragment-level functions relevant to criteria.
- Supporting emergent annotation and clustering of functions across outputs.
- Enabling interactive exploration and verification of evaluations through Evalet.
- Procedure and key techniques:
- Extract criterion-relevant fragments using LLM prompts.
- Label fragments with concise function descriptions and rate their alignment.
- Summarize evaluations into holistic justifications and scores based on fragment ratings.
- Cluster functions into base and super clusters using embedding and hierarchical clustering techniques.
- Visualize functions and clusters to support granular and scalable analysis.
Results
- Concrete findings:
- Functional fragmentation achieved higher recall (90%) in fragment extraction compared to holistic methods (84%).
- Accuracy in identifying higher-quality outputs was improved (80.1% vs. 75.5%).
- User study participants identified 48% more evaluation misalignments using Evalet.
- Advantage over baselines:
- More nuanced evaluations with emergent function annotations.
- Improved interpretability and trust calibration in LLM evaluations.
- Enhanced ability to identify actionable issues in outputs.
- Experiments / evaluation:
- Technical evaluation on Scarecrow and RewardBench datasets using IoU, precision, recall, and accuracy metrics.
- User study (N=10) comparing Evalet against holistic evaluation methods across two tasks: horror story generation and advertisement post creation.
- Limitations and future work:
- Does not account for relative importance of functions.
- Limited testing on larger datasets and real-world workflows.
- Need for additional controls (e.g., fragment size editing) and holistic attributes integration.
Summary
Evalet introduces functional fragmentation to evaluate LLM outputs by dissecting them into fragment-level functions and visualizing their alignment with user-defined criteria. This approach improves interpretability, trust calibration, and identification of actionable issues compared to holistic methods. Technical and user evaluations demonstrate its effectiveness in extracting relevant fragments, assessing overall quality, and supporting nuanced analysis across diverse tasks. Future work aims to refine function prioritization, expand scalability, and integrate fragmented and holistic evaluations into real-world practices.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
From Narrative to Numbers: Evaluating Survey Questionnaires with Large Language Models
IUI '26· Human-LLM Collaboration +2
- 86%
Rationalizer: Leveraging LLM to Support User Providing the Rationales Behind the Rating of Likert Scale Questionnaires
IUI '26· Human-LLM Collaboration +2
- 75%
SemTabla: A Human-in-the-Loop Framework for Semantic Enrichment and Validation of Data Tables
CHI '26· Explainable AI (XAI) +3
- 75%
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
CHI '26· Human-LLM Collaboration +3
- 75%
Integrating Complementary Feature Sets for Human-AI Decision-Making
IUI '26· Human-LLM Collaboration +3
- 75%
Criticality: Scaffolding Decision-Making with Interactive Critical Thinking and Evidence-Based Reasoning Traces
IUI '26· Human-LLM Collaboration +3
- 71%
AI of Oz: Enhancing Wizard of Oz Studies in HCI with AI Assistance for Human Moderation
CHI '26· Human-LLM Collaboration +2
- 67%
LAPS: Automating Hypothesis-Driven Statistical Analysis of Public Survey Using Large Language Models
CHI '26· Human-LLM Collaboration +4
- 67%
Mapping the Design Space of User Experience for Computer Use Agents
IUI '26· Human-LLM Collaboration +4
- 63%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)