Evalet: Evaluating Large Language Models through Functional Fragmentation

Honorable Mention
Human-LLM CollaborationExplainable AI (XAI)User Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingAI/ML Researchers & EngineersHCI ResearchersData Scientists & Analysts

Paper Title

Evalet: Evaluating Large Language Models through Functional Fragmentation

Publication Info

  • Topic area: Evaluation methods for Large Language Models (LLMs)
  • Keywords: Large Language Models, evaluation, functional fragmentation, LLM-as-a-Judge, model behavior, qualitative analysis, fragment-level functions, interactive systems, user study, visualization

Background and Problem

  • Problem / challenge: Current LLM evaluation methods rely on holistic scores that obscure specific elements influencing assessments, making it difficult to identify systemic issues or validate evaluations at scale.
  • Significance: Addressing this limitation is critical for ensuring safe deployment of LLMs, aligning outputs with user goals, and enabling practitioners to trust and act on evaluations effectively.
  • Motivation and related work: Prior approaches provide numeric scores and justifications but require manual review to map evaluations to specific fragments, limiting scalability. Existing fine-grained methods focus on predefined error categories or text spans but lack emergent annotation capabilities.

Solution

  • Proposed approach: Functional fragmentation, a method that decomposes LLM outputs into key fragments, interprets their rhetorical functions relative to evaluation criteria, and visualizes these functions for inspection, rating, and comparison.
  • Novelty:
    1. Disentangling outputs into fragment-level functions relevant to criteria.
    2. Supporting emergent annotation and clustering of functions across outputs.
    3. Enabling interactive exploration and verification of evaluations through Evalet.
  • Procedure and key techniques:
    • Extract criterion-relevant fragments using LLM prompts.
    • Label fragments with concise function descriptions and rate their alignment.
    • Summarize evaluations into holistic justifications and scores based on fragment ratings.
    • Cluster functions into base and super clusters using embedding and hierarchical clustering techniques.
    • Visualize functions and clusters to support granular and scalable analysis.

Results

  • Concrete findings:
    • Functional fragmentation achieved higher recall (90%) in fragment extraction compared to holistic methods (84%).
    • Accuracy in identifying higher-quality outputs was improved (80.1% vs. 75.5%).
    • User study participants identified 48% more evaluation misalignments using Evalet.
  • Advantage over baselines:
    • More nuanced evaluations with emergent function annotations.
    • Improved interpretability and trust calibration in LLM evaluations.
    • Enhanced ability to identify actionable issues in outputs.
  • Experiments / evaluation:
    • Technical evaluation on Scarecrow and RewardBench datasets using IoU, precision, recall, and accuracy metrics.
    • User study (N=10) comparing Evalet against holistic evaluation methods across two tasks: horror story generation and advertisement post creation.
  • Limitations and future work:
    • Does not account for relative importance of functions.
    • Limited testing on larger datasets and real-world workflows.
    • Need for additional controls (e.g., fragment size editing) and holistic attributes integration.

Summary

Evalet introduces functional fragmentation to evaluate LLM outputs by dissecting them into fragment-level functions and visualizing their alignment with user-defined criteria. This approach improves interpretability, trust calibration, and identification of actionable issues compared to holistic methods. Technical and user evaluations demonstrate its effectiveness in extracting relevant fragments, assessing overall quality, and supporting nuanced analysis across diverse tasks. Future work aims to refine function prioritization, expand scalability, and integrate fragmented and holistic evaluations into real-world practices.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222991/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790285
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Data Scientists & Analysts
article
Content Status
Full text indexed
hub
Related Papers
10 related papers