iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision
Authors
Human-LLM CollaborationAI-Assisted Writing & Text GenerationParticipatory DesignPrototyping & User TestingUniversity Professors & ResearchersHCI ResearchersSoftware Engineers & Developers
Paper Title
iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision
Publication Info
- Topic area: Human-AI collaboration for writing evaluation and rubric creation.
- Keywords: Large Language Models, rubric-based feedback, intelligible AI, writing revision, user-defined criteria, explainable AI, scaffolded feedback, meta-rubric, human-AI co-evaluation, skill transfer.
Background and Problem
- Problem / challenge: LLM-generated feedback is often unintelligible, generic, and lacks actionable guidance tailored to user-defined criteria. Rubrics, while effective, are difficult to create and qualify without expertise.
- Significance: Structured, intelligible feedback is crucial for improving writing quality and supporting user-defined criteria in educational and creative contexts.
- Motivation and related work: Prior systems like LLM-Rubric and Prometheus align AI scoring with rubrics but fail to provide criterion-specific explanations or actionable revision guidance. Rubric co-creation has shown promise but lacks systematic qualification and refinement mechanisms.
Solution
- Proposed approach: iRULER—an interactive system for intelligible rubric-based, user-defined LLM evaluation for writing revision and rubric creation.
- Novelty:
- Derivation of six design guidelines for user-defined feedback: Specific, Scaffolded, Justified, Actionable, Qualified, and Refinable.
- Dual-level feedback workflow combining writing revision and rubric creation, recursively applying feedback principles to both artifacts and criteria.
- Integration of meta-rubric evaluation to qualify and refine user-defined rubrics.
- Procedure and key techniques:
- Writing Revision: Users interact with rubrics to receive scores, explanations (Why, Why Not), and actionable revision suggestions (How To).
- Rubric Creation: Users design rubrics and receive meta-rubric feedback for iterative improvement.
- Recursive workflow: Rubrics are refined based on their application to writing artifacts, ensuring alignment with user-defined goals.
Results
- Concrete findings:
- Writing quality improvement: iRULER significantly increased scores (Δ Score: Text-LLM = 16.1, Rubric-LLM = 23.8, iRULER = 30.8).
- Rubric quality improvement: iRULER scored highest (Text-LLM = 76.8, Rubric-LLM = 79.7, iRULER = 86.3).
- Skill transfer: iRULER led to greater unaided revision improvement (Δ Score Post-Pre: Text-LLM = 4.17, Rubric-LLM = 6.33, iRULER = 9.73).
- Advantage over baselines:
- iRULER outperformed Text-LLM and Rubric-LLM in writing quality, rubric quality, perceived helpfulness, correctness, and user confidence.
- Efficiency: iRULER required fewer iterations (Text-LLM = 3.38, Rubric-LLM = 2.69, iRULER = 2.09).
- Experiments / evaluation:
- Writing Revision (N = 48): Tested feedback impact on essay improvement.
- Rubric Creation (N = 36): Evaluated rubric design quality.
- End-to-End (N = 6): Explored recursive rubric refinement and application.
- Expert validation (N = 2): Confirmed LLM scoring alignment with human judgments (α ≥ 0.80, QWK = 0.88 for Rubric-LLM).
- Limitations and future work:
- Limited quantitative data on recursive workflows; future studies should validate end-to-end processes at scale.
- Long-term skill retention remains unexplored; longitudinal studies are needed.
- Generalizability to subjective and multi-modal domains requires further investigation.
Summary
iRULER introduces a novel framework for intelligible, rubric-based feedback in writing revision and rubric creation, guided by six design principles. Empirical results demonstrate its superiority over baseline systems in improving artifact quality, user confidence, and skill transfer. By enabling recursive refinement of user-defined criteria, iRULER supports personalized, structured human-AI collaboration. Future work should explore its applicability to broader domains and assess long-term impacts on learning outcomes.
Research Questions / Practical Problems
Question signals indexed for this paper.
No question or practical-problem data is indexed for this paper yet.
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790539
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Writing & Text Generation, Participatory Design, Prototyping & User Testing
work
Professions
University Professors & Researchers, HCI Researchers, Software Engineers & Developers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers