Think Together and Work Better: Combining Humans' and LLMs' Think-Aloud Outcomes for Effective Text Evaluation
Authors
Research Background and Problem
-
What problems or challenges did the authors identify?
- Current text evaluation methods based on large language models (LLMs) rely on task-specific prompts, leading to insufficient coverage and unreliable results.
- Checklist methods based on manually designed criteria depend on subjectively defined key evaluation standards, which may introduce bias and lack granularity.
- Solely relying on humans or LLMs for evaluation fails to fully leverage their respective strengths. For example, humans are prone to fatigue when handling large-scale tasks, while LLMs may struggle to capture complex semantic nuances.
-
Why is this problem important? Accurately evaluating the quality of generated text is a core challenge in the fields of text generation and natural language processing, directly impacting the reliability of applications such as education and news generation. Only by combining the flexibility of humans with the consistency of LLMs can automated evaluation methods be improved.
-
Research Motivation and Related Work
- Related literature has explored the limitations of LLM evaluation capabilities, such as prompt-based text scoring (G-Eval) and checklist-based binary scoring frameworks (CheckEval).
- Previous studies have shown that human-AI collaboration can outperform in various domains, such as creative processes and data annotation. This study specifically focuses on integrating the strengths of both for text evaluation tasks.
Solution
-
What methods or solutions did the authors propose? The authors proposed a framework called InteractEval, which integrates the innovative thinking processes of humans and LLMs (Think-Aloud, TA) to generate scoring checklists for more detailed and reliable text evaluation.
-
What are the innovative aspects of this solution?
- For the first time, the "Think-Aloud" process of humans and LLMs is combined to generate evaluation attributes, enhancing the comprehensiveness and reliability of the results.
- A checklist mechanism based on question generation is proposed, using merged attributes to create specific, fine-grained evaluation criteria.
- The complementary strengths of humans and LLMs in generating text attributes are leveraged to achieve more efficient collaboration.
-
What are the implementation steps and key technologies used?
- Attribute Collection: Human experts and LLMs independently generate text attributes for checklist construction through the TA process.
- Checklist Construction:
- Component Extraction: Key themes related to dimensions are extracted from merged attributes using GPT-4.
- Attribute Clustering: Text attributes are grouped to generate theme-related questions.
- Question Generation: Fine-grained "yes/no" questions are created for each component, forming the final checklist.
- Checklist-Based Evaluation: LLMs answer the checklist questions to generate final scores, which are compared with human ratings (ground truth) in the dataset.
- Experiments and Comparative Analysis: The framework's performance is validated on text summarization datasets (SummEval) and essay evaluation datasets (ELLIPSE).
Research Outcomes
-
What specific results were achieved?
- InteractEval outperformed existing baseline methods in both text summarization (SummEval) and essay evaluation (ELLIPSE) tasks.
- Analysis revealed complementary strengths of humans and LLMs across different dimensions:
- Humans excelled in evaluating internal quality (e.g., coherence, fluency).
- LLMs performed better in external alignment (e.g., consistency and relevance).
-
What advantages does it have compared to existing solutions?
- In terms of performance, InteractEval surpassed baseline methods such as ROUGE-L, BLEU, G-Eval, and CheckEval.
- By combining the strengths of humans and LLMs, it reduces biases from single sources while enhancing attribute diversity and granularity.
- Demonstrates high scalability and adaptability across tasks.
-
What were the experimental or evaluation results?
- On the SummEval dataset, InteractEval's scores showed significantly higher correlations (Spearman’s Rho and Kendall’s Tau) with human ratings compared to existing LLM-based methods.
- Scalability validation on the ELLIPSE dataset demonstrated that InteractEval outperformed all existing baseline methods across multiple writing evaluation dimensions.
-
Limitations and Future Directions
- Limitations:
- The method has only been validated on English datasets and has not yet been extended to multilingual scenarios.
- The performance of smaller-scale LLMs has not been thoroughly explored.
- The applicability to other related tasks (e.g., question answering or image caption generation) remains to be verified.
- Future Directions:
- Explore the applicability of the method to a broader range of text-related tasks.
- Evaluate the capabilities of more lightweight LLMs in text evaluation.
- Further integrate human-AI collaboration in both online and offline tasks, optimizing interaction processes.
- Limitations:
Conclusion
This study proposed and validated the InteractEval framework, introducing a text evaluation method that combines the strengths of humans and large language models. Experimental results demonstrate the significant potential of human-LLM collaboration in text evaluation, where their complementary strengths effectively enhance the accuracy of generated text assessments, advancing the development of automated evaluation technologies.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What limitations do current LLM-based text evaluation methods have?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How can a more comprehensive and reliable text evaluation framework combine human and LLM strengths?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How much improvement can a framework combining human and LLM strengths achieve in summarization and essay evaluation tasks?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
Practical Problems
1- Existing text evaluation systems have limited coverage and unreliable results, affecting education and news generation applications.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 83%
The Metacognitive Demands and Opportunities of Generative AI
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 83%
Exploring the Innovation Opportunities for Pre-trained Models
DIS '25· Generative AI (Text, Image, Music, Video) +2
- 80%
CreAItive Collaboration? Users' Misjudgment of AI-Creativity Affects Their Collaborative Performance
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
Finding the Conversation: A Method for Scoring Documents for Natural Conversation Content
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 80%
Fluid Transformers and Creative Analogies: Exploring Large Language Models' Capacity for Augmenting Cross-Domain Analogical Creativity
C&C '23· Generative AI (Text, Image, Music, Video) +1
- 80%
User Modelling for Avoiding Overfitting in Interactive Knowledge Elicitation for Prediction
IUI '18· Human-LLM Collaboration +1
- 80%
DxHF: Providing High-Quality Human Feedback for LLM Alignment with Interactive Decomposition
UIST '25· Human-LLM Collaboration +1
- 71%
Bridging Gulfs in UI Generation through Semantic Guidance
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 71%
Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 71%
PCGEF: A Framework for Diagnosing Subjective Alignment in Human-Centered Persona-Conditioned Generation
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)