RELIC: Investigating Large Language Model Responses using Self-Consistency
Authors
Document Title
RELIC: Investigating Large Language Model Responses using Self-Consistency
Document Information
- Subject Area: Human-Computer Interaction, Natural Language Processing (NLP), Validation of Large Language Model (LLM) Outputs
- Keywords: Natural Language Generation, Human-Computer Interaction, Hallucination Detection, Self-Consistency, Visualization, User Study, Information Verification, AI Trust
Research Background and Problem
-
Problems and Challenges:
- Large Language Models (LLMs) often generate fluent text while confusing fact with fiction, a phenomenon known as "hallucination."
- These inaccurate outputs are often presented in a realistic and convincing manner, which can mislead users and lead to legal and ethical issues.
- Current confidence indicators (e.g., word-level probabilities) fail to provide semantic-level trustworthiness, especially for validating long-form text generation.
-
Research Significance:
- Preventing users from being misled by models is a critical issue in the field of human-computer interaction, particularly as LLMs are widely applied in tasks like writing assistance and search engines.
- Designing tools that help users understand and verify generated content is essential for fostering trustworthy human-AI collaboration.
-
Research Motivation and Related Work:
- Current methods include retrieval-based and consistency-based approaches, but most focus on automated validation rather than empowering end-users.
- Self-consistency has been proven effective for evaluating the consistency of multiple model outputs, indicating their potential reliability.
- This study uniquely aims to explore a user-oriented framework that combines self-consistency with interaction design to help users verify generated text through semantic-level analysis.
Solution
-
Proposed Method:
- Introduced an interactive system called RELIC (Reliable Evaluation of Language Information Consistency) to explore and verify the semantic consistency of LLM outputs.
- Designed a new algorithm based on self-consistency to evaluate semantic consistency and trustworthiness by analyzing diverse samples generated from the same prompt.
-
Innovations:
- Extended self-consistency from a single score to complex long-form text by extracting fine-grained semantic atomic claims for trustworthiness interpretation.
- Integrated Natural Language Inference (NLI), question generation, and answer clustering, introducing interactive operations (e.g., annotation, editing, and brushing) to support user verification of generated content.
-
Implementation Steps:
- Algorithm Design:
- Generate multiple samples and extract semantic atomic claims.
- Use Natural Language Inference (NLI) models to evaluate logical entailment relationships among generated atomic claims.
- Transform claims into natural language questions and retrieve answers from other samples.
- Summarize alternative answers through semantic clustering and generate supporting, refuting, and neutral options.
- Interactive System Design:
- Keyword Annotation: Visualize the consistency distribution of different options using bar charts.
- Brushing for Questioning: Users can select text to generate questions for collecting alternative evidence.
- Editing the Generation: Users can edit the text and view real-time changes in consistency evaluation.
- Algorithm Design:
Research Outcomes
-
Specific Outcomes:
- Developed the RELIC system to support users in verifying and refining LLM-generated content.
- Users can easily interact through the Claims View and Evidence View to explore alternative answers and supporting evidence for generated content.
-
Comparison with Existing Solutions:
- Compared to traditional model interfaces (e.g., ChatGPT UI), the RELIC system significantly improves users' ability to access diverse answers and verify generated content.
- Users can engage in structured interactions rather than relying solely on single-turn Q&A.
-
Experiments and Evaluation:
- User Study:
- 10 participants tested RELIC's functionality through task operations and post-task surveys to assess usability.
- Participants generally found RELIC easy to use (5.7/7), helpful for verification (6.4/7), and effective for improving generated text (5.5/7).
- Algorithm Quantitative Analysis:
- Experiments showed that self-consistency scoring significantly aids in identifying inaccurate generations, achieving an AUROC of 0.856 for the InstructGPT model.
- Increasing the number of samples improves consistency evaluation performance, though with diminishing marginal returns.
- User Study:
-
Limitations and Future Directions:
- Limitations:
- Consistency does not always equate to correctness; some errors may still exhibit high consistency.
- The system's integration with external trusted resources is limited, and additional error correction needs remain unresolved.
- Future Directions:
- Enhance the algorithm by incorporating evidence from external knowledge bases or other model generations to improve system reliability.
- Develop intelligent text editing features to further automate the optimization of generated content.
- Limitations:
Conclusion
This study proposed the RELIC system, which combines a self-consistency algorithm with user interaction design to provide a novel method for verifying the authenticity of LLM outputs. User studies and algorithm evaluations demonstrated the system's usability and practicality. In the future, the system could be extended to various domains, including knowledge retrieval, predictive tasks, and creative generation.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How does self-consistency (evaluating consistency through diverse examples) improve semantic-level trustworthiness verification of LLM outputs?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- How can interactive systems help users verify semantic consistency and authenticity of long-text generation?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- How does the self-consistency mechanism perform during users' verification of generated content?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
Practical Problems
1- Users struggle to judge the authenticity of AI-generated text and are easily misled.Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- 100%
Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
CHI '26· Explainable AI (XAI) +2
- 100%
Transferable XAI: Relating Understanding Across Domains with Explanation Transfer
IUI '26· Explainable AI (XAI) +2
- 86%
Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI
CHI '26· Explainable AI (XAI) +3
- 86%
From Reflection to Repair: A Scoping Review of Dataset Documentation Tools
CHI '26· Explainable AI (XAI) +3
- 83%
Explanations as Mechanisms for Supporting Algorithmic Transparency
CHI '18· Explainable AI (XAI) +1
- 83%
Trends and Trajectories for Explainable, Accountable and Intelligible Systems: An HCI Research Agenda
CHI '18· Explainable AI (XAI) +2
- 71%
Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models
CHI '19· Explainable AI (XAI) +2
- 71%
Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine Learning
CHI '20· Explainable AI (XAI) +2
- 71%
Researching AI Legibility through Design
CHI '20· Explainable AI (XAI) +2
- 71%
No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive ML
CHI '20· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)