RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
Authors
Retrieval-Augmented Generation (RAG) systems have emerged as a promising solution to enhance large language models (LLMs) by integrating external knowledge retrieval with generative capabilities. While significant advancements have been made in improving retrieval accuracy and response quality, a critical challenge remains that the internal knowledge integration and retrieval-generation interactions in RAG systems are largely opaque. This paper introduces RAGTrace, a system designed to analyze retrieval and generation dynamics in RAG-based systems. Informed by a comprehensive literature review and expert interviews, the system supports a multi-level analysis approach, ranging from high-level performance evaluation to fine-grained examination of retrieval relevance, generation fidelity, and cross-component interactions. Unlike conventional evaluation practices that focus on isolated retrieval or generation quality assessments, RAGTrace enables an integrated exploration of retrieval-generation relationships, allowing users to trace knowledge sources and identify potential failure cases. The system's workflow allows users to build, evaluate, and iterate on retrieval processes tailored to their specific domains of interest. The effectiveness of the system is demonstrated through case studies and expert evaluations on real-world RAG applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
CHI '26· Human-LLM Collaboration +2
- 100%
DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows
CHI '26· Human-LLM Collaboration +2
- 83%
Interactive Debugging and Steering of Multi-Agent AI Systems
CHI '25· Human-LLM Collaboration +2
- 83%
DIY: Helping People Assess the Correctness of Natural Language to SQL Systems
IUI '21· Human-LLM Collaboration +2
- 83%
CoPrompter: User-Centric Evaluation of LM Instruction Alignment for Improved Prompt Engineering
IUI '25· Human-LLM Collaboration +2
- 75%
DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models
CHI '24· Human-LLM Collaboration +3
- 71%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 71%
Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
CHI '24· Human-LLM Collaboration +2
- 71%
Dango: A Mixed-Initiative Data Wrangling System using Large Language Model
CHI '25· Human-LLM Collaboration +2
- 71%
OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question Answering
CHI '25· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)