PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
Authors
Paper Title
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
Publication Info
- Topic area: Enhancing trust and verification in scholarly question-answering systems using large language models (LLMs).
- Keywords: LLMs, scholarly QA, provenance, claim-evidence matching, trust calibration, argumentation structures, user study, cognitive load, usability, verification.
Background and Problem
- Problem / challenge: LLM-based scholarly QA systems often produce errors like unsupported claims and omissions. Existing provenance mechanisms, such as source citations, lack granularity, making rigorous verification difficult.
- Significance: Errors in LLM outputs can propagate misinformation in high-stakes academic contexts, undermining trust and reliability in scholarly workflows.
- Motivation and related work: Prior work has focused on coarse-grained attribution and explanations, but these approaches often fail to foster appropriate trust or actionable verification. This paper addresses the gap by introducing a system that aligns with the argumentation structures of scholarly discourse.
Solution
- Proposed approach: PaperTrail, a system that decomposes LLM answers and source documents into claims and evidence, mapping them for granular provenance in scholarly QA.
- Novelty:
- Design and implementation of a claim-evidence interface for scholarly QA.
- A backend architecture combining LLM-based, similarity-based, and retrieval-augmented generation (RAG) methods for claim-evidence extraction.
- Empirical evidence showing the impact of claim-evidence provenance on trust calibration.
- Identification of a trust-behavior gap in LLM reliance under time and cognitive constraints.
- Procedure and key techniques:
- Backend: A three-stage pipeline for claim and evidence extraction:
- Offline paper-level claim-evidence extraction using LLM-based methods.
- Real-time answer-level claim-evidence extraction from LLM-generated responses.
- Real-time claim-evidence matching using RAG for provenance indicators.
- Frontend: A three-panel interface with features like claim coverage indicators, interactive claim cards, and coordinated views to support verification workflows.
- Backend: A three-stage pipeline for claim and evidence extraction:
Results
- Concrete findings:
- Trust in LLM outputs was significantly lower with PaperTrail (mean = 3.68) compared to the baseline (mean = 4.22, p = 0.015).
- No significant difference in reliance on LLM-generated text (PaperTrail: 0.75, Baseline: 0.73, p = 0.313).
- Confidence in task outputs was similar across conditions (PaperTrail: 4.05, Baseline: 4.27, p = 0.525).
- Advantage over baselines:
- PaperTrail provided granular claim-evidence provenance, encouraging caution in LLM trust, unlike the baseline citation-based interface.
- Experiments / evaluation:
- A within-subjects study with 26 researchers performing two scholarly editing tasks.
- Metrics: trust (TXAI scale), reliance (Levenshtein edit distance), confidence, usability (SUPR-Q), and cognitive load (NASA-TLX).
- Tasks: multi-paper synthesis and devil’s advocate review using a Mars exploration corpus.
- Limitations and future work:
- High system latency and interface complexity reduced usability.
- Short task durations limited deep engagement with provenance features.
- Future work should explore adaptive provenance, longitudinal studies, and task-specific verification workflows.
Summary
PaperTrail introduces a novel claim-evidence interface for LLM-based scholarly QA, enabling granular provenance by matching LLM-generated claims with source document evidence. A user study showed that PaperTrail reduced trust in LLM outputs but did not significantly change reliance behaviors, highlighting a trust-behavior gap under time and cognitive constraints. While participants valued the concept of detailed provenance, usability challenges limited its practical impact. The system’s backend architecture also demonstrates potential as a framework for evaluating LLM trustworthiness in scholarly contexts. Future work should focus on improving usability, adaptive provenance, and real-world deployment scenarios.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 75%
From Toil to Thought: Designing for Strategic Exploration and Responsible AI in Systematic Literature Reviews
IUI '26· Explainable AI (XAI) +3
- 75%
Criticality: Scaffolding Decision-Making with Interactive Critical Thinking and Evidence-Based Reasoning Traces
IUI '26· Human-LLM Collaboration +3
- 71%
LLM-based In-situ Thought Exchanges for Critical Paper Reading
IUI '26· Human-LLM Collaboration +2
- 71%
DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-Exploration
UIST '24· Human-LLM Collaboration +2
- 67%
LAPS: Automating Hypothesis-Driven Statistical Analysis of Public Survey Using Large Language Models
CHI '26· Human-LLM Collaboration +4
- 63%
CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language Models
CHI '24· Human-LLM Collaboration +2
- 63%
InterFlow: Designing Unobtrusive AI to Empower Interviewers in Semi-Structured Interviews
CHI '26· Human-LLM Collaboration +3
- 63%
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
CHI '26· Human-LLM Collaboration +2
- 63%
Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
IUI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)