PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A

Human-LLM CollaborationExplainable AI (XAI)User Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingUniversity Professors & ResearchersHCI ResearchersStatisticians & Data Scientists

Paper Title

PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A

Publication Info

  • Topic area: Enhancing trust and verification in scholarly question-answering systems using large language models (LLMs).
  • Keywords: LLMs, scholarly QA, provenance, claim-evidence matching, trust calibration, argumentation structures, user study, cognitive load, usability, verification.

Background and Problem

  • Problem / challenge: LLM-based scholarly QA systems often produce errors like unsupported claims and omissions. Existing provenance mechanisms, such as source citations, lack granularity, making rigorous verification difficult.
  • Significance: Errors in LLM outputs can propagate misinformation in high-stakes academic contexts, undermining trust and reliability in scholarly workflows.
  • Motivation and related work: Prior work has focused on coarse-grained attribution and explanations, but these approaches often fail to foster appropriate trust or actionable verification. This paper addresses the gap by introducing a system that aligns with the argumentation structures of scholarly discourse.

Solution

  • Proposed approach: PaperTrail, a system that decomposes LLM answers and source documents into claims and evidence, mapping them for granular provenance in scholarly QA.
  • Novelty:
    1. Design and implementation of a claim-evidence interface for scholarly QA.
    2. A backend architecture combining LLM-based, similarity-based, and retrieval-augmented generation (RAG) methods for claim-evidence extraction.
    3. Empirical evidence showing the impact of claim-evidence provenance on trust calibration.
    4. Identification of a trust-behavior gap in LLM reliance under time and cognitive constraints.
  • Procedure and key techniques:
    • Backend: A three-stage pipeline for claim and evidence extraction:
      1. Offline paper-level claim-evidence extraction using LLM-based methods.
      2. Real-time answer-level claim-evidence extraction from LLM-generated responses.
      3. Real-time claim-evidence matching using RAG for provenance indicators.
    • Frontend: A three-panel interface with features like claim coverage indicators, interactive claim cards, and coordinated views to support verification workflows.

Results

  • Concrete findings:
    • Trust in LLM outputs was significantly lower with PaperTrail (mean = 3.68) compared to the baseline (mean = 4.22, p = 0.015).
    • No significant difference in reliance on LLM-generated text (PaperTrail: 0.75, Baseline: 0.73, p = 0.313).
    • Confidence in task outputs was similar across conditions (PaperTrail: 4.05, Baseline: 4.27, p = 0.525).
  • Advantage over baselines:
    • PaperTrail provided granular claim-evidence provenance, encouraging caution in LLM trust, unlike the baseline citation-based interface.
  • Experiments / evaluation:
    • A within-subjects study with 26 researchers performing two scholarly editing tasks.
    • Metrics: trust (TXAI scale), reliance (Levenshtein edit distance), confidence, usability (SUPR-Q), and cognitive load (NASA-TLX).
    • Tasks: multi-paper synthesis and devil’s advocate review using a Mars exploration corpus.
  • Limitations and future work:
    • High system latency and interface complexity reduced usability.
    • Short task durations limited deep engagement with provenance features.
    • Future work should explore adaptive provenance, longitudinal studies, and task-specific verification workflows.

Summary

PaperTrail introduces a novel claim-evidence interface for LLM-based scholarly QA, enabling granular provenance by matching LLM-generated claims with source document evidence. A user study showed that PaperTrail reduced trust in LLM outputs but did not significantly change reliance behaviors, highlighting a trust-behavior gap under time and cognitive constraints. While participants valued the concept of detailed provenance, usability challenges limited its practical impact. The system’s backend architecture also demonstrates potential as a framework for evaluating LLM trustworthiness in scholarly contexts. Future work should focus on improving usability, adaptive provenance, and real-world deployment scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223500/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791101
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
University Professors & Researchers, HCI Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
9 related papers