LLM or Human? Perceptions of Trust and Quality in Research Summaries
Authors
Paper Title
LLM or Human? Perceptions of Trust and Quality in Research Summaries
Publication Info
- Topic area: Trust and quality perceptions of LLM-generated versus human-written academic abstracts.
- Keywords: Large Language Models, academic writing, trust, quality evaluation, disclosure, human-AI collaboration, scientific communication, abstract generation, reader perceptions, AI-assisted writing.
Background and Problem
- Problem / challenge: Readers struggle to distinguish between LLM-generated and human-written abstracts, and the impact of LLM involvement on trust and quality perceptions is unclear. There is also a lack of consensus on disclosure norms for LLM use in academic writing.
- Significance: Understanding how LLM-generated content is perceived is critical for guiding disclosure policies, fostering trust in scientific communication, and ensuring equitable evaluation of AI-assisted writing.
- Motivation and related work: Prior studies show mixed effects of AI disclosure on trust and quality judgments, with readers often misclassifying AI-generated content. However, systematic evidence on how LLM involvement and disclosure affect perceptions in academic contexts remains limited.
Solution
- Proposed approach: A mixed-methods survey experiment to evaluate perceptions of human-written, LLM-generated, and LLM-edited abstracts, with and without disclosure of LLM involvement.
- Novelty:
- Empirical evaluation of readers’ ability to detect LLM involvement in abstracts.
- Analysis of how actual and perceived LLM involvement shapes trust and quality judgments.
- Identification of three distinct reader orientations toward LLM-assisted writing.
- Insights into the role of disclosure in influencing perceptions of AI-assisted content.
- Procedure and key techniques:
- Participants (N=69, ML experts) evaluated abstracts under two conditions: with disclosed authorship (information condition) and without (guess condition).
- Abstracts were categorized as human-written, LLM-generated, or LLM-edited.
- Participants rated abstracts on trust, clarity, comprehensiveness, engagement, and conciseness using Likert scales.
- Behavioral trust was assessed through abstract selection tasks.
- Qualitative analysis of participants’ reasoning and orientations toward LLM-generated content.
Results
- Concrete findings:
- Participants could not reliably distinguish between human-written, LLM-generated, and LLM-edited abstracts (classification accuracy near chance).
- LLM-edited abstracts were rated highest for clarity and trustworthiness, while LLM-generated abstracts received the lowest ratings.
- Disclosure of LLM involvement increased trust ratings across all abstract types.
- Advantage over baselines:
- LLM-edited abstracts were preferred by 55% of participants when authorship was disclosed, compared to 27–28% for human-written and LLM-generated abstracts.
- Disclosure mitigated suspicion and allowed participants to focus on content quality rather than authorship speculation.
- Experiments / evaluation:
- Participants evaluated 150 abstracts (50 per type) from ML papers on arXiv.
- Measures included Likert-scale ratings, behavioral trust (abstract selection), and qualitative reasoning analysis.
- Statistical models controlled for participant and task-level effects.
- Limitations and future work:
- Limited generalizability to non-expert audiences and other academic domains.
- Focused on a single LLM (Llama 3.1 8B) and fixed prompts.
- Future work should explore diverse populations, multiple LLMs, longitudinal effects, and real-world behaviors like citation choices.
Summary
This study investigates how readers perceive trust and quality in human-written, LLM-generated, and LLM-edited abstracts. Participants could not reliably identify LLM involvement but preferred LLM-edited abstracts for their clarity and trustworthiness, especially when authorship was disclosed. Disclosure of LLM use increased trust across all abstract types, challenging assumptions of an AI disclosure penalty. The study identifies three reader orientations—Disclosure Advocates, Pragmatic Skeptics, and Optimists—highlighting diverse attitudes toward LLM-assisted writing. These findings inform disclosure policies and the design of AI tools to enhance trust and quality in scientific communication.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
IUI '26· Human-LLM Collaboration +2
- 71%
Authorship Drift: How Self-Efficacy and Trust Evolve During LLM-Assisted Writing
CHI '26· Human-LLM Collaboration +2
- 71%
Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
CHI '26· Human-LLM Collaboration +2
- 71%
DraftMarks: Enhancing Transparency in Human-AI Co-Writing Through Interactive Skeuomorphic Process Traces
CHI '26· Human-LLM Collaboration +2
- 71%
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
CHI '26· Human-LLM Collaboration +2
- 71%
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
CHI '26· Human-LLM Collaboration +2
- 71%
A Multimodal Investigation of Controllability and Cognitive Load in Interactive Machine Learning
IUI '26· Human-LLM Collaboration +2
- 71%
Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
IUI '26· Human-LLM Collaboration +2
- 67%
User Modelling for Avoiding Overfitting in Interactive Knowledge Elicitation for Prediction
IUI '18· Human-LLM Collaboration +1
- 67%
DxHF: Providing High-Quality Human Feedback for LLM Alignment with Interactive Decomposition
UIST '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)