Beyond Accuracy: Experts See AI Fact-Checks as Accurate but Less Useful
Authors
Paper Title
Beyond Accuracy: Experts See AI Fact-Checks as Accurate but Less Useful
Publication Info
- Topic area: Evaluation of LLM-generated fact-checking reports by media experts.
- Keywords: Large Language Models, fact-checking, misinformation, media experts, accuracy, usefulness, logic, trust, heuristic-systematic model, AI-generated content.
Background and Problem
- Problem / challenge: While LLMs show promise in automating fact-checking tasks, their ability to generate reliable and useful fact-checking reports remains uncertain. Prior research has focused on classification accuracy but has not extensively explored expert evaluations of LLM-generated reports.
- Significance: Understanding expert perceptions is critical for designing AI tools that can effectively support professionals in combating misinformation, a growing societal issue.
- Motivation and related work: Existing methods for fact-checking rely heavily on human editors, which are costly and limited in scalability. Prior studies have examined LLMs' accuracy in detecting misinformation but have not addressed expert evaluations of their reasoning, usefulness, and trustworthiness. This paper addresses this gap by focusing on media professionals' assessments.
Solution
- Proposed approach: A 2x2 between-subjects online experiment evaluating fact-checking reports generated by humans and LLMs, with or without source disclosure.
- Novelty:
- Focus on expert evaluations rather than layperson assessments.
- Analysis of multiple dimensions (accuracy, usefulness, logic, hallucination, engagement, relevance) of LLM-generated reports.
- Application of the Heuristic-Systematic Model (HSM) to understand expert trust dynamics.
- Integration of qualitative insights to complement quantitative findings.
- Procedure and key techniques:
- Participants (N=274) were randomly assigned to one of four conditions (human/LLM source, disclosed/undisclosed).
- Participants rated eight fact-checking reports on six dimensions using a 7-point Likert scale.
- Reports were generated using GPT-4 Turbo, prompted to mimic PolitiFact's style.
- Post-study questionnaires collected qualitative feedback and demographic data.
- Mixed-effects regression models analyzed quantitative data, while thematic coding analyzed qualitative responses.
Results
- Concrete findings:
- LLM-generated reports were rated as accurate (M=5.45) and logical (M=5.40) as human-written reports (M=5.60 for accuracy, M=5.58 for logic).
- LLM-generated reports were perceived as significantly less useful (M=5.40) than human-written reports (M=5.71, p<.05).
- Party affiliation influenced perceived logicalness, with Republicans and Democrats differing in their trust dynamics.
- Advantage over baselines: LLM-generated reports matched human reports in accuracy and logic but lagged in usefulness, highlighting areas for improvement.
- Experiments / evaluation:
- Dataset: 120 fact-checking reports on health, civics, science, and politics, balanced for truthfulness.
- Metrics: Accuracy, usefulness, logic, hallucination, engagement, relevance.
- Tools: GPT-4 Turbo with zero-shot prompting.
- Limitations and future work:
- Limited generalizability due to U.S.-centric expert sample and static human report conditions.
- Potential priming effects and evolving LLM capabilities (e.g., GPT-5).
- Future work should explore advanced LLM designs, interface transparency, and cross-cultural evaluations.
Summary
This study evaluates how media experts perceive LLM-generated fact-checking reports compared to human-written ones. While LLM reports were rated as accurate and logical as human-authored ones, they were perceived as significantly less useful, particularly when the source was undisclosed. Party affiliation influenced perceptions of logicalness, and qualitative feedback highlighted concerns about source authenticity, reasoning, and AI limitations. These findings underscore the need for improved LLM design and transparency to build trust among professionals. The work provides actionable insights for integrating AI into fact-checking workflows and advancing human-AI collaboration.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
Causal Perception in Question-Answering Systems
CHI '21· Explainable AI (XAI) +1
- 67%
Dialogues with AI Reduce Beliefs in Misinformation but Build No Lasting Discernment Skills
CHI '26· Misinformation & Fact-Checking +1
- 63%
Large Language Model (LLM)-driven Adversarial Social Influences in Online Information Spread: Risks and Interventions
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)