Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection
Paper Title
Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection
Publication Info
- Topic area: Human strategies in detecting deepfake videos and implications for media literacy.
- Keywords: Deepfake detection, multimodal strategies, media literacy, visual cues, audio cues, knowledge-based strategies, calibration, disinformation, human-AI collaboration.
Background and Problem
- Problem / challenge: Humans struggle to reliably detect deepfake videos, often exhibiting overconfidence in their judgments, and existing interventions lack a clear understanding of natural detection strategies.
- Significance: Deepfakes pose threats to trust, reputations, and political stability, making effective detection critical for combating disinformation.
- Motivation and related work: Prior research has explored AI-human collaboration and qualitative analyses of detection strategies but has not quantitatively examined how humans combine visual, audio, and knowledge cues in practice or how these strategies interact.
Solution
- Proposed approach: Quantitative study analyzing human detection strategies for deepfake videos, focusing on accuracy, confidence, and calibration across visual, audio, and knowledge-based cues.
- Novelty:
- Empirical evidence on human accuracy and calibration in detecting real vs. deepfake videos.
- Quantitative analysis of multimodal strategies and their combinations.
- Mapping cue interactions to identify diagnostic and misleading patterns.
- Procedure and key techniques:
- Study with 195 participants evaluating 20 videos (10 real, 10 deepfake).
- Participants judged authenticity, rated confidence, and reported cues used.
- Metrics included accuracy and Expected Calibration Error (ECE).
- Statistical tests (t-tests, ANOVA) and association rule mining analyzed strategy effectiveness and cue interactions.
Results
- Concrete findings:
- Overall accuracy: 76% (real videos: 79%, deepfakes: 72%).
- Calibration error (ECE): Higher for deepfakes (0.32) than real videos (0.24).
- Visual appearance cues were effective for deepfakes (accuracy: 84%), while audio cues (e.g., vocal tone) were more diagnostic for real videos (accuracy: 81%).
- Advantage over baselines:
- Multimodal strategies improved real video detection but did not significantly enhance deepfake detection.
- Combining visual, audio, and knowledge cues boosted accuracy and calibration for real videos.
- Experiments / evaluation:
- Participants evaluated videos featuring public figures across diverse topics (e.g., politics, entertainment).
- Cue combinations analyzed via association rule mining and network graphs.
- Limitations and future work:
- Restricted participant demographics (young adults).
- Limited diversity in deepfake manipulations.
- Predefined cue checklist may influence spontaneous strategy reporting.
- Future work should expand populations, stimuli, and unprompted methods.
Summary
This study quantitatively analyzed human strategies for detecting deepfake videos, revealing that multimodal integration of visual, audio, and knowledge cues improves detection of real videos but offers limited benefits for deepfakes. Visual appearance and intuition were effective for spotting deepfakes, while audio cues and prior knowledge supported real video detection. Calibration analyses highlighted overconfidence in deepfake judgments. Findings suggest media literacy interventions should focus on teaching effective cue use and recognizing misplaced confidence. Future research should broaden demographic and stimulus diversity to refine detection strategies.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 60%
Causal Perception in Question-Answering Systems
CHI '21· Explainable AI (XAI) +1
- 60%
Dialogues with AI Reduce Beliefs in Misinformation but Build No Lasting Discernment Skills
CHI '26· Misinformation & Fact-Checking +1
Based on Jaccard similarity of research subtopics & professions (≥60%)