Uncovering Human Traits in Determining Real and Spoofed Audio: Insights from Blind and Sighted Individuals
Authors
Explainable AI (XAI)Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Deepfake & Synthetic Media DetectionAssistive Technology SpecialistsHCI Researchers
Title of the Paper
Uncovering Human Traits in Determining Real and Spoofed Audio: Insights from Blind and Sighted Individuals
Paper Information
- Subject Areas: Human-AI Collaboration, Human-Computer Interaction, Audio Authenticity Detection, Countermeasures Against Spoofed Audio
- Keywords: Audio Perception, Generative AI, Deepfake Audio, Blind Individuals, Sighted Individuals, Text-to-Speech (TTS), Real Audio, Spoofed Audio, Replay Attacks, Audio Watermarking, AI Collaboration
Research Background and Problem
- Problem or Challenge:
- With the advancement of generative artificial intelligence, it is becoming increasingly difficult to distinguish between real human speech and machine-generated speech.
- Blind individuals cannot rely on visual cues (e.g., facial expressions or inconsistencies in lighting in videos) and may depend more on audio to judge authenticity.
- Significance:
- The proliferation of spoofed audio technologies could impact the security of domains such as banking and smart voice assistants.
- Exploring the differences in audio perception between blind and sighted individuals can help optimize countermeasures and influence the design of future audio technologies.
- Motivation and Related Work:
- Existing automatic speech verification systems face challenges posed by generative AI's ability to produce spoofed audio.
- Previous research has emphasized the relationship between visual impairment and auditory abilities, but few studies have explored how blind individuals perceive audio authenticity.
Solution
- Method:
- Two experimental studies were conducted to explore how blind and sighted individuals judge audio authenticity.
- In the first study, 12 blind participants classified 63 challenging audio samples, and interviews were conducted to understand their decision-making criteria.
- In the second study, the sample size was expanded to 60 participants (30 blind and 30 sighted individuals) to analyze 96 audio samples.
- Innovations:
- Identified unique human advantages in judging audio authenticity, such as focusing on pronunciation, tone, breathing sounds, and emotional cues.
- Explored cognitive model differences between blind and sighted individuals, emphasizing the role of visual cues in perceiving generated audio.
- Implementation Steps:
- First experiment: Provide audio samples, record decision-making processes, and gather participants' confidence levels and challenges.
- Second experiment: Conduct a larger-scale online survey, collecting responses and open-ended feedback from participants.
- Combine quantitative data analysis with text analysis to summarize key findings and propose design and policy recommendations.
Research Findings
- Specific Findings:
- Both blind and sighted individuals were able to identify audio authenticity, with their accuracy significantly outperforming state-of-the-art machine learning models (e.g., the ASSERT algorithm, which had 0% accuracy on challenging samples).
- The two groups exhibited different performance levels in identifying various types of spoofed audio:
- Blind participants performed better with TTS-generated audio (accuracy: 91%).
- Sighted participants excelled in detecting deepfake audio (accuracy: 71%).
- Both groups struggled to identify replay attack audio (average accuracy below 50%).
- Advantages:
- Blind individuals tended to rely on their extensive experience with using text-to-speech tools, enabling high accuracy in identifying TTS-generated audio.
- Sighted individuals used visualized mental models, such as associating audio with specific facial expressions or pronunciation styles, which helped them better detect deepfake audio.
- Limitations and Future Directions:
- The current study's dataset was not fully balanced, particularly with fewer deepfake and TTS samples.
- The complexity and prevalence of replay audio in real-world scenarios require further research.
- Future work should focus on designing binary classifiers based on human traits and perceptible audio watermarking technologies.
- Explore human-AI collaboration mechanisms (e.g., uncertainty quantification or delayed learning architectures) to optimize spoofed audio detection.
Design and Social Impact
- Design Inspirations:
- Improve assistive tools (e.g., screen readers) by incorporating more realistic deepfake voices to enhance user experience.
- Develop human-perceptible watermarking mechanisms to help users quickly distinguish AI-generated content.
- Leverage expert judgment in audio authenticity to inspire recruitment or collaboration in relevant fields.
- Social Impact:
- As spoofing technologies advance, both blind and sighted individuals may face increased cognitive burdens, highlighting the need for security designs that reduce user verification stress.
- Use AI-assisted tools to create employment opportunities for blind individuals, such as training them to become expert audio authenticity evaluators.
This study deepens the understanding of machine-manipulated audio and provides critical insights for future intelligent interactions, technology design, and public policy.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How do visual and non-visual cues affect blind and sighted people's ability to judge audio authenticity?Category: Blind and Low-Vision AccessibilitySimilar questionsarrow_forward
- How do blind and sighted people differ in identifying different types of forged audio (e.g., TTS-generated, deepfake, and replay-attack audio)?Category: Blind and Low-Vision AccessibilitySimilar questionsarrow_forward
- How can human traits inform the design of forged audio detection systems?Category: Blind and Low-Vision AccessibilitySimilar questionsarrow_forward
lightbulb
Practical Problems
1- Users struggle to distinguish real from AI-generated audio, increasing security risk.Category: Blind and Low-Vision AccessibilitySimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642817
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Deepfake & Synthetic Media Detection
work
Professions
Assistive Technology Specialists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers