Uncovering Human Traits in Determining Real and Spoofed Audio: Insights from Blind and Sighted Individuals

Explainable AI (XAI)Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Deepfake & Synthetic Media DetectionAssistive Technology SpecialistsHCI Researchers

Title of the Paper

Uncovering Human Traits in Determining Real and Spoofed Audio: Insights from Blind and Sighted Individuals

Paper Information

  • Subject Areas: Human-AI Collaboration, Human-Computer Interaction, Audio Authenticity Detection, Countermeasures Against Spoofed Audio
  • Keywords: Audio Perception, Generative AI, Deepfake Audio, Blind Individuals, Sighted Individuals, Text-to-Speech (TTS), Real Audio, Spoofed Audio, Replay Attacks, Audio Watermarking, AI Collaboration

Research Background and Problem

  • Problem or Challenge:
    • With the advancement of generative artificial intelligence, it is becoming increasingly difficult to distinguish between real human speech and machine-generated speech.
    • Blind individuals cannot rely on visual cues (e.g., facial expressions or inconsistencies in lighting in videos) and may depend more on audio to judge authenticity.
  • Significance:
    • The proliferation of spoofed audio technologies could impact the security of domains such as banking and smart voice assistants.
    • Exploring the differences in audio perception between blind and sighted individuals can help optimize countermeasures and influence the design of future audio technologies.
  • Motivation and Related Work:
    • Existing automatic speech verification systems face challenges posed by generative AI's ability to produce spoofed audio.
    • Previous research has emphasized the relationship between visual impairment and auditory abilities, but few studies have explored how blind individuals perceive audio authenticity.

Solution

  • Method:
    • Two experimental studies were conducted to explore how blind and sighted individuals judge audio authenticity.
    • In the first study, 12 blind participants classified 63 challenging audio samples, and interviews were conducted to understand their decision-making criteria.
    • In the second study, the sample size was expanded to 60 participants (30 blind and 30 sighted individuals) to analyze 96 audio samples.
  • Innovations:
    • Identified unique human advantages in judging audio authenticity, such as focusing on pronunciation, tone, breathing sounds, and emotional cues.
    • Explored cognitive model differences between blind and sighted individuals, emphasizing the role of visual cues in perceiving generated audio.
  • Implementation Steps:
    • First experiment: Provide audio samples, record decision-making processes, and gather participants' confidence levels and challenges.
    • Second experiment: Conduct a larger-scale online survey, collecting responses and open-ended feedback from participants.
    • Combine quantitative data analysis with text analysis to summarize key findings and propose design and policy recommendations.

Research Findings

  • Specific Findings:
    • Both blind and sighted individuals were able to identify audio authenticity, with their accuracy significantly outperforming state-of-the-art machine learning models (e.g., the ASSERT algorithm, which had 0% accuracy on challenging samples).
    • The two groups exhibited different performance levels in identifying various types of spoofed audio:
      • Blind participants performed better with TTS-generated audio (accuracy: 91%).
      • Sighted participants excelled in detecting deepfake audio (accuracy: 71%).
    • Both groups struggled to identify replay attack audio (average accuracy below 50%).
  • Advantages:
    • Blind individuals tended to rely on their extensive experience with using text-to-speech tools, enabling high accuracy in identifying TTS-generated audio.
    • Sighted individuals used visualized mental models, such as associating audio with specific facial expressions or pronunciation styles, which helped them better detect deepfake audio.
  • Limitations and Future Directions:
    • The current study's dataset was not fully balanced, particularly with fewer deepfake and TTS samples.
    • The complexity and prevalence of replay audio in real-world scenarios require further research.
    • Future work should focus on designing binary classifiers based on human traits and perceptible audio watermarking technologies.
    • Explore human-AI collaboration mechanisms (e.g., uncertainty quantification or delayed learning architectures) to optimize spoofed audio detection.

Design and Social Impact

  • Design Inspirations:
    • Improve assistive tools (e.g., screen readers) by incorporating more realistic deepfake voices to enhance user experience.
    • Develop human-perceptible watermarking mechanisms to help users quickly distinguish AI-generated content.
    • Leverage expert judgment in audio authenticity to inspire recruitment or collaboration in relevant fields.
  • Social Impact:
    • As spoofing technologies advance, both blind and sighted individuals may face increased cognitive burdens, highlighting the need for security designs that reduce user verification stress.
    • Use AI-assisted tools to create employment opportunities for blind individuals, such as training them to become expert audio authenticity evaluators.

This study deepens the understanding of machine-manipulated audio and provides critical insights for future intelligent interactions, technology design, and public policy.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/146813/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642817
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Deepfake & Synthetic Media Detection
work
Professions
Assistive Technology Specialists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers