Voice Presentation Attack Detection through Text-Converted Voice Command Analysis

Intelligent Voice Assistants (Alexa, Siri, etc.)Voice AccessibilityDeepfake & Synthetic Media DetectionCybersecurity EngineersAI/ML Researchers & EngineersPrivacy Policy Makers

Voice assistants are quickly being upgraded to support advanced, security-critical commands such as unlocking devices, checking emails, and making payments. In this paper, we explore the feasibility of using users' text-converted voice command utterances as classification features to help identify users' genuine commands, and detect suspicious commands. To maintain high detection accuracy, our approach starts with a globally trained attack detection model (immediately available for new users), and gradually switches to a user-specific model tailored to the utterance patterns of a target user. To evaluate accuracy, we used a real-world voice assistant dataset consisting of about 34.6 million voice commands collected from 2.6 million users. Our evaluation results show that this approach is capable of achieving about 3.4% equal error rate (EER), detecting 95.7% of attacks when an optimal threshold value is used. As for those who frequently use security-critical (attack-like) commands, we still achieve EER below 5%.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/5659/2019

AdRecommended

Learn AI Coding at CodeNow

At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2019
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Voice Accessibility, Deepfake & Synthetic Media Detection
work
Professions
Cybersecurity Engineers, AI/ML Researchers & Engineers, Privacy Policy Makers
article
Content Status
Abstract only
hub
Related Papers
0 related papers