WavoID: Robust and Secure Multi-modal User Identification via mmWave-voice Mechanism

Eye Tracking & Gaze InteractionBrain-Computer Interface (BCI) & NeurofeedbackPasswords & AuthenticationSoftware Engineers & DevelopersCybersecurity EngineersAI/ML Researchers & EngineersPrivacy Policy Makers

Title of the Paper

WavoID: Robust and Secure Multi-modal User Identification via mmWave-voice Mechanism

Paper Information

  • Subject Area: Multi-modal user identification, integration of mmWave and voice signals
  • Keywords: User authentication, voice recognition, mmWave sensing, multi-modal fusion, user identity, noise robustness, security

Research Background and Problem

  • Problems or Challenges:

    1. Traditional voice-based user identification systems are vulnerable to attacks (e.g., replay attacks, impersonation attacks, and adversarial attacks) due to the propagation characteristics of sound.
    2. Single-modal voice recognition systems are unreliable under noise interference and user movement.
    3. Although multi-modal systems incorporate additional information, they rarely address defenses against multi-modal attacks or the lack of intrinsic correlation in fusion methods.
  • Significance: With the widespread adoption of voice-interactive devices, such as smart speakers and financial systems, it is crucial to provide user identification systems with robust and secure capabilities.

  • Motivation and Related Work:

    1. Shortcomings: Voice recognition systems are susceptible to environmental noise or impersonation attacks, leading to poor accuracy; wireless signal-based systems (e.g., WiFi) have limitations and lack robustness.
    2. Inspiration: mmWave radar excels in sensing subtle vibrations and resisting noise, while voice signals can compensate for motion interference in mmWave signals.
    3. Identification systems should meet the following requirements:
      • No need for auxiliary operations.
      • Robustness under adverse conditions (e.g., noise, motion interference).
      • Effective defense against both multi-modal and single-modal attacks.

Solution

  • Method/Solution: The authors propose a multi-modal user identification system, WavoID, which integrates dual-modal features from millimeter-wave (mmWave) sensing and voice input to overcome the limitations of traditional single-modal methods.

  • Innovations:

    1. Integration of complementary features from mmWave-sensed vocal cord vibrations and recorded voice signals.
    2. Introduction of a cross-modal fine-grained waveform estimation method to remove noise and a dual-modal liveness detection module to enhance anti-spoofing capabilities.
    3. Use of response map generation and cross-modal filtering to strengthen feature representation.
    4. Adoption of a specialized network architecture (CBANet) to improve user identification performance.
  • Implementation Steps and Key Techniques:

    1. Fine-grained Waveform Estimation Module: Utilizes signal decomposition and correlation matrices to extract useful sub-signals from mmWave and voice data, eliminating noise, motion interference, etc.
    2. Dual-modal Liveness Detection Module: Extracts biological liveness features from mmWave (e.g., reflection coefficients) and frequency domain features from voice (e.g., CQCC), and detects attacks through similarity comparison.
    3. Cross-modal Feature Fusion Module: Employs discriminative correlation filtering (DCF) to generate response maps and uses segmented convolution for fusion, enhancing robustness against interference.
    4. User Identification Network (CBANet): Leverages residual networks and attention modules to learn and identify multi-modal inputs.

Research Outcomes

  • Specific Results:

    1. WavoID achieved over 98% identification accuracy and an average Equal Error Rate (EER) of 1.24% on a dataset of 100 users.
    2. In common attack scenarios (e.g., replay, impersonation, adversarial attacks), WavoID successfully detected and rejected over 99% of malicious samples.
  • Advantages Over Existing Solutions:

    1. The multi-modal system outperforms single-modal methods (e.g., voice-based or mmWave-based systems).
    2. Demonstrates high robustness in complex scenarios (e.g., noise interference and motion interference).
    3. The proposed dual-modal detection mechanism and fusion approach significantly enhance system security.
  • Experimental and Evaluation Results:

    1. In experiments, even under strong noise conditions (e.g., 70dB of various background noise) or user motion interference (e.g., walking, phone usage), the system maintained 97-99% identification accuracy.
    2. Exhibited high security detection capabilities against different attacks (e.g., single-modal or combined attacks), with attack rejection rates close to 99%.
  • Limitations and Future Directions:

    1. Current hardware implementation faces challenges in complexity and cost, requiring more affordable and compact hardware solutions.
    2. Experiments were conducted in laboratory settings; future work should validate the system's performance in open, real-world scenarios.
    3. Further optimization is needed to reduce latency for real-time requirements and explore deployment potential on edge devices or mobile platforms.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126761/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606775
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
9 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Brain-Computer Interface (BCI) & Neurofeedback, Passwords & Authentication
work
Professions
Software Engineers & Developers, Cybersecurity Engineers, AI/ML Researchers & Engineers, Privacy Policy Makers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers