Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity Recognition

Human Pose & Activity RecognitionBrain-Computer Interface (BCI) & NeurofeedbackPrivacy Perception & Decision-MakingSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Paper Title

Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity Recognition

Paper Information

  • Research Area: Privacy-preserving activity recognition, generation and learning based on millimeter-wave radar signals
  • Keywords: Human activity recognition, Doppler radar, dataset generation, cross-domain translation, privacy preservation, synthetic data, deep learning

Research Background and Problem Statement

  • Identified Problem: Millimeter-wave Doppler radar, as a privacy-preserving activity recognition sensor, is low-cost and provides rich signals. However, the lack of large-scale training datasets limits its further development in deep learning applications.
  • Importance of the Problem: Privacy concerns are particularly prominent in modern sensing systems. Doppler radar is promising due to its privacy-preserving characteristics but cannot compete with computer vision and audio recognition technologies, which benefit from abundant training data resources.
  • Research Motivation and Related Work:
    • Existing work primarily uses microphones, cameras, and motion capture devices for activity recognition, which suffer from severe privacy issues and high costs.
    • Doppler radar has gained attention for its noise tolerance and privacy-preserving features, but its adoption is hindered by the lack of data.
    • While attempts to synthesize Doppler data (e.g., from point clouds or motion capture data) exist, they are limited in data sources and lack sufficient conversion accuracy compared to video-based generation.

Proposed Solution

  • Proposed Method:
    • Develop a software pipeline to generate synthetic Doppler radar data from ordinary videos.
    • Create a dataset of Doppler signals approximating daily activities by transforming video-based human activities into 3D meshes and viewpoints, ultimately generating realistic synthetic signals using deep learning encoder-decoder models.
  • Innovations:
    • Construct the first Doppler data generation framework using videos as input, sourcing data from video libraries (e.g., YouTube and structured video datasets) instead of custom devices.
    • Compared to prior work using motion capture and depth point clouds, this approach covers more scenarios and provides richer data types.
    • Introduce a method for training with a mix of synthetic data and a small amount of real sensor data.
  • Key Techniques and Implementation Steps:
    1. 3D Mesh Fitting: Extract human 3D meshes from videos using deep learning models (e.g., VIBE).
    2. Viewpoint Synthesis: Set virtual viewpoints around the user to enhance data diversity.
    3. Radar Cross-Section and Radial Velocity Calculation: Simulate local motion characteristics for each video frame.
    4. Visibility and Occlusion Handling: Filter out occluded or back-facing mesh points.
    5. Initial Doppler Signal Generation: Construct coarse-grained signal histograms based on radial velocity.
    6. Signal Refinement: Optimize signals using an encoder-decoder model to approximate real data characteristics.
    7. Activity Classification Training: Train an activity recognition model using synthetic data and VGG-16 network.

Research Outcomes

  • Specific Results: The model trained on synthetic Doppler data achieved 81.4% accuracy in a 12-class activity recognition task. When combined with a small amount of real data, accuracy improved to 95.9%, approaching or even surpassing methods relying entirely on real data (90.2%).
  • Comparison with Existing Solutions:
    • Reduced manual data collection time compared to traditional Doppler radar systems.
    • Demonstrated excellent transferability in cross-user experiments.
  • Evaluation Results:
    • Training with synthetic data only: Accuracy 81.4%.
    • Training with all real data: Accuracy 90.2%.
    • Training with a mix of synthetic and limited real data: Accuracy improved to 95.9%.
  • Limitations and Future Directions:
    • The current method cannot handle dynamic backgrounds or scenarios involving sensor movement.
    • Modeling complex signal interactions in multi-user scenarios remains challenging.
    • Future work could incorporate more advanced generative adversarial networks and temporal processing models (e.g., RNNs) to optimize recognition methods.
    • Long-term storage of privacy-sensitive data also requires special attention.

Notes

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47222/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445138
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human Pose & Activity Recognition, Brain-Computer Interface (BCI) & Neurofeedback, Privacy Perception & Decision-Making
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers