HRTF Estimation in the Wild

Haptic WearablesImmersion & Presence ResearchBiosensors & Physiological Monitoring

Title of the Paper

HRTF Estimation in the Wild

Paper Information

  • Subject Area: Spatial audio technology and personalized Head-Related Transfer Function (HRTF) inference
  • Keywords: Spatial audio, Head-Related Transfer Function (HRTF), virtual reality, sound localization, machine learning, binaural recording, head tracking, personalized audio technology

Research Background and Problem Statement

  • Main Issues or Challenges:

    • Head-Related Transfer Function (HRTF) is crucial for creating immersive spatial audio experiences, but its acquisition requires specialized equipment and expensive testing processes.
    • Due to variations in human anatomy, HRTFs differ across individuals. Using generic HRTFs can lead to inaccurate sound localization and poor auditory experiences.
    • Traditional methods for personalized HRTF measurement are time-consuming, costly, and require specialized laboratory environments, such as anechoic chambers and high-precision scanning equipment.
  • Significance of the Research:

    • Personalized HRTFs are critical in fields such as virtual reality, gaming, cinematic sound effects, and music, as they significantly enhance audio immersion and localization accuracy.
    • Addressing the complexity and cost of HRTF measurement would enable broader adoption on consumer devices like headphones or earbuds.
  • Motivation and Related Work:

    • The widespread use of earbud devices (e.g., Apple AirPods) provides an opportunity to capture data using simple devices in everyday environments.
    • Existing methods have limitations:
      1. Dependence on expensive equipment or specialized environments.
      2. Reliance on cumbersome methods such as scanning and measurement or user-involved quantitative tasks, which prevent passive measurement.
      3. Inability to effectively utilize recorded data from natural environments for personalized HRTF inference.
    • This study proposes a novel method to generate personalized HRTFs using everyday binaural recordings and head-tracking data combined with machine learning, aiming to overcome these barriers.

Solution

  • Proposed Method or Solution:

    • Utilize binaural recordings and head tracking from headphones, combined with a machine learning model (based on a U-Net deep network), to infer personalized HRTFs from head movements and environmental sound variations in natural settings.
    • Revolutionize traditional methods based on controlled sound sources and laboratory measurements by simplifying data collection to everyday head movements of earbud users.
    • Core approach: Learn sound filtering variations from different directions (predicting ILDs) and combine this with user head-tracking information to generate target HRTFs.
  • Innovations:

    • Data Source Innovation: Collect data in ordinary daily environments (instead of anechoic chambers) without requiring additional equipment or complex head scanning.
    • Model Design: Data is directly captured from binaural recordings via headphones, actively leveraging natural head movements as conditional input.
    • Automation: Minimize user involvement, eliminating the need for users to answer complex questions or carry depth sensors.
  • Implementation Steps and Key Techniques:

    1. Data Collection: Use earbuds with built-in binaural microphones and head IMUs to collect audio and head movement data.
    2. Data Modeling: Employ a U-Net model to predict frequency-dependent sound filtering characteristics from captured binaural audio.
    3. Noise Source and Head Movement Matching: During user initialization, locate noise sources and infer personalized HRTFs by analyzing sound source positions and head rotations.
    4. Optimization and Aggregation: Combine prediction data under different environmental conditions, integrating multi-source and multi-angle information to gradually improve HRTF prediction accuracy.
    5. Output Conversion: Combine spectral filtering characteristics with existing interaural time differences (ITDs) to compute the final personalized HRTF and HRIR.

Research Outcomes

  • Specific Achievements:

    • HRTF Accuracy: The predicted HRTFs from the model exhibit high consistency with independently measured anechoic chamber HRTFs, with significantly better log-spectral distortion (LSD) values compared to existing methods.
    • Improved Sound Localization: Users demonstrated superior sound localization performance in virtual audio environments using the HRTFs generated by this study compared to generic HRTFs.
    • Reduced Front-Back Confusion: The HRTFs generated by this method reduced listeners' front-back confusion rate to 14.8%, a notable improvement over generic HRTFs (29.0%).
  • Advantages Over Existing Solutions:

    • Unlike traditional methods relying on 3D head scanning or image analysis, this method is entirely scan-free.
    • Eliminates the need for users to actively participate in complex experimental processes, maximizing data collection simplicity.
    • Enhanced adaptability to noise sources and improved generalization across diverse environments.
  • Experimental or Evaluation Results:

    • Experiments validated the usability of predicted HRTFs in noisy environments and virtual display scenarios.
    • The LSD value was significantly lower than the baseline of traditional methods (model: 4.38 dB, compared to generic HRTF: 7.32 dB).
    • User trials demonstrated significant improvements in directional perception and sound localization accuracy.
  • Limitations and Future Directions:

    • The current method is limited to single stationary sound sources and has not yet been validated in multi-source or dynamic sound source scenarios.
    • Initial manual calibration of the sound source (direction locking) is required, which needs further optimization for full automation.
    • More detailed calibration is needed to model the impact of actual microphone positions in commercial headphones.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126819/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606782
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Haptic Wearables, Immersion & Presence Research, Biosensors & Physiological Monitoring
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
3 related papers