Pantœnna: Mouth Pose Estimation for AR/VR Headsets Using Low-Profile Antenna and Impedance Characteristic Sensing

Eye Tracking & Gaze InteractionAR Navigation & Context AwarenessImmersion & Presence ResearchSocial Robot Interaction

Document Title

Pantœnna: Mouth Pose Estimation for VR/AR Headsets Using Low-Profile Antenna and Impedance Characteristic Sensing

Document Information

  • Subject Area: Mouth pose tracking in virtual reality and augmented reality
  • Keywords: VR/AR, facial expressions, mouth pose tracking, privacy protection, antenna impedance sensing, continuous tracking, biosensing

Research Background and Problem Statement

  • Identified Issues and Challenges:

    • Accurately capturing users' mouth poses in virtual reality (VR) and augmented reality (AR) is critical for providing high-fidelity immersive experiences, such as multimodal input, expression reproduction, and speech recognition.
    • Existing methods, such as camera-based solutions, pose privacy concerns (e.g., capturing users' oral cavity, upper body areas, or even sensitive content), while audio-based methods are limited to detecting speech and cannot track silent expressions.
    • Biosensing methods (e.g., electromyography, EMG) often require direct skin contact and calibration before use.
  • Research Significance:

    • Mouth poses and expressions convey rich emotional and semantic information, which are essential for communication, trust, and user experience in VR/AR.
    • There is a need for a privacy-friendly, calibration-free technology to replace cameras and traditional biosensing methods while supporting both speech and silent expression tracking.
  • Research Motivation and Related Work:

    • Existing mouth tracking technologies have significant shortcomings in terms of privacy, user experience, and continuous tracking.
    • Similar technologies, such as antenna impedance sensing, have proven effective in gesture tracking, demonstrating potential for automated, non-contact sensing applications.

Solution

  • Proposed Method:

    • Developed a mouth pose tracking system (Pantœnna) based on low-profile antennas and impedance characteristic sensing.
    • Utilized antenna sensing to detect changes in impedance characteristics caused by geometric variations in mouth position, thereby inferring mouth poses.
  • Innovations:

    • Introduced antenna impedance sensing to explore a new domain of mouth pose tracking.
    • Designed a dual-mode, cross-polarized low-profile antenna for easy integration.
    • Developed new operational modes, including dual-mode antenna sensing, polarization detection, and transmission data (S21) utilization.
    • Enabled the system to capture both speech and silent expressions, overcoming limitations of existing technologies.
  • Implementation Steps and Key Techniques:

    1. Antenna Design and Parameter Simulation:
      • Developed a compact low-profile slot antenna and optimized its impedance matching network.
      • Used a dual-polarized antenna structure for higher sensitivity.
    2. Signal Acquisition and Processing:
      • Recorded the antenna's reflection coefficient (S11) and transmission coefficient (S21), generating feature vectors from frequency-domain data.
    3. Machine Learning Prediction Model:
      • Predicted three-dimensional mouth key points using an efficient random forest regression method.
    4. User Experiments and Evaluation:
      • Designed and implemented a prototype system to capture mouth expression and speech motion data for validation.

Research Outcomes

  • Specific Results:

    • Developed a VR/AR mouth pose tracking technology based on antenna impedance sensing.
    • Achieved an average accuracy of 2.6 mm without requiring users to wear skin sensors or undergo extensive calibration.
    • Generated three-dimensional mouth key point data suitable for expression animation or intelligent interaction interfaces.
  • Comparative Advantages:

    • Compared to existing methods, Pantœnna significantly outperforms camera-based solutions in privacy protection.
    • The system can detect both speech and silent expressions, addressing limitations of audio-based methods.
    • Modular design facilitates integration into existing devices, with ultra-low-cost materials and lightweight hardware.
  • Experimental and Evaluation Results:

    • Continuous Mouth Pose Tracking: Achieved an average error of 2.6 mm, maintaining high accuracy across wearing sessions.
    • Discrete Expression Classification Accuracy: 96.3% (within-session) and 91.1% (cross-session).
    • User Identification: Achieved a 99.5% recognition rate among 12 simulated users.
  • Limitations and Future Directions:

    • The current prototype requires further optimization to reduce size, improve directionality, and minimize environmental interference.
    • Test datasets introduced some errors due to adjustments in head-mounted equipment; future work could expand to larger and more diverse populations.
    • Suggested exploration of faster hardware and multi-band antennas to enhance real-time performance.

Conclusion

  • Pantœnna proposes a novel method for privacy-safe, secure, and high-accuracy mouth pose tracking without relying on cameras or microphones.
  • System validation demonstrates significant advantages in dynamic tracking and privacy protection, with potential for exploring multimodal integration.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126805/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606805
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, AR Navigation & Context Awareness, Immersion & Presence Research, Social Robot Interaction
work
Professions
article
Content Status
Full text indexed
hub
Related Papers
0 related papers