Multimodal Fusion for Driver Referencing: A Comparison of Pointing to Objects Inside and Outside the Vehicle

In-Vehicle Haptic, Audio & Multimodal FeedbackHand Gesture RecognitionVoice User Interface (VUI) DesignAutomotive Manufacturers & Vehicle DesignersAutonomous Driving Engineers & Test Drivers

Title of the Paper

Multimodal Driver Referencing: A Comparison of Pointing to Objects Inside and Outside the Vehicle

Paper Information

  • Research Area: Intelligent User Interaction, Driver Behavior Studies, Gesture and Gaze Integration
  • Keywords: Multimodal Fusion, Natural User Interaction, Gaze Tracking, Head Pose, Gesture Recognition

Research Background and Problem Statement

  • Problems or Challenges:

    1. How to accurately detect the driver's pointing intention, regardless of whether the target is inside or outside the vehicle.
    2. Single modalities (e.g., eye movement, head pose, or gestures) are limited in their ability to demonstrate consistent performance across different scenarios.
    3. Existing technologies rarely address the differences between internal and external contexts in an integrated manner.
  • Significance:

    • The development of Natural User Interaction (NUI) in autonomous driving and intelligent driving assistance can significantly enhance user experience.
    • Multimodal interaction (voice, gaze, gestures) helps reduce the distraction risks associated with touch-based interaction.
  • Research Motivation and Related Work:

    • Related studies indicate that combining multiple interaction methods (e.g., gestures and voice) can enhance the naturalness and accuracy of in-vehicle interactions.
    • However, there is a lack of research comparing pointing behaviors inside and outside the vehicle, with most studies focusing on simulated experiments and lacking validation in real-world environments.

Proposed Solution

  • Methods or Solutions:

    1. A two-stage multimodal fusion framework is proposed:
      • Stage 1: Use a shallow Convolutional Neural Network (CNN) to classify whether the driver's pointing target is inside or outside the vehicle.
      • Stage 2: Based on the classification results from Stage 1, apply a suitable deep CNN model for multimodal fusion to infer the driver's pointing direction.
    2. Use voice as a trigger to mark the reference event time points.
    3. Multimodal inputs include finger pointing, gaze tracking, and spatial position and direction of head pose.
  • Innovations:

    1. Simultaneously address the comparison of pointing behaviors for targets inside and outside the vehicle.
    2. Quantify and analyze driver pointing behaviors in real vehicles and environments for the first time.
    3. Propose a multimodal model fusion strategy that significantly improves pointing accuracy based on deep learning.
  • Implementation Steps and Technical Details:

    1. Data Collection: Collect pointing event data from drivers in real vehicles using contactless sensors (e.g., gesture cameras and vision cameras), targeting 12 regions inside the vehicle and 5 landmarks outside.
    2. Feature Extraction: Extract 3D spatial position and directional information from finger endpoints, gaze, and head pose.
    3. Data Preprocessing: Handle missing data using linear interpolation and normalize coordinate axes.
    4. Model Training: Train CNN models based on the collected dataset, optimizing for classification and angle regression tasks.
    5. Experimental Setup: Test and evaluate single modalities, multimodal combinations, and cross-dataset generalization capabilities.

Research Results

  • Specific Results:

    • The proposed framework achieved 98.6% classification accuracy in distinguishing between targets inside and outside the vehicle.
    • The multimodal fusion model significantly improved the accuracy of pointing direction prediction:
      • For targets inside the vehicle, the Mean Angular Deviation (MAD) decreased from 6.1° (single modality) to 2.5°.
      • For targets outside the vehicle, the MAD decreased from 9.3° (single modality) to 7°.
    • Comparison between single and multimodal approaches showed that using a single modality (e.g., gaze or gestures) could not achieve optimal performance in both scenarios, while the fusion method performed better.
  • Advantages Comparison:

    • Multimodal fusion significantly outperformed single modalities in both accuracy and hit rate.
    • The framework effectively addressed differences in driver behavior in real-world environments and compensated for the limitations of single modalities, such as occlusion or missing data.
  • Experimental or Evaluation Results:

    • Results for single modalities showed significant differences between inside and outside the vehicle, with target distance and pose variations having a major impact on accuracy.
    • Individual driver analysis revealed that fusion results showed significant improvements in accuracy and hit rate, though some participants performed poorly in specific environments.
    • Cross-dataset learning validated the model's dependency on specific environmental conditions; while joint training improved generalization, it sacrificed some precision.
  • Limitations and Future Directions:

    • Limitations:
      1. Experiments were conducted only in stationary vehicles, without considering the potential impact of dynamic vehicle conditions.
      2. The model relies on the reliability of sensor data, and inaccuracies in gesture recognition could significantly affect identification accuracy.
    • Future Directions:
      1. Validate the framework's effectiveness in dynamic environments (moving vehicles or driving scenarios).
      2. Incorporate more semantic information into multimodal fusion, such as combining voice content with contextual background.
      3. Explore personalized models to adapt to different drivers' behavior patterns and habits.

Conclusion

This study demonstrates how a multimodal fusion-based natural interaction system can accurately predict drivers' pointing targets, offering significant implications and potential applications for enhancing intelligent driving experiences.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/79938/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3490099.3511142
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
In-Vehicle Haptic, Audio & Multimodal Feedback, Hand Gesture Recognition, Voice User Interface (VUI) Design
work
Professions
Automotive Manufacturers & Vehicle Designers, Autonomous Driving Engineers & Test Drivers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers