Understanding Driving Distractions: A Multimodal Analysis on Distraction Characterization

Automated Driving Interface & Takeover DesignIn-Vehicle Haptic, Audio & Multimodal FeedbackHuman Pose & Activity RecognitionAutonomous Driving Engineers & Test Drivers

Title of the Paper

Understanding Driving Distractions: A Multimodal Analysis on Distraction Characterization

Paper Information

  • Research Area: Multimodal analysis of driving distraction behaviors
  • Keywords: distracted driving, machine learning, physiological signal processing, action unit analysis, multimodal interaction, multimodal dataset

Research Background and Issues

  • Identified Challenges or Problems:

    • Distracted driving is a major cause of traffic accidents, but traditional computer vision-based methods fail to comprehensively capture distraction behaviors, particularly cognitive distractions.
    • Although several studies have attempted to address distraction detection, there are still limitations in recognizing driving distractions caused by cognitive and emotional factors.
  • Significance of the Research:

    • Distracted driving causes numerous traffic accidents and significant economic losses annually. For instance, in the United States, distracted driving results in economic losses of up to $15.433 billion per year.
    • Addressing the issue of distracted driving can not only improve road safety but also support the development of smarter and more personalized driver assistance systems.
  • Motivation and Related Work:

    • Existing studies primarily focus on detecting driving distractions through facial features and eye tracking; however, these methods are limited in handling individualized cognitive and emotional distractions.
    • Physiological signals have shown advantages in analyzing cognitive states, but their integration with multimodal data has not been thoroughly studied.

Solution

  • Method/Solution:

    • A novel multimodal driving distraction behavior dataset is proposed, covering information channels such as visual, acoustic, near-infrared, thermal infrared, physiological, and linguistic data.
    • The dataset includes multimodal data from 45 participants in a simulated driving environment, recording four different types of distractions (three cognitive distractions and one physical distraction).
  • Innovations:

    • The dataset captures not only physical distraction behaviors but also specifically designed cognitive and emotional distraction tasks.
    • Distraction detection and recognition experiments tested the effectiveness of visual and physiological signals, while exploring the potential of multimodal fusion.
  • Implementation Steps and Key Technologies:

    • Data collection involved a simulated driving environment with "free driving" and tasks designed to induce four types of distractions:
      1. Text input (physical, cognitive, and visual distractions).
      2. N-Back task (cognitive load).
      3. Listening to the radio and commenting (cognitive and emotional induction).
      4. Voice interaction with simulated GPS (cognitive frustration and emotion-induced distraction).
    • OpenFace was used to extract facial action unit (AU) features of drivers.
    • Physiological signals were collected, including blood volume pulse (BVP), skin temperature, galvanic skin response, and respiration rate.
    • Statistical features were extracted using a sliding window technique, and multimodal information fusion was attempted (early fusion and late fusion).

Research Outcomes

  • Specific Findings:

    • Physiological signals outperformed visual modalities, especially in distinguishing cognitive distractions.
    • Multimodal methods achieved the best performance in distraction recognition tasks, with late fusion methods demonstrating stable results across all tasks.
    • Different types of distractions were significantly correlated with specific visual action units and physiological features, e.g., emotion-induced distractions were directly associated with high-frequency BVP signals.
  • Advantages:

    • Experiments demonstrated that multimodal methods can capture both physical and cognitive distraction characteristics, with fusion methods outperforming single modalities in classification performance.
    • Selected action units and physiological frequency features were particularly important for detecting and classifying distraction behaviors.
  • Experimental or Evaluation Results:

    • Visual modality performed best in the text input task (F1 score exceeding 0.85).
    • Multimodal late fusion achieved an average F1 score of 0.94 in binary classification tasks and 0.67 in four-class classification tasks.
    • Physiological signals showed stable performance in more complex cognitive interference classification tasks, while visual modality performance declined.
  • Limitations and Future Directions:

    • Current experiments are based on simulated driving environments, and distraction types and intensities in real-world settings may differ.
    • Future work is recommended to explore additional physiological signal channels (e.g., heart rate variability and EEG) and test the dataset's cross-domain robustness to enhance model generalization.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/57966/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3397481.3450635
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Automated Driving Interface & Takeover Design, In-Vehicle Haptic, Audio & Multimodal Feedback, Human Pose & Activity Recognition
work
Professions
Autonomous Driving Engineers & Test Drivers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers