MUST: Smartwatch-based Multimodal Framework for Predicting Driver State and Takeover Performance

Automated Driving Interface & Takeover DesignIn-Vehicle Haptic, Audio & Multimodal FeedbackSmartwatches & Fitness BandsAutomotive Manufacturers & Vehicle DesignersAutonomous Driving Engineers & Test Drivers

Paper Title

MUST: Smartwatch-based Multimodal Framework for Predicting Driver State and Takeover Performance

Publication Info

  • Topic area: Driver monitoring and takeover performance in conditionally autonomous vehicles using wearable sensors.
  • Keywords: SAE Level 3 automation, smartwatch sensing, multimodal fusion, takeover readiness, driver state prediction, affective computing, CARLA simulator, physiological signals, human-machine interaction, adaptive interfaces.

Background and Problem

  • Problem / challenge: Existing driver monitoring systems face limitations in practicality and reliability, with intrusive physiological sensors, fragile vision-based methods, and simplistic multimodal fusion strategies failing to accurately predict takeover readiness in diverse real-world scenarios.
  • Significance: Ensuring safe and timely takeover in autonomous vehicles is critical to prevent accidents caused by poor situational awareness, slow reaction times, and emotional disruptions during takeover requests (TORs).
  • Motivation and related work: Prior research has explored factors influencing takeover performance, such as non-driving-related tasks (NDRTs) and emotional states, but lacks scalable, unobtrusive, and predictive solutions validated in realistic driving conditions. Physiological sensors and vision-based systems are impractical or unreliable, while multimodal learning approaches often fail to model the interplay between emotion and behavior effectively.

Solution

  • Proposed approach: MUST (Multimodal Unified Smartwatch-based Takeover), a framework leveraging smartwatch signals (PPG and IMU) combined with vehicle telemetry and pre-survey data to predict driver state and takeover performance in real time.
  • Novelty:
    1. Smartwatch-based driver monitoring using unobtrusive wearable sensors integrated with vehicle telemetry.
    2. Asymmetric causal fusion mechanism modeling delayed cross-attention between motion and emotion features.
    3. Enhanced prediction of takeover metrics (TOT and ACT) by integrating behavioral and affective state inferences.
    4. Validation in 13 diverse CARLA-based scenarios with 48 participants, demonstrating robust performance under dynamic hazards.
  • Procedure and key techniques:
    • Stage 1: Modality-specific expert encoders process PPG, IMU, vehicle telemetry, and contextual data into a shared latent space.
    • Stage 2: Fusion block employs causal cross-attention and Feature-wise Linear Modulation (FiLM) to integrate motion and emotion features asymmetrically.
    • Outputs are directed to task-specific heads for motion prediction, affect estimation, and alignment enforcement.
    • Evaluated using CARLA simulator with synchronized smartwatch data and multimodal alerts.

Results

  • Concrete findings:
    • TOT prediction accuracy: 91.4%.
    • ACT regression RMSE: 2.3 seconds.
    • NDRT classification accuracy: 95%.
    • Valence estimation accuracy: 68%; arousal estimation accuracy: 51%.
    • Real-time inference latency: ~10.3 ms (97 FPS) on NVIDIA RTX 4090.
  • Advantage over baselines: Comparable TOT accuracy to camera and physiology-based systems, but using unobtrusive wearable sensing. Unlike prior methods, MUST models full ACT trajectories for richer safety assessments.
  • Experiments / evaluation:
    • Conducted with 48 participants across 13 hazard-driven scenarios involving vehicles, pedestrians, and obstacles.
    • Multimodal alerts issued 5 seconds before hazards; data synchronized at 100 Hz.
    • Ablation studies confirmed the importance of PPG for emotion modeling and IMU for NDRT classification.
  • Limitations and future work:
    • Static simulator lacks inertial forces, reducing ecological validity.
    • Cohort limited to East Asian participants; broader demographic validation needed.
    • Reduced robustness in rare scenarios like oscillatory mode switching.
    • Future work includes on-road validation, diverse participant recruitment, and privacy-preserving edge processing.

Summary

MUST introduces a smartwatch-based multimodal framework for predicting driver state and takeover performance in conditionally autonomous vehicles. By leveraging unobtrusive wearable sensors and asymmetric causal fusion, it integrates behavioral and affective signals to enhance predictions of takeover metrics like TOT and ACT. Validated in diverse CARLA scenarios, MUST achieves high accuracy while addressing limitations of prior systems. Future research aims to expand demographic diversity, validate in naturalistic settings, and refine adaptive strategies for ethical and inclusive deployment.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222900/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791703
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Automated Driving Interface & Takeover Design, In-Vehicle Haptic, Audio & Multimodal Feedback, Smartwatches & Fitness Bands
work
Professions
Automotive Manufacturers & Vehicle Designers, Autonomous Driving Engineers & Test Drivers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers