HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI

Teleoperation & TelepresenceHuman Pose & Activity RecognitionHand Gesture RecognitionAutonomous Driving Engineers & Test DriversPhysicians, Nurses & CliniciansAI/ML Researchers & Engineers

Paper Title

HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI

Publication Info

  • Topic area: Long-range human-robot interaction (HRI) focusing on command source identification (CSI).
  • Keywords: Human-robot interaction, command source identification, optical-inertial fusion, wearable IMU, robot-mounted camera, long-range interaction, gesture recognition, multi-user scenarios, public-space robotics, spectral motion analysis.

Background and Problem

  • Problem / challenge: Long-range HRI scenarios introduce challenges such as visual ambiguity, unnatural interaction methods, and sensor noise, making CSI difficult in multi-user environments.
  • Significance: Reliable CSI is essential for enabling intuitive and robust interactions in public spaces, such as summoning service robots or directing drones from a distance.
  • Motivation and related work: Prior work largely focuses on near-range HRI or requires unnatural gestures, expensive hardware, or pre-deployed infrastructure. Existing visual-inertial fusion methods degrade at long distances due to synchronization issues and noise. This paper addresses these gaps by proposing a robust CSI system for long-range, multi-user HRI.

Solution

  • Proposed approach: HiSync, an optical-inertial fusion framework that aligns robot-mounted camera optical flow with hand-worn IMU signals to identify command sources in long-range HRI.
  • Novelty:
    1. Introduction of a spectral-domain optical-inertial fusion framework for CSI.
    2. Development of CSINet with modules like Quality-Aware Feature Modulation, IMU-Anchored Cross-Modal Attention, and Scale-Aware Multi-Window Fusion.
    3. Creation of the first large-scale multimodal dataset for long-range CSI.
    4. Validation of HiSync on real-robot deployments in dynamic environments.
  • Procedure and key techniques:
    1. Extract spectral motion features from robot-mounted cameras and wearable IMUs.
    2. Use CSINet to align and match cross-modal features, leveraging quality-aware modulation, multi-window fusion, and attention mechanisms.
    3. Evaluate the system on curated datasets and real-world scenarios, including adversarial settings with mimics and bystanders.

Results

  • Concrete findings: HiSync achieves 97.82% accuracy from 3–34 m, outperforming the previous SOTA by 27.3%. At 34 m, HiSync maintains 94.31% accuracy compared to the baseline's 43.88%.
  • Advantage over baselines: HiSync outperforms vision-only and optical-inertial baselines, especially at long distances, with up to 26.30% higher accuracy than VIPL and robust performance under temporal noise.
  • Experiments / evaluation: Evaluated on a multimodal dataset (9465 sequences, 452,055 frames) and real-robot deployments. Metrics include CSI accuracy, response time, and subjective usability scores. Ablation studies confirm the importance of spectral features and multi-window fusion.
  • Limitations and future work: Limited ecological validity due to low crowd density in test environments; sensitivity to robot ego-motion; untested multi-IMU scenarios. Future work involves extending to dynamic robot motion, natural micro-gestures, and high-density venues.

Summary

HiSync introduces a novel optical-inertial fusion framework for robust command source identification in long-range human-robot interaction. By aligning wearable IMU signals with robot-mounted camera optical flow in the spectral domain, HiSync achieves high accuracy (up to 94.31% at 34 m) and outperforms existing methods. The system is validated on a large multimodal dataset and real-robot deployments, demonstrating usability and scalability in dynamic environments. Future work aims to address limitations such as crowd density and robot motion, expanding applicability to public spaces and multi-user scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223475/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790345
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
12 authors
sell
Subtopics
Teleoperation & Telepresence, Human Pose & Activity Recognition, Hand Gesture Recognition
work
Professions
Autonomous Driving Engineers & Test Drivers, Physicians, Nurses & Clinicians, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers