Don’t Look Now: Audio/Haptic Guidance for 3D Scanning of Landmarks

Honorable Mention
In-Vehicle Haptic, Audio & Multimodal FeedbackVibrotactile Feedback & Skin StimulationInteractive Data VisualizationContext-Aware Computing

Document Title

Don’t Look Now: Audio/Haptic Guidance for 3D Scanning of Landmarks

Document Information

  • Topic Area: User experience-driven 3D scanning technology, integrating audio and haptic feedback to optimize object scanning in public environments
  • Keywords: 3D scanning, guidance, audio feedback, haptic feedback, visual feedback, augmented reality (AR), user engagement, user safety, mesh reconstruction accuracy, multimodal feedback

Research Background and Problem

  • Problem or Challenge:
    • Users face challenges such as keeping the scanning target centered in the frame and avoiding walking disruptions when using smartphones to scan objects or landmarks in public spaces.
    • Traditional screen-based feedback, primarily video guidance, may increase user stress, complicate interactions, and cause discomfort or privacy concerns in public settings.
  • Significance:
    • With the proliferation of smartphones equipped with LIDAR and real-time multi-view stereo computation, non-expert users have rising demands and expectations for high-quality 3D scanning. These scans can be applied in fields such as virtual reality (VR), archaeology, medical education, and cultural heritage preservation.
  • Research Motivation and Related Work:
    • Current 3D scanning support tools are not yet mature in terms of usability and handling high-traffic environments. For example, visual feedback often lacks immediacy and suffers from drift errors.
    • This research aims to rethink feedback modalities by integrating multimodal feedback (especially audio and haptic) to reduce users' reliance on screens during scanning.

Solution

  • Proposed Solution:
    • Designed and developed a 3D scanning application that incorporates audio and haptic guidance, optimizing user interaction through real-time audio prompts and vibration feedback.
    • Improved initialization algorithms and anti-drift algorithms, including a real-time image segmentation model based on convolutional neural networks (e.g., Google’s MagicTouch technology) and 3D bounding box estimation.
  • Innovations:
    • Introduced audio and haptic guidance modes to replace screen feedback, enabling users to perform “eyes-free” operations during scanning and avoid visual distractions.
    • Implemented the Landmark Tracker drift correction algorithm to significantly reduce real-time scanning frame drift issues.
    • Enhanced the integration of algorithms and interface design, such as adapting dynamic audio rhythms to screen changes and real-time adjustments to vibration frequency and intensity.
  • Implementation Steps and Techniques:
    • The user interface initiates intelligent segmentation by tapping on the target object to generate a 3D bounding box.
    • During the process, audio pitch and haptic intensity dynamically adjust to guide users in maintaining optimal scanning posture.
    • Finally, a local mesh is generated, allowing users to zoom and rotate to view the resulting POI model.

Research Outcomes

  • Specific Results:
    • Compared to traditional visual guidance, the audio/haptic guidance interface significantly improved user engagement (particularly in terms of user experience and reward feedback), extended scanning duration, and produced higher-quality meshes.
    • The application designed with solely audio/haptic feedback allowed users to maintain better awareness of their surroundings in outdoor scanning scenarios, reducing social evaluation pressure (e.g., being perceived as invading privacy by onlookers).
  • Advantages:
    • Scanning under visual screen guidance resulted in lower mesh accuracy, whereas audio/haptic guidance produced higher-quality and more complete 3D meshes.
    • Significantly reduced drift issues, with drift occurrence decreasing from 83% to 18%.
  • Limitations and Future Directions:
    • The audio/haptic guidance mode did not significantly reduce overall user stress, and some users reported needing to grip the device more tightly due to the lack of direct visual feedback.
    • Testing conditions did not include the simultaneous use of audio, haptic, and visual multimodal feedback.
    • Outdoor environmental variables (e.g., crowd density) may potentially impact user experience.
    • Future research directions could include comparing effectiveness across different scenarios (e.g., large buildings, indoor spaces) and studying long-term adaptation to the guidance modes.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147615/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642271
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
In-Vehicle Haptic, Audio & Multimodal Feedback, Vibrotactile Feedback & Skin Stimulation, Interactive Data Visualization, Context-Aware Computing
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
0 related papers