Document Title

OmniSense: Exploring Novel Input Sensing and Interaction Techniques on Mobile Devices with an Omni-Directional Camera

Document Information

  • Domain: Human-Computer Interaction, multimodal sensing input technologies
  • Keywords: Omni-directional, 360° camera, input sensing, interaction technique, mobile devices, computer vision, virtual reality, user experience, gesture recognition, contextual awareness

Research Background and Problem Statement

  • Problem and Challenges: With the growing popularity of 360-degree cameras in the consumer market, their primary use remains focused on capturing images and video content. However, existing mobile devices face limitations in input sensing and interaction techniques, particularly lacking the ability to provide multimodal sensing capabilities within a single device.
  • Significance: The full field of view (FoV) of 360-degree cameras offers comprehensive visual information, including user finger movements, body postures, and surrounding environments. This opens up new possibilities for addressing user needs in mobile interactions, especially by extending input capabilities on general-purpose sensors.
  • Motivation and Related Work: Current technologies primarily rely on adding additional hardware to achieve specific functionalities, failing to deliver a "single-sensor multifunctional" solution. The authors propose leveraging built-in 360-degree cameras and computer vision techniques to explore their potential in input recognition and interaction design, addressing limitations of previous technologies.

Solution

  • Methods and Design Space: The authors propose the OmniSense framework, based on 360° cameras on mobile devices, exploring a sensing design space divided into three pillars:
    • Near Device Interaction: Detecting grip styles, finger touch positions, air hovering, and behind-device actions.
    • Around Device Interaction: Tracking body postures, limb movements, and the spatial relationship between the device and the user.
    • Surrounding Device and Context-Aware Sensing: Capturing environmental information, object detection, hazard awareness, etc.
  • Innovations:
    • Supports multimodal input sensing using a single sensor.
    • Operates flexibly without requiring specific directional cameras, leveraging computer vision techniques.
    • Practical applications span entertainment, user interface optimization, safety alerts, and more.
  • Implementation Steps and Key Technologies:
    1. Employ mainstream deep learning methods (e.g., EfficientNet, OpenPose) for sensing and object recognition.
    2. Collect and annotate data, combining depth information from depth cameras to train models.
    3. Implement screen capture and real-time video transmission to bypass restricted camera APIs.
    4. Post-process sensing data and develop specific application prototypes.

Research Outcomes

  • Specific Results:
    • Developed a real-time interaction system based on 360° cameras, encompassing functionalities such as finger tracking, body posture recognition, and environmental sensing.
    • Created 13 representative applications, including virtualized video conferencing, hazard warning systems, gesture-based gaming interactions, and dynamic user interface adjustments.
  • Advantages Compared to Existing Solutions:
    • Unlike existing technologies relying on specialized sensors, OmniSense achieves multimodal sensing capabilities using a single camera.
    • Provides real-time processing and user feedback at a lower cost than multi-sensor solutions.
  • Experimental or Evaluation Results:
    • Performance evaluations using deep learning methods achieved high accuracy, with grip recognition tasks reaching 99.55% accuracy and finger hover tracking precision exceeding 90% in 2D regression tasks.
    • User trials indicated that hazard warning and virtualized interaction applications were most favored, while finger-based near-device interactions received average ratings.
  • Limitations and Future Directions:
    • Limitations: Camera API restrictions remain unresolved; finger-based sensing techniques are considered impractical; real-time processing demands high mobile device performance.
    • Future Directions: Introduce more application scenarios for 360° cameras, such as vital sign monitoring and identity verification; optimize user interfaces and energy consumption; expand camera designs to reduce size.

Conclusion and Potential Impact

OmniSense demonstrates the immense potential of utilizing 360° cameras to achieve multimodal input sensing with a single sensor. Through comprehensive exploration of interaction design spaces, it provides new theoretical perspectives and technical solutions, serving as a valuable reference for researchers and manufacturers in the HCI field to develop novel mobile interaction technologies. The authors' work lays a solid foundation for future research while calling for further development of comprehensive device applications to address existing technological limitations.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95930/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3580747
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
10 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Immersion & Presence Research, 360° Video & Panoramic Content
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers