Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue Integration

Full-Body Interaction & Embodied InputMixed Reality WorkspacesPhysical-Digital Hybrid InteractionUI/UX DesignersHCI Researchers

Paper Title

Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue Integration

Publication Info

  • Topic area: Mixed reality (MR) object selection using multimodal probabilistic inference.
  • Keywords: Mixed reality, object selection, probabilistic inference, pointing, grasping gestures, Bayesian integration, spatial ambiguity, semantic ambiguity, multimodal interaction.

Background and Problem

  • Problem / challenge: Existing MR object selection methods rely on single cues or deterministic multi-cue fusion, which degrade under spatial or semantic ambiguities (e.g., cluttered scenes or similar object shapes). No single cue reliably supports diverse MR scenarios.
  • Significance: Efficient and robust object selection is crucial for MR applications like 3D design, gaming, and collaborative tasks. Addressing ambiguities can improve usability and accuracy in such environments.
  • Motivation and related work: Prior work has explored directional cues (e.g., raycasting) and gestural cues (e.g., grasping), but both have limitations under specific ambiguities. Deterministic multi-cue approaches lack flexibility when dominant cues fail. This paper builds on these insights by introducing a probabilistic framework to integrate complementary cues.

Solution

  • Proposed approach: Point&Grasp, a probabilistic multi-cue integration framework combining directional pointing and grasping gestures via Bayesian inference.
  • Novelty:
    1. Introduction of a probabilistic framework for integrating spatial and semantic cues for MR object selection.
    2. Development of the Out-of-Reach Grasping (ORG) dataset, capturing in-reach and out-of-reach grasping gestures with annotated compatibility labels.
    3. Creation of a pose-aware gesture–object likelihood model trained on the ORG dataset.
    4. Validation of Point&Grasp through user studies, demonstrating superior performance over single-cue baselines and state-of-the-art methods.
  • Procedure and key techniques:
    1. Directional cues modeled as Gaussian likelihoods based on ray endpoints relative to object centers.
    2. Gestural cues modeled via a neural network estimating gesture–object compatibility using hand joint positions and object point clouds.
    3. Bayesian fusion of directional and gestural likelihoods to compute posterior probabilities for target inference.
    4. Implementation in VR with real-time feedback and a null gesture for cases without semantic cues.

Results

  • Concrete findings:
    • Point&Grasp achieved higher accuracy and faster selection times than single-cue baselines (Point, Grasp).
    • Outperformed BubbleRay and matched Expand in high-spatial ambiguity while avoiding Expand’s overhead in low-spatial layouts.
    • Demonstrated robust performance under both spatial and semantic ambiguities.
  • Advantage over baselines:
    • Faster selection times and higher completion rates compared to single-cue methods.
    • Comparable or superior performance to state-of-the-art techniques (BubbleRay, Expand) across various ambiguity conditions.
  • Experiments / evaluation:
    • Two user studies:
      1. Study 1 compared Point&Grasp with single-cue baselines under varying spatial and semantic ambiguities.
      2. Study 2 benchmarked Point&Grasp against BubbleRay and Expand.
    • Metrics: selection time, trial completion rate, and cue agreement analysis.
    • Participants: 12–19 users across studies, with controlled VR setups and randomized conditions.
  • Limitations and future work:
    • Limited dataset size and diversity; future work could expand the ORG dataset for better generalization.
    • Need for personalization of cue weighting for individual users.
    • Exploration of additional modalities (e.g., gaze, speech) and downstream tasks (e.g., object manipulation).

Summary

This paper introduces Point&Grasp, a probabilistic framework for out-of-reach object selection in MR, integrating directional pointing and grasping gestures via Bayesian inference. The method leverages complementary cues to address spatial and semantic ambiguities, achieving robust performance across diverse scenarios. User studies validated its superiority over single-cue baselines and competitive performance against state-of-the-art techniques. The work contributes a novel dataset (ORG) and a gesture–object likelihood model, paving the way for scalable multimodal interaction techniques in MR. Future research could enhance dataset diversity, personalization, and integration of additional modalities.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223410/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790836
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Full-Body Interaction & Embodied Input, Mixed Reality Workspaces, Physical-Digital Hybrid Interaction
work
Professions
UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers