Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue Integration
Authors
Paper Title
Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue Integration
Publication Info
- Topic area: Mixed reality (MR) object selection using multimodal probabilistic inference.
- Keywords: Mixed reality, object selection, probabilistic inference, pointing, grasping gestures, Bayesian integration, spatial ambiguity, semantic ambiguity, multimodal interaction.
Background and Problem
- Problem / challenge: Existing MR object selection methods rely on single cues or deterministic multi-cue fusion, which degrade under spatial or semantic ambiguities (e.g., cluttered scenes or similar object shapes). No single cue reliably supports diverse MR scenarios.
- Significance: Efficient and robust object selection is crucial for MR applications like 3D design, gaming, and collaborative tasks. Addressing ambiguities can improve usability and accuracy in such environments.
- Motivation and related work: Prior work has explored directional cues (e.g., raycasting) and gestural cues (e.g., grasping), but both have limitations under specific ambiguities. Deterministic multi-cue approaches lack flexibility when dominant cues fail. This paper builds on these insights by introducing a probabilistic framework to integrate complementary cues.
Solution
- Proposed approach: Point&Grasp, a probabilistic multi-cue integration framework combining directional pointing and grasping gestures via Bayesian inference.
- Novelty:
- Introduction of a probabilistic framework for integrating spatial and semantic cues for MR object selection.
- Development of the Out-of-Reach Grasping (ORG) dataset, capturing in-reach and out-of-reach grasping gestures with annotated compatibility labels.
- Creation of a pose-aware gesture–object likelihood model trained on the ORG dataset.
- Validation of Point&Grasp through user studies, demonstrating superior performance over single-cue baselines and state-of-the-art methods.
- Procedure and key techniques:
- Directional cues modeled as Gaussian likelihoods based on ray endpoints relative to object centers.
- Gestural cues modeled via a neural network estimating gesture–object compatibility using hand joint positions and object point clouds.
- Bayesian fusion of directional and gestural likelihoods to compute posterior probabilities for target inference.
- Implementation in VR with real-time feedback and a null gesture for cases without semantic cues.
Results
- Concrete findings:
- Point&Grasp achieved higher accuracy and faster selection times than single-cue baselines (Point, Grasp).
- Outperformed BubbleRay and matched Expand in high-spatial ambiguity while avoiding Expand’s overhead in low-spatial layouts.
- Demonstrated robust performance under both spatial and semantic ambiguities.
- Advantage over baselines:
- Faster selection times and higher completion rates compared to single-cue methods.
- Comparable or superior performance to state-of-the-art techniques (BubbleRay, Expand) across various ambiguity conditions.
- Experiments / evaluation:
- Two user studies:
- Study 1 compared Point&Grasp with single-cue baselines under varying spatial and semantic ambiguities.
- Study 2 benchmarked Point&Grasp against BubbleRay and Expand.
- Metrics: selection time, trial completion rate, and cue agreement analysis.
- Participants: 12–19 users across studies, with controlled VR setups and randomized conditions.
- Two user studies:
- Limitations and future work:
- Limited dataset size and diversity; future work could expand the ORG dataset for better generalization.
- Need for personalization of cue weighting for individual users.
- Exploration of additional modalities (e.g., gaze, speech) and downstream tasks (e.g., object manipulation).
Summary
This paper introduces Point&Grasp, a probabilistic framework for out-of-reach object selection in MR, integrating directional pointing and grasping gestures via Bayesian inference. The method leverages complementary cues to address spatial and semantic ambiguities, achieving robust performance across diverse scenarios. User studies validated its superiority over single-cue baselines and competitive performance against state-of-the-art techniques. The work contributes a novel dataset (ORG) and a gesture–object likelihood model, paving the way for scalable multimodal interaction techniques in MR. Future research could enhance dataset diversity, personalization, and integration of additional modalities.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
"In VR, everything is possible!": Sketching and Simulating Spatially-Aware Interactive Spaces in Virtual Reality
CHI '20· Full-Body Interaction & Embodied Input +2
- 67%
HOICraft: In-Situ VLM-based Authoring Tool for Part-Level Hand-Object Interaction Design in VR
CHI '26· Full-Body Interaction & Embodied Input +2
- 67%
Beyond Links: Exploring Visual Representations of Multi-View Relations in Mixed Reality
CHI '26· Mixed Reality Workspaces +2
- 67%
Unbounded: Object-Boundary Interactions in Mixed Reality
CHI '26· Mixed Reality Workspaces +2
- 67%
Draped Surfaces: A Contour-Adaptive Interface Overlaid on the Physical Environment for Mixed Reality Workspaces
CHI '26· Mixed Reality Workspaces +1
- 67%
GestureCanvas: A Programming by Demonstration System for Prototyping Compound Freehand Interaction in VR
UIST '23· Hand Gesture Recognition +2
- 67%
TwinSpin: A Virtual Ball in a VR Controller Enabling In-Hand 3DoF Rotation
UIST '25· Shape-Changing Interfaces & Soft Robotic Materials +2
- 60%
Evaluating the Combination of Visual Communication Cues for HMD-based Mixed Reality Remote Collaboration
CHI '19· Mixed Reality Workspaces
- 60%
GraV: Grasp Volume Data for the Design of One-Handed XR Interfaces
DIS '24· Full-Body Interaction & Embodied Input +1
Based on Jaccard similarity of research subtopics & professions (≥60%)