Gazeify Then Voiceify: Physical Object Referencing Through Gaze and Voice Interaction with Displayless Smart Glasses

Best Paper
Eye Tracking & Gaze InteractionVoice User Interface (VUI) DesignContext-Aware ComputingSoftware Engineers & DevelopersUI/UX DesignersAI/ML Researchers & Engineers

Smart glasses enhance interactions with the environment by using head-mounted cameras to observe the user’s viewpoint , but lack the visual feedback used for common interactions. We introduce "Gazeify then Voiceify", a multimodal approach allowing object selection via gaze and voice using displayless smart glasses. Users can select a physical object with their gaze, and the system generates a digital mask and a voice description of the object's semantics. Users can further correct errors through free-form conversation. To demonstrate our approach, we develop an interactive system by integrating advanced object segmentation and detection with a visual-language model. User studies reveal that participants achieve correct gaze selection in 53% of the task trials and use voice disambiguation to correct 58% remaining errors. Participants also rated the system as likable, useful and easy to use.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/226641/2026

AdRecommended

Learn AI Coding at CodeNow

At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2026
emoji_events
Award
Best Paper
group
Authors
9 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Voice User Interface (VUI) Design, Context-Aware Computing
work
Professions
Software Engineers & Developers, UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Abstract only
hub
Related Papers
3 related papers