Gazeify Then Voiceify: Physical Object Referencing Through Gaze and Voice Interaction with Displayless Smart Glasses
Best PaperAuthors
Smart glasses enhance interactions with the environment by using head-mounted cameras to observe the user’s viewpoint , but lack the visual feedback used for common interactions. We introduce "Gazeify then Voiceify", a multimodal approach allowing object selection via gaze and voice using displayless smart glasses. Users can select a physical object with their gaze, and the system generates a digital mask and a voice description of the object's semantics. Users can further correct errors through free-form conversation. To demonstrate our approach, we develop an interactive system by integrating advanced object segmentation and detection with a visual-language model. User studies reveal that participants achieve correct gaze selection in 53% of the task trials and use voice disambiguation to correct 58% remaining errors. Participants also rated the system as likable, useful and easy to use.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
Enhancing Mobile Voice Assistants with WorldGaze
CHI '20· Eye Tracking & Gaze Interaction +1
- 67%
Stretch Gaze Targets Out: Experimenting with Target Sizes for Gaze-Enabled Interfaces on Mobile Devices
CHI '25· Eye Tracking & Gaze Interaction +1
- 63%
Understanding Gaze-Based Identification in VR Through Preattentive Processing and Binocular Rivalry
CHI '26· Eye Tracking & Gaze Interaction +3
Based on Jaccard similarity of research subtopics & professions (≥60%)