VoiceAlign: A Shimming Layer for Enhancing the Usability of Legacy Voice User Interface Systems
Authors
Voice user interfaces (VUIs) are rapidly transitioning from accessibility features to mainstream interaction modalities. Yet most operating systems' built-in voice commands remain underutilized despite possessing robust technical capabilities. Through our analysis of four commercial VUI systems and a formative study with 16 participants, we found that fixed command formats require exact phrasing, restrictive timeout mechanisms discard input during planning pauses, and insufficient feedback hampers multi-step interactions. To address these challenges, we developed VoiceAlign, an adaptive shimming layer that mediates between users and legacy VUI systems. VoiceAlign intercepts natural voice commands, transforms them to match the required syntax using a large language model, and transmits these adapted commands through a virtual audio channel that remains transparent to the underlying system. In our evaluation with 12 participants, VoiceAlign reduced command failures by half, required 25% fewer commands per task, and significantly lowered cognitive and temporal demands when paired with an existing legacy VUI system. Furthermore, we created a synthetic dataset informed by our studies and fine-tuned a small language model that achieves over 90% accuracy with 200 ms response time when served locally, eliminating dependence on third-party APIs while enabling real-time interaction on edge devices. This work demonstrates how modern AI techniques can unlock the underutilized potential of legacy VUI systems without requiring system modifications, offering a practical solution without replacing existing infrastructure.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Enabling Conversational Interaction with Mobile UI using Large Language Models
CHI '23· Voice User Interface (VUI) Design +1
- 83%
ONYX: Assisting Users in Teaching Natural Language Interfaces Through Multi-Modal Interactive Task Learning
CHI '23· Voice User Interface (VUI) Design +2
- 83%
Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning
UIST '21· Voice User Interface (VUI) Design +1
- 71%
Planning for Natural Language Failures with the AI Playbook
CHI '21· Human-LLM Collaboration +2
- 71%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 71%
ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
CHI '24· Voice User Interface (VUI) Design +2
- 71%
The AI Memory Gap: Users Misremember What They Created With AI or Without
CHI '26· Human-LLM Collaboration +2
- 71%
CoAutoML: User Interface Framework for Machine Learning Novices using LLM-based AutoML and Test-Driven Machine Teaching
IUI '26· AutoML Interfaces +2
- 67%
Validating AI-Generated Code with Live Programming
CHI '24· Human-LLM Collaboration +1
- 67%
Content-Driven Local Response: Supporting Sentence-Level and Message-Level Mobile Email Replies With and Without AI
CHI '25· Voice User Interface (VUI) Design +1
Based on Jaccard similarity of research subtopics & professions (≥60%)