SIA: A Framework for Context-Aware Intent Clarification in Speech-Driven Immersive Analytics
Authors
The rise of generative AI has increased attention to voice interfaces. In immersive analytics, we conceptualize this trend as Speech-driven Immersive Analytics. While speech interfaces enable natural interactions, users, especially novices, still face a learning curve in articulating analytic intent and exploring data during the foraging phase. Prior work has primarily addressed these challenges through multimodal interaction or textual disambiguation. We introduce a context-aware Speech-driven Immersive Analytics framework (SIA) as a speech-oriented approach that leverages speech acts to convey actionable intent. This framework (SIA) was designed based on a formative study, a prototype development, three technical studies, and a user study. By extracting speech acts from utterances, SIA infers analytic tasks and embodiment tendencies, then integrates them with spatial, chart, and data context to generate feedforward: previews of potential actions and outcomes. The formative study identified user needs. The technical studies demonstrated that SIA improved the inference quality, enabling context-aware feedforward generation. The user study highlighted that the SIA-based prototype was responsive and intuitive, and feedforward helped users learn during the onboarding phase of data exploration. In particular, the user study identified which feedforward elements participants referenced and how they applied them when expressing intent in immersive analytics. Our key technical findings emphasize that the ensemble model, embedded in the Uncertainty Estimator, improves accuracy and stabilizes task inference. The Projector's context summary was critical in generating context-aware feedforward. Based on these results, we discuss future research directions for intelligent Speech-driven Immersive Analytics.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
🌳-generAItor: Tree-in-the-loop Text Generation for Language Model Explainability and Adaptation
IUI '25· Explainable AI (XAI) +2
- 63%
VoiceAlign: A Shimming Layer for Enhancing the Usability of Legacy Voice User Interface Systems
IUI '26· Voice User Interface (VUI) Design +2
Based on Jaccard similarity of research subtopics & professions (≥60%)