Geno: A Developer Tool for Authoring Multimodal Interaction on Existing Web Applications
Authors
Supporting voice commands in applications presents significant benefits to users. However, adding such support to existing GUIbased web apps is effort-consuming with a high learning barrier, as shown in our formative study, due to the lack of unified support for creating multimodal interfaces. We present Geno—a developer tool for adding the voice input modality to existing web apps without requiring significant NLP expertise. Geno provides a high-level workflow for developers to specify functionalities to be supported by voice (intents), create language models for detecting intents and the relevant information (parameters) from user utterances, and fulfill the intents by either programmatically invoking the corresponding functions or replaying GUI actions on the web app. Geno further supports multimodal references to GUI context in voice commands (e.g., “move this [event] to next week” while pointing at an event with the cursor). In a study, developers with little NLP expertise were able to add multimodal voice command support for two existing web apps using Geno.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
Using Bayes' Theorem for Command Input: Principle, Models, and Applications
CHI '20· Voice User Interface (VUI) Design +1
- 67%
Enabling Conversational Interaction with Mobile UI using Large Language Models
CHI '23· Voice User Interface (VUI) Design +1
- 67%
Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning
UIST '21· Voice User Interface (VUI) Design +1
- 60%
What’s The Talk on VUI Guidelines? A Meta-Analysis of Guidelines for Voice User Interface Design
CUI '23· Voice User Interface (VUI) Design
Based on Jaccard similarity of research subtopics & professions (≥60%)