Multimodal Error Correction for Speech-to-Text in a Mobile Office Automated Vehicle: Results From a Remote Study
Authors
Automated Driving Interface & Takeover DesignVoice User Interface (VUI) DesignHuman-LLM CollaborationAutonomous Driving Engineers & Test DriversSoftware Engineers & DevelopersUI/UX Designers
Document Title
Multimodal Error Correction for Speech-to-Text in a Mobile Office Automated Vehicle: Results From a Remote Study
Document Information
- Subject Area: Human-Computer Interaction, specifically multimodal speech-to-text error correction techniques in autonomous vehicles
- Keywords: Speech-to-text, intelligent text input, error correction, multimodal navigation, user study, remote study, Wizard-of-Oz method, autonomous driving, mobile office
Research Background and Problem
- Problem or Challenge: This paper explores the impact of speech recognition error correction (i.e., the "correction problem") on user experience when performing office tasks via voice input in a future autonomous vehicle "mobile office" environment. Current voice input systems still face recognition error issues, causing users to spend significant time correcting errors.
- Significance: The proliferation of advanced autonomous vehicles (L3<L4<L5) makes "mobile office" scenarios possible, enhancing travel efficiency while meeting the demand for conducting non-driving-related tasks in vehicles. However, poor user experience (UX) could hinder technology acceptance and adoption.
- Research Motivation and Related Work:
- Voice as a primary input method reduces visual distractions while driving but may increase cognitive load.
- Although some studies have explored the potential of multimodal interaction to improve speech-to-text recognition accuracy and speed, research on users' subjective experiences remains limited.
Solution
- Proposed Method or Solution: Introducing multimodal approaches in autonomous vehicles to address speech recognition errors. Specifically, three interaction modes were compared:
- Voice-only (Baseline Condition): Activating the voice assistant using keywords and correcting errors via voice commands.
- Voice and Touchpad (VaT): Navigating to select errors using a touchpad, then specifying corrections via voice.
- Voice and Gestures (VaG): Using gestures (e.g., swiping left or right with the palm to select words), clenching a fist to confirm selection, and then correcting via voice.
- Innovative Aspects:
- Introducing touchpads and mid-air gestures as multimodal interactions combined with voice, aligning with the existing layout of in-vehicle devices.
- Studying real user experiences through a remote "Wizard-of-Oz" method, mitigating health risks during the pandemic.
- Implementation Steps and Key Technologies:
- The experiment was conducted using remote video conferencing software (ZOOM) and interactive click prototypes.
- A user study process was designed with three experimental conditions, employing the User Experience Questionnaire Short Version (UEQ-S), Technology Acceptance Model (TAM), and semi-structured interviews for data collection and analysis.
Research Findings
- Specific Findings:
- The gesture mode (VaG) scored highest in hedonic quality (aesthetics and enjoyment) but lower in pragmatic quality (practicality).
- The combination of voice and touchpad (VaT) was perceived as fast and easy to use, making it the most preferred mode among participants.
- The voice-only mode performed comparably to VaT in terms of technology acceptance and task ease of use.
- Advantages:
- VaT combines the benefits of voice interaction and precise target selection, improving task efficiency for error correction.
- VaG demonstrated the potential of gesture interaction for enhancing enjoyment, which could be valuable in entertainment or non-task-oriented scenarios.
- Experimental or Evaluation Results:
- Most participants preferred the voice-and-touchpad mode.
- Interviewed users prioritized practicality for productive tasks but might favor the playful nature of gestures in entertainment settings.
- The study effectively collected data remotely, avoiding health risks, but requires high-fidelity setups for further validation.
- Limitations and Future Directions:
- The fidelity of remote studies was limited, as users could not experience real tactile feedback or driving simulation.
- The study did not deeply explore the performance of the technology under emergency takeover scenarios in highly autonomous vehicles.
- Future research should validate findings in driving simulation environments and investigate user experiences in complex voice task scenarios (e.g., text formatting).
Conclusion and Outlook
- This study confirms that a multimodal approach combining touchpads and voice is more accepted in productive mobile office scenarios than single-mode interactions. It also validates the potential entertainment applications of gesture interaction. In the future, high-fidelity testing platforms and broader user group evaluations will be essential to enrich research conclusions and drive application development. Additionally, the socio-technical impacts of mobile office scenarios, such as privacy and noise pollution, require further discussion.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How do multimodal interaction modalities affect UX in speech-to-text error correction during mobile work in autonomous vehicles?Category: Mobile Touch and Micro-Gesture InputSimilar questionsarrow_forward
- Is combined voice and touchpad interaction more efficient than voice alone for speech-to-text error correction?Category: Mobile Touch and Micro-Gesture InputSimilar questionsarrow_forward
- Can voice and gesture interaction improve users' interaction enjoyment in mobile work scenarios?Category: Mobile Touch and Micro-Gesture InputSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Users are inefficient at correcting speech recognition errors during mobile work in autonomous vehicles.Category: Mobile Touch and Micro-Gesture InputSimilar questionsarrow_forward
- 67%
Content-Driven Local Response: Supporting Sentence-Level and Message-Level Mobile Email Replies With and Without AI
CHI '25· Voice User Interface (VUI) Design +1
- 67%
Vajra: Step-by-step programming with natural language
IUI '19· Voice User Interface (VUI) Design +1
- 67%
ILuvUI: Instruction-tuned LangUage-Vision modeling of UIs from Machine Conversations
IUI '25· Voice User Interface (VUI) Design +1
- 67%
Automated Driving System HMIs: “Clear and Unambiguous”?
AutoUI '24· Automated Driving Interface & Takeover Design
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3490099.3511131
At a Glance
fact_checkPaper Snapshot
dataset
Source
IUI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Automated Driving Interface & Takeover Design, Voice User Interface (VUI) Design, Human-LLM Collaboration
work
Professions
Autonomous Driving Engineers & Test Drivers, Software Engineers & Developers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers