Multimodal Error Correction for Speech-to-Text in a Mobile Office Automated Vehicle: Results From a Remote Study

Automated Driving Interface & Takeover DesignVoice User Interface (VUI) DesignHuman-LLM CollaborationAutonomous Driving Engineers & Test DriversSoftware Engineers & DevelopersUI/UX Designers

Document Title

Multimodal Error Correction for Speech-to-Text in a Mobile Office Automated Vehicle: Results From a Remote Study

Document Information

  • Subject Area: Human-Computer Interaction, specifically multimodal speech-to-text error correction techniques in autonomous vehicles
  • Keywords: Speech-to-text, intelligent text input, error correction, multimodal navigation, user study, remote study, Wizard-of-Oz method, autonomous driving, mobile office

Research Background and Problem

  • Problem or Challenge: This paper explores the impact of speech recognition error correction (i.e., the "correction problem") on user experience when performing office tasks via voice input in a future autonomous vehicle "mobile office" environment. Current voice input systems still face recognition error issues, causing users to spend significant time correcting errors.
  • Significance: The proliferation of advanced autonomous vehicles (L3<L4<L5) makes "mobile office" scenarios possible, enhancing travel efficiency while meeting the demand for conducting non-driving-related tasks in vehicles. However, poor user experience (UX) could hinder technology acceptance and adoption.
  • Research Motivation and Related Work:
    • Voice as a primary input method reduces visual distractions while driving but may increase cognitive load.
    • Although some studies have explored the potential of multimodal interaction to improve speech-to-text recognition accuracy and speed, research on users' subjective experiences remains limited.

Solution

  • Proposed Method or Solution: Introducing multimodal approaches in autonomous vehicles to address speech recognition errors. Specifically, three interaction modes were compared:
    1. Voice-only (Baseline Condition): Activating the voice assistant using keywords and correcting errors via voice commands.
    2. Voice and Touchpad (VaT): Navigating to select errors using a touchpad, then specifying corrections via voice.
    3. Voice and Gestures (VaG): Using gestures (e.g., swiping left or right with the palm to select words), clenching a fist to confirm selection, and then correcting via voice.
  • Innovative Aspects:
    • Introducing touchpads and mid-air gestures as multimodal interactions combined with voice, aligning with the existing layout of in-vehicle devices.
    • Studying real user experiences through a remote "Wizard-of-Oz" method, mitigating health risks during the pandemic.
  • Implementation Steps and Key Technologies:
    • The experiment was conducted using remote video conferencing software (ZOOM) and interactive click prototypes.
    • A user study process was designed with three experimental conditions, employing the User Experience Questionnaire Short Version (UEQ-S), Technology Acceptance Model (TAM), and semi-structured interviews for data collection and analysis.

Research Findings

  • Specific Findings:
    • The gesture mode (VaG) scored highest in hedonic quality (aesthetics and enjoyment) but lower in pragmatic quality (practicality).
    • The combination of voice and touchpad (VaT) was perceived as fast and easy to use, making it the most preferred mode among participants.
    • The voice-only mode performed comparably to VaT in terms of technology acceptance and task ease of use.
  • Advantages:
    • VaT combines the benefits of voice interaction and precise target selection, improving task efficiency for error correction.
    • VaG demonstrated the potential of gesture interaction for enhancing enjoyment, which could be valuable in entertainment or non-task-oriented scenarios.
  • Experimental or Evaluation Results:
    • Most participants preferred the voice-and-touchpad mode.
    • Interviewed users prioritized practicality for productive tasks but might favor the playful nature of gestures in entertainment settings.
    • The study effectively collected data remotely, avoiding health risks, but requires high-fidelity setups for further validation.
  • Limitations and Future Directions:
    • The fidelity of remote studies was limited, as users could not experience real tactile feedback or driving simulation.
    • The study did not deeply explore the performance of the technology under emergency takeover scenarios in highly autonomous vehicles.
    • Future research should validate findings in driving simulation environments and investigate user experiences in complex voice task scenarios (e.g., text formatting).

Conclusion and Outlook

  • This study confirms that a multimodal approach combining touchpads and voice is more accepted in productive mobile office scenarios than single-mode interactions. It also validates the potential entertainment applications of gesture interaction. In the future, high-fidelity testing platforms and broader user group evaluations will be essential to enrich research conclusions and drive application development. Additionally, the socio-technical impacts of mobile office scenarios, such as privacy and noise pollution, require further discussion.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/79996/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3490099.3511131
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Automated Driving Interface & Takeover Design, Voice User Interface (VUI) Design, Human-LLM Collaboration
work
Professions
Autonomous Driving Engineers & Test Drivers, Software Engineers & Developers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers