BeParrot: Efficient Interface for Transcribing Unclear Speech via Respeaking

Intelligent Voice Assistants (Alexa, Siri, etc.)Conversational ChatbotsPhysicians, Nurses & CliniciansAmazon Mechanical Turk Workers

Title of the Paper

BeParrot: Efficient Interface for Transcribing Unclear Speech via Respeaking

Paper Information

  • Research Areas: Artificial Intelligence, Human-Computer Interaction, Speech Recognition, and Transcription Technologies
  • Keywords: Respeaking, Speech Transcription, Automatic Speech Recognition (ASR), Human-Computer Interaction (HCI), Error Correction, Transcription Tools, Dictation Interface Optimization, Noisy Speech Processing, User Studies, Echo Speech

Research Background and Problem Statement

  • Identified Issues or Challenges: Current speech transcription methods, such as post-correction approaches based on Automatic Speech Recognition (ASR), require clear speech input and perform poorly with unclear speech (e.g., in high-noise or reverberant environments). While Respeaking technology is effective, it imposes high cognitive demands on operators, requiring professional training and making it difficult for novice users to adopt.
  • Significance of the Research: Efficient transcription of unclear speech is crucial for enhancing video accessibility, training speech models, and analyzing language production. However, optimization research targeting such speech scenarios remains limited.
  • Motivation and Related Work: By introducing traditional Respeaking methods into speech transcription tasks and designing a novel user interface, the authors aim to lower the barrier to using Respeaking, enabling novice users to efficiently participate in the transcription of unclear speech.

Solution

  • Proposed Solution or Method:
    • Develop an interface called BeParrot, which incorporates two key features—“parameter adjustment” and “pronunciation feedback”—to assist novice users in completing Respeaking tasks.
    • Parameter Adjustment: Dynamically adjusts the playback speed and length of speech segments based on user interaction history.
    • Pronunciation Feedback: Provides users with phoneme-level suggestions based on error records to help improve their pronunciation.
  • Innovations:
    1. Introduced automatic adjustment and customizable playback speed and segment length functionality.
    2. Provided real-time pronunciation feedback to help users refine their speaking habits.
    3. Applied Respeaking to various unclear speech scenarios and successfully reduced task complexity.
  • Implementation Steps and Techniques:
    • Task segmentation: Break speech data into short segments for users to transcribe step-by-step.
    • Integration of two ASR models: An offline model for generating reference results from complete audio and a streaming ASR model for real-time Respeaking transcription.
    • Interface design includes speech playback, parameter adjustment sliders, real-time recognition result display, and error feedback tools.

Research Outcomes

  • Specific Results:
    • Using BeParrot, users achieved a 32.2% increase in transcription speed for unclear speech (e.g., noisy or reverberant audio) without compromising accuracy.
    • The study demonstrated that even novice users could efficiently perform Respeaking tasks with BeParrot without requiring extensive training.
  • Advantages Over Existing Solutions:
    • Compared to traditional post-correction methods, BeParrot significantly reduces transcription time for unclear speech.
    • Helps novice users overcome the high cognitive burden of Respeaking.
    • Enhances user experience and efficiency through personalized adjustment and feedback features.
  • Experimental or Evaluation Results:
    • The 60 participants in the experiment showed significant reductions in transcription time using BeParrot, particularly in noisy or historical recording scenarios.
    • Experimental results confirmed the effectiveness of the parameter adjustment feature in optimizing playback speed and segment length.
  • Limitations and Future Directions:
    • The current approach relies on high-accuracy ASR models and has limited support for low-resource languages.
    • It has yet to be tested on more complex speech data, such as multi-speaker dialogues or long-duration recordings.
    • Future research could explore adaptive capabilities of online ASR models to further enhance customization and optimize user feedback display mechanisms.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/79976/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3490099.3511164
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Conversational Chatbots
work
Professions
Physicians, Nurses & Clinicians, Amazon Mechanical Turk Workers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers