Situating Automatic Speech Recognition Development within Communities of Under-heard Language Speakers

Multilingual & Cross-Cultural Voice InteractionDeveloping Countries & HCI for Development (HCI4D)User Research Methods (Interviews, Surveys, Observation)

Title of the Paper

Situating Automatic Speech Recognition Development within Communities of Under-heard Language Speakers

Paper Information

  • Domain: Speech Recognition, Human-Computer Interaction (HCI), Computer Interaction
  • Keywords: Automatic Speech Recognition, Community Participation, Low-resource Languages, Mixed Languages, Speech Data Collection, Crowdsourced Transcription, Digital Equity, isiXhosa

Research Background and Problem

  • Problems and Challenges: Despite the expanding application of Automatic Speech Recognition (ASR) technology in mainstream languages, many underrepresented languages (e.g., isiXhosa) remain unsupported by these technologies. For instance, existing ASR systems rely on large amounts of high-quality speech data and precise manual transcription, often dominated by large corporations, which excludes these communities from the development process.
  • Significance of the Problem: Small language and multilingual mixed communities require tailored solutions to support their unique language usage patterns. However, existing unsupervised methods demand large datasets and are prohibitively expensive for small communities. Additionally, these methods may exacerbate the neglect of local data ownership, raising ethical and security concerns.
  • Research Motivation: To explore a community-driven approach for developing more equitable ASR systems that better support under-heard language communities and reflect their everyday language practices in speech technology.

Solution

  • Method/Tool Introduction:

    • Developed a "toolkit" comprising the following components: 1) SpeechBox, a public information device suitable for deployment in community spaces; 2) TranscriptTool, a mobile-based crowdsourced transcription application; 3) a feedback system, including an ASR model demonstrator, to share data usage outcomes with community members.
    • Data collection utilized localized hardware to incentivize community members to provide speech data, combined with feedback loops to showcase technological progress.
  • Innovations:

    • Embedded speech recognition development within communities to collect diverse, unstructured language data through effective public participation.
    • Designed a training and evaluation framework tailored to isiXhosa and its mixed-language characteristics, reducing reliance on standardized language norms.
    • Proposed a community feedback-driven ASR evaluation and optimization method, addressing the limitations of traditional metrics (gold standards and error rates).
  • Implementation Steps:

    1. Deployed SpeechBox in internet cafes and convenience stores in the Langa area of South Africa to collect community stories related to COVID-19.
    2. Mobilized community participation in speech transcription tasks using TranscriptTool to address challenges of multilingual mixing and noisy environments.
    3. Developed a multilingual ASR model integrating isiXhosa, English, and other South African languages, optimizing for code-switching and conversational transcription performance.
    4. Built an ASR demonstrator to allow community members to access the collected COVID-19 database via voice queries and rate system feedback.

Research Outcomes

  • Specific Results:

    • Collected 318 community stories with a total duration of nearly two hours, providing valuable data for studying natural language use in low-resource languages.
    • Reduced the ASR model's character error rate from 51.2% to 27.7%. While the error rate remains relatively high, the model supports the community's understanding of the "gist" in key contexts.
    • Developed TranscriptTool, which effectively enabled community members to perform accurate transcription of challenging audio under limited technological conditions.
    • The demonstrator showcased the potential of combining voice queries with semantic search. During the deployment, users completed 750 queries, with 71% returning results.
  • Advantages:

    • Speech recognition and digital participation were embedded in the dynamic language practices of the community, rather than being confined to laboratory dataset development.
    • The data collection, transcription, and model optimization processes highlighted the diversity of language use, better reflecting community needs.
    • Data governance and feedback loop mechanisms enhanced the community's agency in the technology development process.
  • Experiments and Evaluation Results:

    • During the deployment of the speech query demonstrator, 245 queries successfully returned matching results, with 66% of feedback rated as "relevant." This indicates community acceptance and recognition of the practical application of the ASR system.
    • Speech and transcription data revealed that the highly dynamic and non-standardized nature of isiXhosa poses significant challenges to defining traditional ASR "gold standards."
  • Limitations and Future Directions:

    • Limitations:
      • Background noise and multi-speaker recordings increased transcription difficulty.
      • Community members faced hardware constraints (e.g., old phones, cracked screens) when using the tools.
      • The current model still struggles with code-switching and language boundary handling.
    • Future Directions:
      • Develop and test more sophisticated code-switching algorithms for small languages like isiXhosa.
      • Conduct community-led ASR system optimization research to define more flexible and reliable "good enough" transcription standards for users.
      • Explore the scalability of these community-driven methods and tools in broader low-resource language contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96417/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581385
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
9 authors
sell
Subtopics
Multilingual & Cross-Cultural Voice Interaction, Developing Countries & HCI for Development (HCI4D), User Research Methods (Interviews, Surveys, Observation)
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
2 related papers