Identifying an Aurally Distinct Phrase Set for Text Entry Techniques

Honorable Mention
Voice AccessibilityAugmentative & Alternative Communication (AAC)Assistive Technology SpecialistsAmazon Mechanical Turk Workers

Document Title

Identifying an Aurally Distinct Phrase Set for Text Entry Techniques

Document Information

  • Subject Area: Human-Computer Interaction, Non-Visual Text Entry Systems, Text-to-Speech Technology
  • Keywords: text entry, accessibility, text-to-speech (TTS), synthesized speech, input techniques, phrases corpus

Research Background and Problem

  • Identified Issues or Challenges:

    • For blind or low-vision users, traditional text entry tasks require visual interaction, limiting accessibility.
    • Existing phrase sets used in text entry research do not ensure clarity when provided via speech.
    • There is a lack of a standardized phrase corpus to test non-visual text entry techniques.
  • Significance:

    • Non-visual text entry techniques enable blind or low-vision users to input text more efficiently on portable devices.
    • An accurate phrase corpus is fundamental for evaluating the performance of such techniques and designing solutions to enhance usability.
  • Research Motivation and Related Work:

    • Investigating how to create or identify aurally distinct phrase sets for use in the non-visual text entry domain.
    • Building on existing text entry phrase sets (e.g., MacKenzie and Soukoref's 500-phrase set) to optimize their intelligibility in speech scenarios.

Solution

  • Proposed Method or Solution:

    • The research team utilized an existing phrase set, generating audio files using text-to-speech systems from Microsoft and Apple.
    • Listening and transcription experiments were conducted with participants recruited from Amazon Mechanical Turk and the target population (blind or low-vision users) to identify aurally distinct phrases.
  • Innovations:

    • Proposed a data-driven phrase selection method based on speech generation and human participation.
    • Identified error types in phrases (e.g., spelling errors, spacing errors, comprehension errors) to establish guidelines for future corpus design.
    • Developed a corpus of 92 phrases with low comprehension error rates.
  • Implementation Steps and Techniques:

    • Speech Generation: Generated audio files for 500 phrases using Apple macOS and Microsoft Windows text-to-speech technologies, with both male and female voices.
    • Experiment 1:
      • Recruited participants from Amazon Mechanical Turk, collecting approximately 17,421 transcription data points.
      • Analyzed errors using Levenshtein distance and manually categorized error types.
      • Identified 96 phrases with no comprehension errors.
    • Experiment 2:
      • Recruited 80 blind or low-vision users for auditory testing.
      • Validated the auditory clarity of the phrases, resulting in a final set of 92 phrases with no comprehension errors.

Research Outcomes

  • Specific Outcomes:

    • Identified a set of 92 aurally distinct phrases suitable for non-visual text entry research.
    • Proposed nine transcription error categories, including comprehension errors, spelling errors, extra space errors, numeric errors, etc., and established classification standards.
  • Advantages Over Existing Solutions:

    • Compared to the original phrase set, the selected phrases significantly reduced comprehension errors in speech transcription.
    • Provided a replicable method for speech generation and error classification to support future corpus expansion and applied research.
  • Experimental or Evaluation Results:

    • First Experiment: 96 phrases showed no comprehension errors across multiple speech versions, with an average accuracy rate as high as 77.51%.
    • Second Experiment: 92 phrases were validated by blind or low-vision users, with accuracy rates significantly improved and error rates reduced to 23.85%.
  • Limitations and Future Directions:

    • Limitations:
      • Small participant sample size may not represent broader user behavior.
      • The identified corpus is limited to English and needs to support other languages in the future.
      • Manual coding may introduce subjective bias.
    • Future Directions:
      • Develop broader, cross-linguistic aurally distinct phrase corpora.
      • Explore the impact of different speech configurations in text-to-speech technology on user comprehension.
      • Test the new phrases in specific devices and non-visual interaction scenarios, such as optimizing speech rate adjustments.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68769/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501897
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Voice Accessibility, Augmentative & Alternative Communication (AAC)
work
Professions
Assistive Technology Specialists, Amazon Mechanical Turk Workers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers