Comparison of Methods for Evaluating Complexity of Simplified Texts among Deaf and Hard-of-Hearing Adults at Different Literacy Levels

Voice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)Speech-Language Pathologists & AudiologistsSpecial Education TeachersSocial Workers

Title of the Paper

Comparison of Methods for Evaluating Complexity of Simplified Texts among Deaf and Hard-of-Hearing Adults at Different Literacy Levels

Paper Information

  • Subject Area: Accessibility Design and Text Simplification Research
  • Keywords: Automatic Text Simplification, Methodological Research, Deaf and Hard-of-Hearing, Accessibility Design, Text Complexity Evaluation, Reading Comprehension, Response Scoring, Subjective Evaluation, Low Literacy, Information Technology

Research Background and Problem

  • Issues or Challenges Identified by the Authors:

    • The Deaf and Hard-of-Hearing (DHH) population exhibits diverse reading abilities, and low literacy levels may lead to difficulties in reading complex texts.
    • Automatic Text Simplification (ATS), as an assistive reading technology, benefits DHH users, but previous research lacks guidance on how to evaluate the complexity of simplified texts for this group.
    • Common metrics in text complexity evaluation (e.g., reading speed, comprehension questions, subjective judgments) have not been validated, making their applicability to DHH users uncertain.
  • Why This Problem is Important:

    • Evaluating the quality of simplified texts is crucial for developing assistive reading tools tailored to DHH users.
    • Without reliable text complexity evaluation, it is challenging to improve the accuracy and user experience of assistive technologies.
  • Motivation and Related Work:

    • The authors reviewed existing studies on DHH users, language technology evaluation methods, and ATS literature, finding that many studies lack explicit methods for evaluating text complexity specific to the DHH population.
    • Drawing on prior experience in evaluating captioning tools and American Sign Language (ASL) animations, this study aims to establish a more reliable framework for ATS evaluation.

Solution

  • Proposed Method or Solution:

    • Compare news articles of varying complexity levels (manually simplified versions) to ensure high-quality corpora.
    • Test multiple evaluation metrics, including objective metrics (e.g., reading speed, success rate in answering comprehension questions) and subjective metrics (e.g., readability and comprehensibility assessments).
  • Innovative Aspects:

    • Identify which evaluation metrics effectively distinguish between texts of different complexity levels, particularly for DHH users with varying literacy levels.
    • Propose grouping readers (based on literacy levels) and controlling experimental design to reveal statistically significant differences in metric effectiveness.
  • Implementation Steps and Key Techniques:

    1. Corpus Selection: Articles were selected from the Newsela platform, which provides news at different complexity levels, ensuring uniform standards for length.
    2. Experimental Metrics:
      • Reading speed (calculated as word count/reading time).
      • Comprehension questions (main factual questions, categorized as high-complexity and low-complexity questions).
      • Score prediction (participants predict their own scores on comprehension questions).
      • Subjective evaluation: including judgments on text comprehensibility and readability.
    3. Method Validation: Participants were divided into high and low literacy groups (WRAT-H and WRAT-L) based on the Wide Range Achievement Test (WRAT).
    4. Experimental Design: A Latin square design was used to randomize article complexity and question order, reducing experimental bias.

Research Findings

  • Specific Findings:

    • Subjective metrics (especially readability evaluation) performed best: This metric effectively distinguished text complexity regardless of participants' literacy levels.
    • Low-complexity comprehension questions performed better in the low-literacy group: They could only differentiate between the lowest and highest complexity conditions.
    • High-complexity comprehension questions and reading speed had limited effectiveness: These are not suitable as core metrics for evaluating text complexity.
  • Comparison with Existing Solutions:

    • Unlike studies that rely solely on objective metrics, this research highlights the superiority of subjective evaluation, particularly in assessing text complexity for the DHH population.
    • In contrast to work evaluating captioning tools or ASL animations, this study finds that challenges in evaluating complex texts are more pronounced among high-literacy readers.
  • Experimental or Evaluation Results:

    • Reading speed and success rates in answering comprehension questions were limited by participants' literacy levels, while subjective metrics proved more effective.
    • Eliminating response bias among readers with different literacy levels is key to enabling cross-study comparisons.
  • Limitations and Future Directions:

    • Limitations:
      • Due to COVID-19 restrictions, eye-tracking data for more complex attention analysis could not be collected.
      • High-complexity texts may not sufficiently challenge the high-literacy group, requiring further validation.
    • Future Directions:
      • Test ATS output texts containing errors and grammatical issues to study the impact of errors on metric validity.
      • Conduct language-specific studies focusing on DHH users who use ASL.
      • Explore the effectiveness of complexity evaluation for other text types (e.g., novels, technical documents).
      • Extend research to other populations with reading difficulties (e.g., non-native readers).

Conclusion

This study provides clear guidance for evaluating text complexity for the DHH population, emphasizing the practicality of subjective evaluation metrics. It recommends that researchers control for participants' literacy levels in their analyses to ensure the reliability of results and the feasibility of cross-study comparisons.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47247/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445038
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)
work
Professions
Speech-Language Pathologists & Audiologists, Special Education Teachers, Social Workers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers