Comparison of Methods for Evaluating Complexity of Simplified Texts among Deaf and Hard-of-Hearing Adults at Different Literacy Levels
Authors
Title of the Paper
Comparison of Methods for Evaluating Complexity of Simplified Texts among Deaf and Hard-of-Hearing Adults at Different Literacy Levels
Paper Information
- Subject Area: Accessibility Design and Text Simplification Research
- Keywords: Automatic Text Simplification, Methodological Research, Deaf and Hard-of-Hearing, Accessibility Design, Text Complexity Evaluation, Reading Comprehension, Response Scoring, Subjective Evaluation, Low Literacy, Information Technology
Research Background and Problem
-
Issues or Challenges Identified by the Authors:
- The Deaf and Hard-of-Hearing (DHH) population exhibits diverse reading abilities, and low literacy levels may lead to difficulties in reading complex texts.
- Automatic Text Simplification (ATS), as an assistive reading technology, benefits DHH users, but previous research lacks guidance on how to evaluate the complexity of simplified texts for this group.
- Common metrics in text complexity evaluation (e.g., reading speed, comprehension questions, subjective judgments) have not been validated, making their applicability to DHH users uncertain.
-
Why This Problem is Important:
- Evaluating the quality of simplified texts is crucial for developing assistive reading tools tailored to DHH users.
- Without reliable text complexity evaluation, it is challenging to improve the accuracy and user experience of assistive technologies.
-
Motivation and Related Work:
- The authors reviewed existing studies on DHH users, language technology evaluation methods, and ATS literature, finding that many studies lack explicit methods for evaluating text complexity specific to the DHH population.
- Drawing on prior experience in evaluating captioning tools and American Sign Language (ASL) animations, this study aims to establish a more reliable framework for ATS evaluation.
Solution
-
Proposed Method or Solution:
- Compare news articles of varying complexity levels (manually simplified versions) to ensure high-quality corpora.
- Test multiple evaluation metrics, including objective metrics (e.g., reading speed, success rate in answering comprehension questions) and subjective metrics (e.g., readability and comprehensibility assessments).
-
Innovative Aspects:
- Identify which evaluation metrics effectively distinguish between texts of different complexity levels, particularly for DHH users with varying literacy levels.
- Propose grouping readers (based on literacy levels) and controlling experimental design to reveal statistically significant differences in metric effectiveness.
-
Implementation Steps and Key Techniques:
- Corpus Selection: Articles were selected from the Newsela platform, which provides news at different complexity levels, ensuring uniform standards for length.
- Experimental Metrics:
- Reading speed (calculated as word count/reading time).
- Comprehension questions (main factual questions, categorized as high-complexity and low-complexity questions).
- Score prediction (participants predict their own scores on comprehension questions).
- Subjective evaluation: including judgments on text comprehensibility and readability.
- Method Validation: Participants were divided into high and low literacy groups (WRAT-H and WRAT-L) based on the Wide Range Achievement Test (WRAT).
- Experimental Design: A Latin square design was used to randomize article complexity and question order, reducing experimental bias.
Research Findings
-
Specific Findings:
- Subjective metrics (especially readability evaluation) performed best: This metric effectively distinguished text complexity regardless of participants' literacy levels.
- Low-complexity comprehension questions performed better in the low-literacy group: They could only differentiate between the lowest and highest complexity conditions.
- High-complexity comprehension questions and reading speed had limited effectiveness: These are not suitable as core metrics for evaluating text complexity.
-
Comparison with Existing Solutions:
- Unlike studies that rely solely on objective metrics, this research highlights the superiority of subjective evaluation, particularly in assessing text complexity for the DHH population.
- In contrast to work evaluating captioning tools or ASL animations, this study finds that challenges in evaluating complex texts are more pronounced among high-literacy readers.
-
Experimental or Evaluation Results:
- Reading speed and success rates in answering comprehension questions were limited by participants' literacy levels, while subjective metrics proved more effective.
- Eliminating response bias among readers with different literacy levels is key to enabling cross-study comparisons.
-
Limitations and Future Directions:
- Limitations:
- Due to COVID-19 restrictions, eye-tracking data for more complex attention analysis could not be collected.
- High-complexity texts may not sufficiently challenge the high-literacy group, requiring further validation.
- Future Directions:
- Test ATS output texts containing errors and grammatical issues to study the impact of errors on metric validity.
- Conduct language-specific studies focusing on DHH users who use ASL.
- Explore the effectiveness of complexity evaluation for other text types (e.g., novels, technical documents).
- Extend research to other populations with reading difficulties (e.g., non-native readers).
- Limitations:
Conclusion
This study provides clear guidance for evaluating text complexity for the DHH population, emphasizing the practicality of subjective evaluation metrics. It recommends that researchers control for participants' literacy levels in their analyses to ensure the reliability of results and the feasibility of cross-study comparisons.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Which metrics can effectively distinguish text at different complexity levels, especially for deaf and hard-of-hearing users with varying literacy levels?Category: Reading Behavior, Attention, and Eye-Tracking AnalysisSimilar questionsarrow_forward
- Are subjective evaluations (e.g., readability and comprehensibility) more suitable than objective metrics (e.g., reading speed and comprehension Q&A) for assessing simplified text complexity?Category: Reading Behavior, Attention, and Eye-Tracking AnalysisSimilar questionsarrow_forward
- Does grouping users by literacy level (e.g., WRAT tests) in experimental design improve reliability of text complexity assessment?Category: Reading Behavior, Attention, and Eye-Tracking AnalysisSimilar questionsarrow_forward
Practical Problems
1- Deaf and hard-of-hearing users struggle to read traditionally complex text, affecting assistive reading tool precision.Category: Reading Behavior, Attention, and Eye-Tracking AnalysisSimilar questionsarrow_forward
Based on Jaccard similarity of research subtopics & professions (≥60%)