Methods for Evaluating the Fluency of Automatically Simplified Texts with Deaf and Hard-of-Hearing Adults at Various Literacy Levels
Authors
Voice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Universal & Inclusive DesignSpecial Education TeachersAssistive Technology Specialists
Title of the Paper
Methods for Evaluating the Fluency of Automatically Simplified Texts with Deaf and Hard-of-Hearing Adults at Various Literacy Levels
Bibliographic Information
- Subject Area: Automatic Text Simplification (ATS) and Accessibility Design
- Keywords: Automatic Text Simplification, Accessibility, Deaf and Hard-of-Hearing, Methodological Research, Reading Fluency, Text Simplification Evaluation
Research Background and Problem Statement
-
Identified Problems or Challenges:
- Automatic Text Simplification (ATS) may introduce grammatical or semantic errors, negatively affecting the fluency and integrity of the text.
- There is a lack of clear guidelines for evaluating the fluency of texts simplified for Deaf and Hard-of-Hearing (DHH) users when assessing ATS outputs.
- Previous studies primarily relied on "expert readers" with high literacy levels to evaluate the fluency of simplified texts, lacking methods tailored to specific target user groups.
-
Why It Matters:
- Literacy levels among DHH individuals vary widely, with some experiencing low literacy rates, making it crucial to evaluate text simplification for specific user groups.
- Designing evaluation methods tailored to target users can help develop more user-centric accessible reading tools.
-
Related Work and Research Motivation:
- Previous research has explored methods for evaluating text complexity for DHH individuals but has not explicitly addressed how to comprehensively evaluate the fluency (grammatical correctness) and fidelity (semantic accuracy) of simplified texts.
- This study aims to fill this methodological gap, particularly by exploring how to use appropriate evaluation metrics to assess text fluency among DHH individuals.
Proposed Solution
-
Methods or Solutions:
- Conduct rigorously designed experiments using multiple metrics (reading speed, comprehension question responses, subjective ratings, etc.) to reliably measure the fluency of simplified texts among DHH groups.
- Investigate the applicability of these metrics and potential "literacy bias" issues across DHH individuals with varying literacy levels (high vs. low).
- Design and implement methods for generating controlled text sets to ensure strict experimental comparisons of fluency levels.
-
Innovations:
- Proposed new methods for evaluating fluency, combining metrics such as reading speed and subjective grammaticality ratings to directly measure the fluency of simplified texts.
- Introduced target user groups beyond high-literacy readers (e.g., DHH individuals) into fluency evaluation research.
- Developed a research framework tailored to DHH individuals to guide future studies on text fluency evaluation and ATS performance.
-
Implementation Steps and Key Techniques:
- Text Generation and Validation:
- Selected six science education articles from the Newsela platform; generated low-complexity texts using two state-of-the-art ATS models and human rewrites.
- Conducted subjective evaluations of the generated texts (grammatical correctness, complexity).
- Used a Python-based greedy algorithm to generate experimental texts with controlled fluency levels.
- Participant Grouping:
- Recruited 29 DHH adults and categorized them into high (WRAT-H) and low (WRAT-L) literacy groups.
- Evaluation Methods:
- Measured reading speed, comprehension question scores, and subjective metrics (e.g., ratings of text readability, understandability, and grammatical correctness).
- Analyzed the discriminative power of these metrics and potential literacy bias.
- Text Generation and Validation:
Research Outcomes
-
Specific Findings:
- Effectiveness of Reading Speed and Grammar Judgments:
- Reading speed and subjective grammaticality ratings reliably distinguished "high," "medium," and "low" fluency texts, effective across both literacy groups (WRAT-H and WRAT-L).
- Literacy Bias and High Literacy Effects:
- Subjective evaluations of understandability and readability were only effective in distinguishing text fluency among high-literacy DHH readers.
- Comprehension question scores were sensitive to literacy levels (with high-literacy participants scoring significantly higher) but failed to differentiate fluency levels.
- Disrecommended Metrics:
- Comprehension questions were ineffective in distinguishing text fluency but could help maintain participant engagement during experiments.
- Implicit Findings:
- The scientific and complexity-controlled nature of the materials may have suppressed the effectiveness of certain potential metrics.
- Effectiveness of Reading Speed and Grammar Judgments:
-
Advantages Over Existing Solutions:
- Current evaluation methods focus primarily on text complexity or benchmark scores from expert readers, with limited research on the impact of target user groups (e.g., DHH adults) and varying literacy levels.
- This study innovatively introduced metrics such as reading speed and subjective grammaticality ratings to address fluency evaluation, offering a more comprehensive and practical approach.
-
Experimental or Evaluation Results:
- Reading Speed: Performed well across both literacy groups, effectively distinguishing "low" and "high" fluency texts.
- Subjective Grammar Ratings: Partially effective in distinguishing fluency levels (low vs. high).
- Readability and Understandability: Effective for high-literacy participants but not for low-literacy participants.
- Comprehension Questions: Failed to distinguish fluency levels and exhibited significant "literacy bias."
-
Limitations and Future Directions:
- Sample Size Limitation: Some metrics may not have shown significance due to insufficient statistical power.
- Single Text Domain: The study focused only on scientific texts; future research should validate findings across other text types (e.g., stories, technical documents).
- Metric Expansion: Future studies could explore additional evaluation methods, such as eye-tracking or Cloze tests, and integrate multimodal approaches.
- Research Expansion: Introduce semantic fidelity evaluation and design ATS interface adjustments to align with DHH user preferences.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can automatic text simplification fluency be evaluated to accommodate different literacy levels of deaf and hard-of-hearing users?Category: Deaf Communication and Sign Language RecognitionSimilar questionsarrow_forward
- Which evaluation metrics (e.g., reading speed, grammar judgment) best distinguish fluency levels across text difficulties?Category: Deaf Communication and Sign Language RecognitionSimilar questionsarrow_forward
- How does literacy level affect deaf and hard-of-hearing users' subjective evaluations and performance of text fluency?Category: Deaf Communication and Sign Language RecognitionSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Deaf users with varying literacy levels struggle to fluently read simplified text.Category: Deaf Communication and Sign Language RecognitionSimilar questionsarrow_forward
- 67%
Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing Individuals
CHI '24· Voice Accessibility +2
- 67%
Weaving Sound Information to Support Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User
CHI '25· Voice Accessibility +2
- 67%
"Make your voice high and squeak": Engaging Students with Visual Impairments in Collaborative, Participatory Sonic Interaction Design
DIS '25· Intelligent Voice Assistants (Alexa, Siri, etc.) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3517566
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Universal & Inclusive Design
work
Professions
Special Education Teachers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
3 related papers