Beyond Technical Metrics: Understanding the Gap Between AI Performance and Deaf User Experience in Chinese Natural Sign Language Generation
Authors
Paper Title
Beyond Technical Metrics: Understanding the Gap Between AI Performance and Deaf User Experience in Chinese Natural Sign Language Generation
Publication Info
- Topic area: AI-generated sign language evaluation and Deaf user experience.
- Keywords: AI-generated sign language, Deaf user experience, Chinese Natural Sign Language, participatory design, accessibility technology, cognitive load, linguistic quality, cultural authenticity.
Background and Problem
- Problem / challenge: Current evaluation metrics for AI-generated sign language focus on technical benchmarks but fail to capture Deaf users' comprehension, cognitive load, and cultural authenticity. This leads to a mismatch between system performance and real-world usability.
- Significance: Addressing this gap is critical for ensuring that AI-generated sign language systems meet the communicative needs of Deaf users, enhancing accessibility and inclusion.
- Motivation and related work: Prior work has relied on gloss-based and word-order mapping methods, which fail to capture CNSL’s spatial grammar and situational meaning. Evaluation paradigms have been researcher-defined, limiting Deaf users' agency in shaping criteria. This paper builds on participatory design principles to redefine evaluation frameworks.
Solution
- Proposed approach: A participatory evaluation framework co-developed with Deaf co-researchers, incorporating seven dimensions that prioritize user comprehension, cognitive load, and cultural authenticity.
- Novelty:
- Development of a seven-dimensional evaluation framework integrating objective comprehension tests, subjective measures, and CNSL-specific criteria.
- Discovery of inverse complexity effects, where linguistic complexity reduces cognitive load for human signers but increases it for AI-generated signing.
- Methodological insights for inclusive HCI, emphasizing participatory design in evaluation development.
- Procedure and key techniques:
- Creation of a scene-aware CNSL generation prototype with three stages: text-to-HamNoSys, HamNoSys-to-pose, and pose-to-video synthesis.
- Five workshops with Deaf co-researchers to iteratively define evaluation dimensions, scales, and criteria.
- User evaluation study with 24 Deaf participants using blinded, randomized video comparisons of AI-generated and human signing.
Results
- Concrete findings:
- AI-generated videos scored lower on comprehensibility (objective: 24.5% vs. human 37.6%; subjective: AI 33.0% Bad vs. human 34.4% Good).
- Mental demand was higher for AI-generated videos (mean 38.3 vs. human 27.9).
- Acceptance ratings favored human signing (mean 47.1 vs. AI 36.8).
- Linguistic quality dimensions (fluency, spatial grammar, non-manual markers, expressiveness) consistently showed human signing outperforming AI-generated signing.
- Advantage over baselines:
- Improved technical metrics (e.g., symbol accuracy 80.6%, SSIM 0.865) did not translate into better user outcomes, highlighting the disconnect between technical benchmarks and real-world usability.
- Experiments / evaluation:
- 54 videos (27 human-AI pairs) evaluated across declarative, interrogative, and negative sentence types.
- Mixed-effects models analyzed ratings for seven dimensions, revealing significant Source × Sentence Type interactions.
- Qualitative interviews captured user feedback on comprehension challenges and cognitive fatigue.
- Limitations and future work:
- Evaluation focused on single sentences rather than extended discourse.
- Findings based on one prototype; validation with alternative models is needed.
- Recruitment bias toward younger, digitally engaged signers; future studies should include broader demographics.
- Hearing researchers led the process; future work should explore Deaf-led research structures.
Summary
This study highlights the gap between technical benchmarks and Deaf user experience in AI-generated CNSL videos. A participatory evaluation framework co-developed with Deaf co-researchers revealed dimensions—comprehensibility, cognitive load, and cultural authenticity—that conventional metrics overlook. User evaluations showed systematic disadvantages for AI-generated signing compared to human signing, particularly for complex sentence types. The findings emphasize the need for Deaf-centered evaluation criteria and suggest practical adjustments for improving sign language generation systems. Future work should extend evaluation to discourse-level interactions, validate across architectures, and prioritize Deaf-led research processes.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)