Towards AI-driven Sign Language Generation with Non-manual Markers

Honorable Mention
Voice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Speech-Language Pathologists & Audiologists

Research Background and Issues

  • What problems or challenges did the authors identify?
    Current sign language generation systems fail to meet user needs in several critical aspects, such as:

    • Poor translation quality: Systems often struggle to accurately handle complex grammatical structures.
    • Lack of facial expressions and body language: These elements are essential grammatical markers in sign language.
    • Insufficient visual and motion quality: Existing solutions struggle to produce smooth, natural, and high-quality video content.
  • Why is this issue important?
    Sign language is a vital communication method for the Deaf and Hard of Hearing (DHH) community, but existing technologies fail to adequately meet their communication needs, hindering the widespread adoption and practical application of sign language technologies.

  • Research Motivation and Related Work
    The authors reviewed recent advancements in sign language generation, translation technologies, and related artificial intelligence research to address existing bottlenecks in sign language generation and propose more natural, comprehensive, and high-quality solutions.

Solution

  • What methods or solutions did the authors propose?
    A modular sign language generation system was proposed, with the following key objectives:

    • Capturing manual markers (hand movements) and non-manual markers (facial expressions) in sign language.
    • Generating coherent and smooth sign language videos, rather than simple concatenations of hand movements.
    • Integrating state-of-the-art large language models (LLMs) and video generation technologies to enhance translation quality and visual expressiveness.
  • What are the innovations of this solution?

    1. Modular Architecture: Divided into three modules:
      • English text to sign language representation (including manual and non-manual information).
      • Sign language representation to skeletal motion sequences.
      • Skeletal motion to sign language video generation.
    2. Non-manual Marker Processing: Extracting facial language features such as eyebrow movements and conditional tones from English text using LLM-based analysis.
    3. Motion Matching: Optimizing skeletal motion generation; employing "motion matching" techniques to ensure smooth transitions between movements.
    4. High-quality Video Generation: Utilizing advanced image generation networks to directly produce realistic dynamic videos.
  • What are the implementation steps and key technologies used?

    1. Module 1: English Text to Sign Language Representation
      Using the GPT-4o model with few-shot and zero-shot methods to generate sign language representations and extract non-manual information. Data was cleaned and standardized to support more consistent translations.
    2. Module 2: Sign Language Representation to Skeletal Motion Sequences
      Leveraging a signing dictionary from the ASLLRP dataset and generating smooth motion sequences using motion matching techniques, while incorporating facial expression enhancements.
    3. Module 3: Skeletal Motion to Video Frame Generation
      Applying a U-Net model to transform skeletal motions into high-quality sign language videos with rich visual details.

Research Outcomes

  • What specific results were achieved?

    1. In the text-to-sign language representation translation task, the BLEU-4 score reached 0.276, significantly surpassing the highest result in existing literature (0.191).
    2. In the sign language video generation task, the optimization techniques significantly improved visual quality and motion naturalness.
  • What advantages does it have compared to existing solutions?

    • Improved Translation Performance: The proposed system not only generates more accurate sign language representations but also handles non-manual markers to enhance overall comprehension.
    • Enhanced Visual Quality: The video generation module effectively reduces blurriness and motion discontinuity by improving skeletal visualization and using high-resolution data.
    • Optimized User Feedback: Through user studies, the system captured the needs and pain points of DHH users, which informed iterative improvements.
  • What were the experimental or evaluation results?

    1. In user studies, approximately 53.8% of videos were rated as "acceptable" or higher in terms of semantic accuracy; non-manual markers significantly improved the recognition of translation results.
    2. While visual quality still requires improvement, the modular design allows for independent optimization of each component in the future.
    3. Motion naturalness performed well in user evaluations, particularly for videos generated directly from sign language data.
  • Limitations and Future Directions

    • Limitations:
      • The current dataset is small and suffers from issues such as inconsistent annotations and visual quality, limiting the model's generalizability.
      • The system cannot handle context-dependent sign language content, such as spatial references and role switching.
      • The video generation module still needs improvement in maintaining temporal consistency.
    • Future Directions:
      • Develop larger, standardized sign language datasets to improve the robustness of translation models.
      • Introduce more non-manual linguistic features, such as mouth shapes and body postures.
      • Explore long-term memory mechanisms for contextualized sign language generation.
      • Enhance video quality by adopting more advanced diffusion models to address detail loss.

Conclusion

This paper presents a modular sign language generation system that significantly improves the translation of English text into sign language and enhances the naturalness and visual quality of sign language videos by incorporating both manual and non-manual information. Although there is still a gap in fully meeting the needs of the DHH community, the research outcomes provide promising directions for developing more practical and natural sign language generation technologies.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189085/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713855
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
10 authors
sell
Subtopics
Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)
work
Professions
Speech-Language Pathologists & Audiologists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers