ELMI: Interactive and Intelligent Sign Language Translation of Lyrics for Song Signing

Voice User Interface (VUI) DesignConversational ChatbotsVoice AccessibilityGenerative AI (Text, Image, Music, Video)Human-LLM CollaborationSpeech-Language Pathologists & AudiologistsMusicians, DJs & Sound Designers

Research Background and Issues

  • Issues and Challenges: The authors identified four major challenges in translating song lyrics into sign language:

    1. Semantic Translation: Understanding the meaning of lyrics, especially those with complex forms, or content specific to certain cultures or languages.
    2. Syntactic Translation: Selecting appropriate sign language symbols while addressing grammatical differences between sign language and English.
    3. Expressive Translation: Conveying the emotions of the lyrics through facial expressions and body language to the audience.
    4. Rhythmic Translation: Synchronizing sign language translation with the rhythm of the song, particularly for fast-paced songs, which poses a significant challenge for many translators.
  • Research Significance: Translating songs into sign language is an important means of making music more accessible to d/Deaf (deaf/sign language users) communities, integrating elements of language, culture, and art. Existing research predominantly focuses on music perception and accessibility, with limited attention to designing accessible support systems from artistic and cultural perspectives.

  • Relevant Background: Previous studies on sign language translation have primarily focused on formal communication scenarios, such as news broadcasting or daily conversations. Research on lyric translation, especially the development of creative support tools, remains scarce.

Solution

  • Methods and Tools: The authors proposed and developed an interactive intelligent lyric translation tool—ELMI (Explore Lyrics and Music Interactively). This tool integrates real-time lyric-video synchronization, large language model (LLM)-based conversational functionality, and performance guidance to help users address the four identified challenges.

  • Core Innovations:

    1. Real-time Lyric and Video Synchronization: Lyrics are highlighted line by line, combined with video playback, enabling users to match rhythm and emotion more effectively.
    2. LLM-based Assisted Discussion: Users can engage with ELMI to discuss semantics, translation methods, emotional expression, and rhythm synchronization, fostering creative decision-making.
    3. Fine-grained Line-by-line Translation: Users can input "gloss" (textual representation of sign language) line by line, with ELMI providing real-time improvement suggestions.
  • Implementation Steps and Techniques:

    1. Obtain standard lyrics and timestamps, and align lyrics with video content using automatic speech recognition (ASR) to generate word-by-word annotations.
    2. Users can select song segments line by line, input and edit their translations with rich visual feedback (e.g., dynamic highlighting or expression suggestions).
    3. During interactions with ELMI, the GPT-4-based conversational system generates specific suggestions related to "semantic, syntactic, emotional, and rhythmic translation."

Research Outcomes

  • Specific Results:

    1. ELMI has been proven to help users more efficiently capture the deeper meanings of lyrics and performance emotions, significantly improving translation independence and confidence.
    2. Regardless of hearing background, participants successfully used ELMI to translate lyrics while incorporating personal styles to produce diverse translations and performances.
  • Advantages:

    1. Facilitates multimodal (visual, auditory) collaboration, making music translation more accessible to d/Deaf users.
    2. The AI-driven feedback mechanism both encourages users and provides constructive improvement suggestions during interactions.
    3. Offers an "all-in-one" translation workspace, integrating lyrics, videos, and translation operations.
  • Experiments and Evaluation:

    1. For popular songs like "Butter" (BTS), research demonstrated that ELMI's real-time assistance system significantly alleviated difficulties in the translation process.
    2. Chi-square tests and analysis revealed that d/Deaf users focused more on rhythm synchronization, while hearing users prioritized emotional expression.
  • Limitations and Future Directions:

    1. Multiple participants noted that ELMI's output sometimes overly relied on English grammatical order or lacked deeper semantic expression, necessitating improvements in the model's understanding of ASL grammar and culture.
    2. The tool lacks support for multi-line perspectives (e.g., generalized translation across related lines).
    3. Integration with sign language dictionaries should be enhanced to provide more intuitive symbol representations (e.g., videos or examples).

Conclusion

Through the development of ELMI, the authors introduced greater cultural sensitivity and technological support for sign language lyric translation. Future work could focus on enhancing adaptability to cross-cultural contexts and deeper understanding of ASL, while expanding sign language input and expressive capabilities through multimodal information such as video inputs.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189454/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713973
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Voice User Interface (VUI) Design, Conversational Chatbots, Voice Accessibility, Generative AI (Text, Image, Music, Video)
work
Professions
Speech-Language Pathologists & Audiologists, Musicians, DJs & Sound Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers