FAME: Exploring Expressive Facial Avatars for Lyrical and Non-Lyrical Music Visualization for d/Deaf Individuals
Authors
Paper Title
FAME: Exploring Expressive Facial Avatars for Lyrical and Non-Lyrical Music Visualization for d/Deaf Individuals
Publication Info
- Topic area: Music accessibility for d/Deaf and Hard of Hearing (DHH) individuals using expressive avatars.
- Keywords: Accessibility, d/Deaf, music visualization, facial avatars, lip-sync, captions, rhythm, emotion, non-lyrical music.
Background and Problem
- Problem / challenge: Existing music visualization tools for DHH individuals, such as captions and abstract visualizers, fail to fully convey emotional depth, rhythm, and structural nuances of music. Lip-sync systems are often cognitively demanding and lack expressiveness, while non-lyrical music remains underexplored.
- Significance: Enhancing music accessibility for DHH individuals is critical for inclusivity, enabling richer engagement with music’s emotional, rhythmic, and lyrical elements.
- Motivation and related work: Prior research has explored abstract visualizers, haptic systems, and avatar-based lip-sync tools, but these approaches lack co-design with DHH users, emotional expressiveness, and support for non-lyrical music. This paper builds on these gaps by exploring avatar-based systems that integrate multimodal cues.
Solution
- Proposed approach: FAME (Facial Avatar for Musical Expression), a system that visualizes music through expressive facial animations, synchronized captions, and instrument highlights, with lip-sync for lyrics and scat-singing for melodies.
- Novelty:
- Combines captions, instrument cues, and avatar performance to convey musical elements.
- Introduces scat-singing avatars for non-lyrical music, maintaining visual engagement.
- Emphasizes stylized, emotionally expressive avatars to avoid the uncanny valley and improve clarity.
- Explores avatars as performers, interpreters, and companions for music accessibility.
- Procedure and key techniques:
- Formative study with 9 DHH participants to identify design requirements.
- Development of FAME, integrating captions, instrument highlights, and avatar animations.
- Two-phase exploratory study with 12 DHH participants to evaluate FAME against a visualizer baseline (ViTune) and gather feedback on its features.
Results
- Concrete findings:
- FAME achieved 95.8% accuracy in matching music to visualizations, outperforming ViTune (66.7% accuracy, p = 0.026).
- Participants rated FAME highly for conveying lyrics (Median = 4), emotion (Median = 4), and rhythm (Median = 4), but less effective for melody (Median = 2).
- Captions were deemed essential, with a median usefulness rating of 5.
- Advantage over baselines:
- Improved comprehension of lyrical and emotional elements compared to abstract visualizers.
- Higher alignment accuracy and emotional engagement, particularly for lyrical music.
- Experiments / evaluation:
- Comparison phase: Participants matched FAME and ViTune visualizations to audio across four songs.
- Application phase: Participants explored FAME’s features (avatars, captions, instrument highlights) across diverse musical genres.
- Metrics: Accuracy, response time, Likert-scale ratings, and thematic analysis of interviews.
- Limitations and future work:
- Limited participant diversity; most were already musically engaged.
- Focused on high-arousal genres, with less exploration of ambient or low-arousal music.
- Online study environment may not reflect real-world contexts like concerts.
- Manual preprocessing limits scalability; future work should enable real-time automation.
Summary
This paper introduces FAME, a facial avatar-based system designed to enhance music accessibility for DHH individuals by visualizing lyrics, rhythm, and emotion. Through iterative design and evaluation, FAME demonstrated significant improvements in musical comprehension and emotional engagement compared to baseline visualizers. Key contributions include integrating captions and instrument highlights, introducing scat-singing for non-lyrical music, and emphasizing expressive clarity over photorealism. Future work should address scalability, expand to diverse musical genres, and explore real-world applications in social contexts like concerts and karaoke.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)