SpellRing: Recognizing Continuous Fingerspelling in American Sign Language using a Ring

Foot & Wrist InteractionVoice AccessibilityMotor Impairment Assistive Input TechnologiesSpeech-Language Pathologists & AudiologistsSpecial Education TeachersDisability Service Providers

Research Background and Problem

  • Problems and Challenges:

    • American Sign Language (ASL) fingerspelling is a crucial component of ASL, particularly for spelling proper nouns, names, and technical terms. In natural fingerspelling, the complexity of hand shapes, palm orientations, and motion variations poses significant challenges for continuous recognition tasks.
    • Traditional computer vision-based methods are limited by high costs, privacy concerns (e.g., requiring cameras), and poor portability. Existing wearable devices (e.g., data gloves, wristbands) often require bulky hardware and significantly interfere with users' gestural behavior.
    • Currently, most devices can only recognize isolated letters and fail to accommodate the natural continuous fingerspelling behavior of fluent sign language users.
  • Significance:

    • Fingerspelling recognition is fundamental to building comprehensive ASL recognition systems, providing accessible text input methods for the hearing impaired and enhancing their interaction with digital devices (e.g., smartphones, VR devices).
  • Research Motivation and Related Work:

    • The authors summarize the requirements for sign language recognition devices: accuracy, minimal interference, and natural user behavior. They note that current systems have yet to fully meet these criteria.
    • To address the portability and privacy issues of traditional methods, the authors explore the potential of lightweight devices (e.g., rings) and integrate active acoustic sensing with inertial measurement unit (IMU) technology.

Solution

  • Methodology and Innovation:

    • A smart ring named SpellRing is proposed, enabling real-time recognition of continuous fingerspelling with a single device. This is the first single-ring system to combine active acoustic sensing and IMU sensors.
    • The system incorporates a Connectionist Temporal Classification (CTC) algorithm, allowing recognition of complete words without requiring individual letter annotations in fingerspelling.
    • A dual-modal sensing technique is employed: active acoustic sensing captures hand shape features, while the IMU tracks hand movements and palm orientation, generating fine-grained motion data.
  • Implementation Steps:

    1. Hardware Design:
      • The ring integrates a microphone, speaker, and IMU module, utilizing a flexible printed circuit board (FPCB) for lightweight design.
    2. Data Processing:
      • Acoustic sensor data (differential echo features) and IMU data (hand rotational angular velocity) are normalized and input into a deep learning model.
      • A multimodal deep learning pipeline is developed to fuse data from different sensors, with the final output generated by the CTC model.
    3. Model Training:
      • A two-stage training approach is tested, involving cross-user pretraining and single-user fine-tuning, to enhance the model's performance in user-dependent settings.
    4. Language Model Correction:
      • An N-gram language model is used to further correct potential recognition errors, improving phrase prediction accuracy in real-world contexts.

Research Outcomes

  • Specific Results:

    • In offline recognition of 1,164 words, SpellRing achieved a Top-1 accuracy of 82.45% and a Top-5 accuracy of 92.42%.
    • In real-time phrase-level testing, the system achieved a word error rate (WER) of 0.099 after language model correction.
    • The system demonstrated strong cross-user generalization while maintaining a compact hardware design.
  • Advantages:

    • Compared to existing multi-ring systems (e.g., FingerSpeller achieving 91% accuracy with five rings), SpellRing delivers excellent recognition performance using only a single ring.
    • The system minimally interferes with users' natural fingerspelling behavior, while the hardware is more lightweight and portable.
    • A sensor fusion framework is proposed, effectively mitigating the limitations of single-modal sensors.
  • Experiments and Evaluation:

    • User studies involved 20 participants, including 13 fluent ASL users and 7 beginners:
      • Fluent users exhibited relatively lower accuracy (84.06%), reflecting the model's challenges in adapting to rapid fingerspelling behavior.
      • Beginners performed better (94.38%) due to their clearer gestures, which improved model performance.
    • The language model significantly reduced error rates in phrase prediction, highlighting the potential advantages of incorporating linguistic context.
  • Limitations and Future Directions:

    • Limitations:
      • User-independent recognition performance remains low (48.42%); the system struggles with rapid fingerspelling (up to 8 letters per second).
      • The current vocabulary is limited (1,164 words), failing to fully encompass the richness of sign language expressions.
      • The device may be affected by environmental noise, requiring further optimization of acoustic processing mechanisms.
    • Future Directions:
      • Expand the dataset and develop more advanced pretraining models to improve cross-user recognition performance.
      • Integrate recognition of two-handed gestures and extend to complete ASL sentences and non-manual markers.
      • Enhance hardware design by utilizing flexible circuit boards and curved batteries for a fully ring-shaped device.
      • Strengthen collaboration with the hearing-impaired and sign language users to ensure device optimization aligns with the real needs of target users.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188879/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713721
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
11 authors
sell
Subtopics
Foot & Wrist Interaction, Voice Accessibility, Motor Impairment Assistive Input Technologies
work
Professions
Speech-Language Pathologists & Audiologists, Special Education Teachers, Disability Service Providers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers