Multimodal Silent Speech-based Text Entry with Word-initials Conditioned LLM

Electrical Muscle Stimulation (EMS)Hand Gesture RecognitionHuman-LLM Collaboration

Although exhibiting great potential in enabling seamless communication between humans and conversational agents, large vocabulary recognition is still challenging for silent speech interfaces. In this research, we propose a novel interaction technique that combines silent speech and typing to enable more efficient text entry while preserving privacy. This technique allows users to use abbreviated phrase input while still ensuring high accuracy by leveraging visual information. By fine-tuning a large language model with a visual speech encoder, we condition the models to decode the speech content with word initials as hints. Evaluations on existing datasets show that our model can reduce the Word Error Rate from 20.3% to 9.19%, compared to state-of-the-art visual speech recognition models. Results from a user study demonstrated significant improvements in input speed and keystroke saving. Participants reported that our prototype, LipType, leads to an overall lower perceived workload, particularly in the effort and physical demand dimension.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/cui/204387/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3719160.3736612
At a Glance

Paper Snapshot

fact_check
dataset
Source
CUI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Electrical Muscle Stimulation (EMS), Hand Gesture Recognition, Human-LLM Collaboration
work
Professions
—
article
Content Status
Abstract only
hub
Related Papers
3 related papers