helpResearch questionCaptions, Speech Transcription, and Deaf Communication Assistance
How can real-time flowing speech be segmented word-by-word in time to generate accurate captions?UIST '24Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition