Captions must name the speaker and the sound that matters
Aliases: speaker identification · sound-effect captions · closed captions · non-speech captions
What it is
Mute the clip. The line of type still has to answer two questions: who is speaking, and what that non-speech sound just did. Captions that only transcribe dialogue turn a multi-speaker scene into unowned sentences, and drop the doorbell, the gunshot, the alarm, the off-screen laugh. They serve people who cannot hear, cannot hear well, or cannot turn sound on right now. They live on the video timeline. They are not a document you search.
Why it happens
Auditory information in a video comes in at least three layers: words, speaker identity, and event sounds. Words can become glyphs. Identity can sometimes be recovered from a visible mouth. The moment someone is off-camera, on the other end of a phone, or overlapping after a cut, the visual cue is gone and the caption has to carry the name.
Key sound effects have no words to transcribe. They move events: a door, a countdown, a laugh track. Leaving them out leaves the event out. Continuous room tone usually stays out of captions not because it is not sound, but because it does not change the next judgement. A theme that suddenly identifies a character is not decoration.
Studying it
Watch captions with the audio off. Run two item types, not one satisfaction score: speaker attribution and event detection. Ask “who said that” and “did the door sound,” not “did you like it.”
Independent variables: speaker labels present or not, non-speech labels present or not, speaker on-screen or off. Dependent variables: attribution accuracy, misses on critical events, narrator lines misread as character dialogue.
Broadcast caption QC (speaker IDs, non-speech coverage) is a production checklist. It does not replace the mute comprehension test — a checklist can tick “has a label”; the test shows whether the label named the right person.
Where it stops holding
A single on-screen speaker with a clear mouth does not need a name on every line; the labels become noise. A meeting product that already badges the speaker on video may not need the name repeated in the caption. Decorative street bed does not belong line by line. Live captions thin out non-speech labels under time pressure; that is a latency trade, not a licence to skip sound. Tone and spatial direction barely survive as a few words — completeness here means who spoke and which kind of event sounded, not the whole acoustics.
Applying it
- With two or more speakers, label by name or role. Mark voice-over, the far end of a call, and narration as their own voices.
- List the sounds that change understanding or the next action (alarms, doors, laughs, cue tones). Write them as short tags such as
[doorbell], not as a spectrum. - When dialogue and a sound fight for the same slot, keep the sound that changes a judgement, then tighten the dialogue line.
- How to check: watch the whole piece muted. For each line, “who said this?” For each turn, “what just sounded?” Every wrong person and every missing sound is still a dialogue-only transcript pretending to be captions.
Related
- Same group: J4.01.2 Auto-captions misfire badly on specialized talk · J4.01.3 Caption type size, colour and position must be the viewer's to change
- Nearby: D2.12 Equivalence between sound and captions · A3.08 Hearing loss implications for interaction design · J4.02 Transcripts
- Search terms:
caption speaker identification·closed captions·non-speech captions