Caption delay destroys the timing cues that sound carried
Aliases: caption latency · sync · timing cues
What it is
When a caption appears is itself information. The same line, early or late, gives a different cue; if captions and sound are offset, the temporal relations the sound supplied—who spoke first, which line followed which action—are scrambled, and the reader reconstructs a different causal order than the listener.
Why it happens
Timing in speech carries turn-taking, interruption, and causal connection: a reply landing right after a question means answering, while a few seconds' gap may mean hesitation. A lagging caption makes the reader's order disagree with the audio, especially when speakers alternate or when effects interleave with dialogue, and causality gets misattributed. Delay can also surface two temporally non-adjacent lines at once. Caption processing latency comes from the generation pipeline as well as from dwell-time choices made for readability, and the two needs do not always align.
Studying it
Systematically vary caption offset—early, synchronized, and lagged by milliseconds to seconds—and measure accuracy on speaker turns, causal order, and event timing, along with reading load and the share of content the reader cannot keep up with. Variables include speaker count, speech rate, and whether critical effects coincide. Include sequence errors as an outcome, since delay mainly breaks order rather than content.
Where it stops holding
Extending dwell time for readability is necessary, and brief lag is normally harmless; problems appear when delay reverses causal judgments or leaves the reader unsure who is speaking. With slow pacing, a single speaker, and no concurrent effects, timing cues are redundant enough that delay does less damage. Captions produced by live recognition add unpredictability, which differs in kind from a small, stable offset.
Applying it
- Align caption onset with the corresponding audio and put readability dwell on the disappearing end rather than the appearing end.
- Check that order remains readable where speakers alternate or effects interleave with dialogue.
- For live-generated captions, mark the delay or offer a review point so readers do not take misordering as fact.
- Verification: play a rapid turn-taking segment and have caption-only readers write the event order, comparing against audio listeners; disagreement means a timing problem.
Related
- Within the group: D2.12.1 Captions carry the words but rarely restore tone or urgency · D2.12.4 Equivalence is not full replacement; some sound information is lost in transcription
- Adjacent: D3.04 Design consequences of haptic-visual de-synchronization · D2.06.3 Speech output must support interruption and replay
- Search terms:
caption latency·audiovisual sync·temporal cues