Auto-captions misfire badly on specialized talk
Aliases: auto-captions · ASR captions · word error rate · live captions
What it is
The one-click automatic captions a platform draws over a video can look like a usable draft in casual, close-miked, single-speaker talk. Move to a medical lecture, a hearing, a product launch, an academic paper, and proper names and critical negations get swapped for nearby common words. The proposition flips. High error rate is the default on specialized content, not the odd typo.
Why it happens
Recognition is built on models of ordinary talk. Specialized terms sit in the long tail of the training distribution: drug names, statute numbers, product SKUs, people, expansions of acronyms. The model folds them into frequent words that sound close — once not becomes now, the truth-value of the sentence inverts.
The second layer is the room. Cross-talk, distance, a mask, an accent, terms fired at speed, and acoustic confidence falls with them. A domain hotlist can pull back registered product names; unlisted coinages and homophone terms still fail. Live output cannot be edited in time, so the error is welded to those seconds of picture. The reader has no later, correct version to consult.
Treating the engine as a drafting step that a human then edits, and treating the engine as the shipped captions, are different pipelines. The second one, on specialized talk, reliably ships a text that looks like captions and reads like a different lecture.
Studying it
Against a human-corrected reference, compute word error rate (WER), and for specialized talk also proper-name error rate — overall WER is dragged down by function words while every name is still wrong. Live captioning often uses NER (number of errors in recognition): substitutions, insertions and deletions scored separately, with extra weight on critical names.
Independent variables: domain (chat / medicine / launch), domain lexicon on or off, close single speaker versus overlap. Dependent variables: overall WER, name error rate, negation flips; and whether people who only see the auto-captions can answer domain questions.
Three-minute samples are enough to run. Comprehension items must use domain facts. “Did you get the gist” will bless a fatal error.
Where it stops holding
Closed vocabularies — ordering food, a finite command set — can be accurate enough. Using auto-captions as a draft before human edit does not stop the error rate being high; it stops the errors reaching the viewer. Casual talk in good recording conditions sometimes passes. Languages, accents and dialects with thin training data are worse; an English close-mic chat demo does not transfer to a product keynote. Dimensions the engine never emits (who spoke, what that sound was) belong on another card. This one only scores whether the words are the words.
Applying it
- For courses, launches, clinical or legal material, human-edit before shipping. Do not treat the player’s automatic track as captions delivered.
- Keep a domain list: product names, people, recurring terms. Spot-check against this session’s script before release.
- For live, use a stenographer or a short delay for correction. If that is impossible, label the track as automatic and fallible; do not dress it as reviewed.
- How to check: take three minutes that contain at least ten names or one critical negation, and count errors against the script. One negation or drug name that flips the proposition fails the segment.
Related
- Same group: J4.01.1 Captions must name the speaker and the sound that matters · J4.01.3 Caption type size, colour and position must be the viewer's to change
- Nearby: D2.12 Equivalence between sound and captions · A3.08 Hearing loss implications for interaction design
- Search terms:
automatic captions·ASR captions·word error rate