V5.10.4Transcription errors quoted as verbatim speechdesignresearch

Automatic transcription errors get quoted as the original words

Aliases: authority of faulty transcripts · misquotation by ASR · draft status of transcripts

What it is

Automatic speech recognition (ASR) transcripts always contain errors, densest on proper nouns, jargon, accents, and overlapping speech. The problem is not the errors but their treatment: once a transcript lands in a searchable archive, it gets used as a verbatim record — "the minutes say you said this." Words the speaker never uttered appear in first person and get quoted: the injury form unique to transcription errors is that they forge not facts but a record of speech.

Why it happens

Three mechanisms make transcription errors more dangerous than ordinary text errors. First, errors are unevenly distributed: recognition error rates vary systematically with accent, dialect, speech rate, non-nativeness, and terminology density — quantified studies have documented several-fold error-rate gaps across speaker groups — so the system is systematically less faithful to some speakers, and the risk of being misquoted falls unfairly on accented and non-native speakers. Second, text carries a documentary authority readers extend automatically: seeing a sentence in quotation marks implies it was verified, and the transcript is precisely the one thing nobody verified. Third, correction fails after propagation: the moment the erroneous sentence is copied into minutes, tickets, and emails, fixing the source cannot reach the copies — and the speaker is left protesting "I never said that," with the burden of proof inverted onto the person who was mis-transcribed. Together these grant the transcript a finalized status it has not earned.

Studying it

ASR performance evaluation is a mature field: with human-labeled speech corpora, word error rates can be computed per speaker group, and systematic gaps across accents, genders, and age groups are documented in multiple quantitative studies. Extending this to collaboration settings: error-type analysis of meeting transcripts (entity names, numbers, negations, pronoun attribution — negation and number errors being the most damaging since they flip meaning), and tracing the chain by which errors enter citations (transcript → minutes → ticket → confrontation). Methodological caveats: word error rate masks the importance distribution of errors — five percent of errors landing on proper nouns versus on negations have entirely different consequences, so evaluation must weight by semantic consequence; and commercial tools' error rates shift with model versions, so any concrete number must be labeled with model and date rather than cited as a stable constant.

Where it stops holding

An error becomes an injury only through the premise of being taken as fact: transcripts used solely for personal catch-up and discarded same-day have no propagation chain and do no harm; transcripts entering searchable archives and serving as citation sources are the ones needing governance. High-noise conditions (far-field audio, overlapping speakers, telephone quality) steeply raise error rates, making those transcripts closer to summaries than verbatim records. Human minutes contain errors too, but human error carries an accountable corrector — what machine transcription lacks is precisely that responsibility link, and governance must supply accountability, not just accuracy.

Applying it

  • Always label transcripts "machine-generated, unverified," persistently visible in the interface and exports, never blending them into archives in undifferentiated body text.
  • Before high-stakes content (numbers, commitments, name-calling points, objections) moves from transcript into minutes, require confirmation by the speaker; give everyone an entry to review and correct their own utterances, with edit history.
  • Pre-load custom vocabularies (project, product, customer names) into the recognition system to suppress the most lethal error class at the source.
  • Require quotations of transcripts to link the original audio timestamp, making "check the actual words" cheaper than "just trust the text."
  • Verification: periodically sample quoted sentences from transcripts against the original audio to measure misquotation rates; group misquotes by speaker to see whether they concentrate in particular accent or rate groups — a concentration signals a fairness problem requiring a more balanced recognition option.

Related

  • Same group: V5.10.1 Recording requires everyone's informed consent beforehand, with a real option to refuse · V5.10.2 Being recorded changes what people say and how candidly · V5.10.3 Searchable transcripts extend the reach of original speech · V5.10.5 Retention period and access scope must be fixed in advance
  • Nearby: V2.03 Change Awareness · V5.08.3 Searchable text records are the long-term value
  • Search terms: speech recognition error · word error rate · transcription quality · ASR bias

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/V5.10.4