Accuracy gaps exclude particular populations
Aliases: speech equity · disparate WER · availability exclusion
What it is
When a group’s recognition error rate is high enough that tasks routinely fail, voice is no longer a general channel. It is exclusion. Exclusion does not need an explicit deny list: a default that only hears a standard accent and a standard lexicon has the effect of shutting other speakers out of the feature. The issue is not “could be a bit more accurate.” It is who is allowed to finish the same task on this path.
Why it happens
Once a product makes voice the primary path, or the only hands-free path, a WER gap becomes a completion-rate gap. In driving, cooking, or low-vision use, excluded people have no equivalent keyboard to fall back on. Repeated failure also changes behaviour: they shift toward a more “standard” accent, abandon voice, or hand the device to someone whose accent sits closer to the training set. Logs then fill with already-covered speakers, and the gap reinforces itself. Exclusion is an interaction outcome, not only a model metric: the same control succeeds on one breath for some people and loops on respeak for others.
Studying it
Stratify task completion—not only WER—by group: successful submits, retries, abandonment, and whether they switched channels. Field work should include dialect users and speakers far from the advertised accent, not only Mandarin-speaking colleagues. Compare exclusion strength with and without a fallback channel. Ethically, do not treat participants’ failures as spectacle, and state in consent how recordings will be used. Coaching everyone in the lab into a standard accent measures the exclusion away.
Where it stops holding
Occasional recognition failure is not exclusion; exclusion requires a stable gap that blocks a critical task. A voluntary stylistic choice (a theatrical voice) is not the same as an accent or age-related voice the speaker cannot drop. If a product never claimed to support a dialect, say so at the capability boundary rather than marketing “it understands people.” Treating exclusion as only a communications risk misses that the first fact is an unfinished task.
Applying it
- If voice is how a critical task is completed, provide an equivalent path that does not depend on this recognizer until completion rates for the gap group approach the majority.
- State supported languages and varieties in release notes; do not imply that uncovered accents will work if people “just speak more slowly.”
- Watch abandonment by accent or region; treat a rise as an availability incident, not as users who do not know how to talk.
Related
- Same group: C7.04.1 Training-data distribution determines whose speech is recognized well · C7.04.2 Recognition degrades on mixed Chinese–English and dialect words · C7.04.4 Mid-utterance language switches require live language identification · C7.04.5 A wrong language decision forces the rest of the sentence through the wrong phonology · C7.04.6 Dialects share vocabulary with the standard but differ in pronunciation, so they are misread as near-homophones · C7.04.7 Accent and dialect gains depend on training data from those speakers, not on algorithms alone
- Adjacent: C7.06 Hands-Free and Eyes-Free Use · C7.03 Types of Recognition Errors
- Search:
speech equity·demographic exclusion·ASR disparity
Cards in the same group
- C7.04.1Training-data distribution determines whose speech is recognized well
- C7.04.2Recognition degrades on mixed Chinese–English and dialect words
- C7.04.4Mid-utterance language switches require live language identification
- C7.04.5A wrong language decision forces the rest of the sentence through the wrong phonology
- C7.04.6Dialects share vocabulary with the standard but differ in pronunciation, so they are misread as near-homophones
- C7.04.7Accent and dialect gains depend on training data from those speakers, not on algorithms alone