M4.04.1age-stratified word error ratedesignresearch

Word error rate is substantially worse for children and older adults

Aliases: child and older-adult ASR · age-banded WER

What it is

In a classroom, the same command set on a teaching assistant yields a markedly higher word error rate for the children than for the adult at the front. In a rideshare, a passenger in their seventies giving a destination will see the same recognizer swap in the wrong street more often. Age-stratified word error rate is not a footnote: children and older adults sit well below the single curve products quote. The gap is in the acoustic distribution, not in “they don’t know how to use it.” A model tuned on adult-modal speech is a different input for the two tails.

Why it happens

Mainstream ASR trains and decodes around young-adult modal speech: F0, formant centers, glottal source, all calibrated to that cloud. Children’s F0 is higher and the vocal tract shorter, so formants shift up and the same phone lands outside the adult model’s low-frequency mass. Older speakers more often bring unstable vocal-fold vibration, lower energy, and softer consonant closures — a deviation on the other side of the spectrum. The decoder still hunts the nearest word in adult acoustic space; distance grows, substitutions and deletions rise together. That is out-of-distribution input, not extra noise on the same distribution. Classroom and cabin far-field noise push both tails farther, but the split already exists in quiet close-talk, so the whole gap cannot be billed to the room.

Studying it

Run age-stratified WER on one command set: children, younger adults, older adults. Report per-group word error, sentence error, and N — not one merged number. Control vocabulary (the same street names, the same classroom directives) so content difficulty is not tangled with age. Independent variables: age band, far field or not, whether an adult models the phrase. Classroom live use and in-car destinations are the ecological conditions; a lab read list can split acoustic gap from task gap. “The kids thought it was fun” is not a substitute for WER — enjoyment and being heard are different dependents.

Where it stops holding

Systems adapted on children’s speech, or older adults on a long-used close-talk device, can narrow the split; do not assume it is zero. Whisper, post-surgical larynx, voices off the modal gender line also sit off the training center; age is not the only axis. Writing the gap as “older people are slow, children are slurred” moralizes an acoustic-distribution problem and replaces WER with a judgment of the user. Handheld close-talk in hard noise will also collapse adult WER, and relative positions can move — remeasure by scene instead of reusing the classroom delta.

Applying it

  • Any skill whose primary users include children or older adults (classroom assistant, car navigation, hotel-room controls) must report WER for those two groups. A single overall recognition rate is not an acceptance packet.
  • Prefer command words that are acoustically stabler at both tails (avoid street names that hinge on a fine vowel contrast), and give a silent fallback.
  • How to check: take the twenty demo commands and run them with a primary class and with an older passenger group. If sentence error is still well above the internal adult test group, “recognition is ready” does not hold for the tails. Change the model or the channel, not the politeness of the prompt.

Related

  • Same group: M4.04.2 Recognition gaps become de facto exclusion · M4.04.3 Evaluate by cohort, not by the overall number
  • Nearby: M4.09 Recognition differences for children and older adults · C7.04 Accent, dialect, and code-switching · M1.01 When voice-first is appropriate
  • Search terms: age-stratified word error rate · cohort evaluation · child ASR

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/M4.04.1