M3.07.1grouped pauses for digit stringsdesignresearch

Digits and addresses need grouped pauses

Aliases: chunking pause · spoken telephone grouping · OTP readout grouping

What it is

A locker that utters a pickup code as “three nine four eight two one” in one breath hands the listener six equal sounds, not the three groups “39, 48, 21.” Card numbers, one-time codes, house numbers, and numbered road names work the same way: the listener needs a chunking pause at conventional cuts, not a reading of digits as ordinary words. The issue is how a sequence is sliced internally, not the comma-and-period pauses that mark sentence structure.

Why it happens

Digit strings offer almost no semantic cue for the next token. The phonological loop holds few unstructured items; conventional grouping (mobile 3-4-4, codes in pairs, addresses at road / lane / number) packs six to eleven tokens into two or three chunks that recall can keep. The gap closes a chunk: silence after a group lets the listener lift those digits out of the stream and free the loop for the next group. Without the gap, later digits overwrite earlier ones and the middle falls out first — pickup codes fail on the inner pair more often than on the edges.

The cuts have to match groupings the listener already uses. Slicing an eleven-digit mobile number 2-2-2-2-3 supplies enough gaps but the wrong chunks, so rehearsal still fails. In an address, the boundary between a named road and an Arabic house number is a chunk boundary too: do not break the road name; do break before the number.

Studying it

Grouped recall: the same pickup code, card number, or house number in three readings — ungrouped, pauses at conventional cuts, equal gaps at the wrong cuts. Immediate spoken or written report. Dependent measures: whole-string accuracy, inner-digit errors, spontaneous replay requests. Independent variables: string length, digits per group, inter-group silence (on the order of 150–400 ms).

In the wild, align “digit-string playback → user repeat or keypad entry.” Inner-digit errors well above edge errors, with replays clustering on unstructured long strings, point to missing grouping rather than a voice-quality problem. A lab that prints the digits on a card while people listen is measuring proofreading, not auditory chunking.

Where it stops holding

Two digits, known to be two (“which table”), have nothing to group. A number the user just spoke already sits in their own chunks; even a run-on readback can match. A road name the car has spoken all week is one lexical item and does not need internal cuts. Treating grouped pauses as “read digits slowly” lengthens every token without restoring chunk boundaries — slow and uncut, the middle still drops.

Applying it

  • Write a cut template per class: phone, OTP, card, order id, house number. Insert silence only at those boundaries; do not pad every digit equally.
  • Cut addresses at city / district / road / lane-and-number. Finish the road name as one block before the house number; leave a gap between lane and number.
  • When reading back a number the user just said, keep their grouping; do not swap in another scheme.
  • How to check: take real pickup codes and mobiles, ungrouped versus conventionally grouped, then immediate dictation. Inner-digit accuracy should rise on the grouped version. If you only slowed the same cuts, the work is not done.

Related

  • Same group: M3.07.2 First listen and replay want different rates · M3.07.3 Final contour does the turn-yielding
  • Nearby: M3.01 Naturalness and Intelligibility of Synthetic Speech · M3.02 Speech Rate and Prosody · M1.02 Memory Load of Screenless Interaction
  • Search terms: grouped pauses for digit strings · chunking · telephone number grouping

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/M3.07.1