M1.01.2short-command bias of voicedesignresearch

Voice prefers short commands over long procedures

Aliases: one-shot command · short utterance · slot-filling cost

What it is

In spoken dialogue, a short command (ten-minute timer, play this album, lights off in the living room) holds up better than a multi-slot procedure (book a flight: date, flight, seat, ID, payment). Short commands close in one or two turns; long procedures unpack a form into a chain of questions in time. The channel can carry the long chain; the cost curve is unkind. This is not “voice cannot fill forms.” It is that every extra slot is another chance to be misheard, interrupted, or to lose the previous step.

Why it happens

The planning cost of a spoken command sits before the mouth opens: the goal has to be compressed into one utterance in working memory. “Timer, ten minutes” is already in the mouth. “Economy, window if possible, Thursday morning, pay with points” forces a decision about what to say first; omissions wait for the system to chase. Each chase reloads the just-filled material from memory and stacks substitution errors on top. A graphical form leaves unfilled items as empty cells in space; in dialogue an unfilled item is a turn that already slid past, kept only if the system remembers — or if the user remembers what they already said.

Short commands have a second structural edge: repeating the whole utterance on failure is cheap. Fail a long procedure on the third slot and the status of the first two becomes a new conversational burden. “Voice-first” is therefore skewed in length: the shorter the goal, the more it deserves to be the default channel; the longer it gets, the sooner it should be handed to a screen or another path.

Studying it

Implement the same goal as a single-turn command and as multi-turn slot filling. Compare completion, turn count, mid-task abandonment, and slot rollback. Independent variables: number of required slots, how open each slot's vocabulary is (closed list versus open place names), and whether users may pack several slots into one utterance. Do not stop at task success: note the turn at which people switch to “forget it” or pick up the phone.

A cleaner cut on live traffic is to archive sessions by closing length: user turns from wake to task done. If the mass sits at 1–2 turns, the product is being used as a short-command channel. Forcing a path past five turns usually lifts abandonment in the middle. Scripted lab procedures that read like a questionnaire undercount spontaneous interruption and self-repair.

Where it stops holding

Experts chunk a frequent long procedure (the same parameter bundle to a warehouse system every day); “long” has been practised into short. On a screenless device there is no screen to hand off to, so the long procedure is not “a poor fit” so much as “no better path” — tighter slot confirmation and a spoken “start over” become mandatory. One-shot high-stakes decisions (lending, triage) will be lengthened by legal and safety demands even when the user would accept more turns; the short-command bias yields to consequence. Collapsing every capability into “just say it and I will do it” guesses missing slots; the benefit of shortness is eaten by wrong execution.

Applying it

  • Inventory skills. What can close in one or two turns becomes a voice default. What needs three or more required slots defaults to an on-screen form; voice jumps in or fills one or two slots.
  • Let people pack several slots into one utterance, but do not make “say everything at once” the only grammar. Ask for what is missing; do not re-ask what was given.
  • After the second or third chase in a long procedure, offer an exit: “I'll finish the rest on the phone.” Do not keep someone in a fourth turn while pretending this is still a short command.
  • How to check: histogram user turns from wake to done, per skill. A median above three means cut slots or mark a handoff to screen. The slot where abandonment spikes is where length stopped paying.

Related

  • Same group: M1.01.1 Voice earns its keep when hands and eyes are already taken · M1.01.3 Comparison and browsing tasks are a poor fit for voice
  • Nearby: M1.03 Dialogue turns · M2.07 Dialogue flow and state design · M2.13 Multi-intent and compound commands
  • Search terms: short-command bias of voice · slot-filling cost · one-shot command

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/M1.01.2