M4.01.5speech output more exposing than inputdesignresearch

Spoken replies leak more than spoken commands

Aliases: TTS overhearing · loud readback · reply-side disclosure

What it is

In a hotel room someone murmurs “tomorrow’s itinerary” at a phone. The loudspeaker answers at living-room level: “7:40 a.m. flight to Hongqiao, meeting with attorney Wang at three, bring the contract.” The corridor and the next room hear the second sentence. Speech output more exposing than input: the user can omit, hedge, drop to breathy voice; the system, chasing intelligibility, reads slots in full, loud, as complete noun phrases. The leak’s peak is often the reply, not the request.

Why it happens

On the input side the person still holds the encoding: “that meeting,” a whisper, switching to typing halfway. On the output side wording, length, and target level are written by the synthesizer and the product policy, optimized for confirmation and being heard, not for staying inside the walls. Confirmation readback is the sharp case — it rebroadcasts the most sensitive slots (time, name, account tail, message body) at a designed amplitude. Output is also often longer: a short command buys a calendar digest. Thin hotel walls, open benches, car speakerphone: room SNR is enough for the reply to be the clearer end. Barge-in only helps if it arrives before the slot; the slot is often in the first clause.

Studying it

At the same distance in the same room, compare bystander intelligibility of the user’s utterance versus the system’s reply. Script a short command / long readback pair; have someone in the hall or the next seat transcribe; score by slot. Independent variables: speaker level, headphones on or off, full versus clipped readback. In diaries ask “what do you think was heard — what you said or what it said.” If most point at the reply, the asymmetry is already in the wild. Do not take the owner’s “I spoke very quietly” as the evidence.

Where it stops holding

Headphones all the way, and the system off the loudspeaker, shut this output-side asymmetry; input-side exposure remains. Screenless driving that must hear a confirm has a safety reason for intelligible readback; heading cannot be mumbled for the bystander’s sake. Public-address and radio output is meant for many listeners. Collapsing every long reply to “done” blinds the owner too. The asymmetry holds when the system is louder or more complete than the user; a shouted command can reverse it.

Applying it

  • Slots that carry names, addresses, calendar detail, or message bodies should not be read in full over a loudspeaker by default. Prefer “added an item at three tomorrow” or a lock-screen summary.
  • In public scenes or with no headset, use a short confirm; expand only when the user asks to hear it.
  • Barge-in must kill the speaker before the first slot finishes, not after the whole digest.
  • How to check: door closed, someone in the hotel corridor, play a real itinerary reply. If the hall can write down the lawyer’s surname or the flight number and cannot catch the user’s short command, output is the exposing channel and readback policy has to change.

Related

  • Same group: M4.01.1 Speaking to a dialogue system hands the content to everyone in earshot · M4.01.2 Addressing a device in public is still socially marked · M4.01.3 Social cost drives people to abandon voice in public · M4.01.4 Bystanders pay a cost for being forced to overhear · M4.01.6 Headphones privatize the reply, not the request
  • Nearby: M3.11 Trimming speech output · M2.08 Explicit and implicit confirmation · C7.07 Privacy visibility of voice input
  • Search terms: speech output more exposing than input · readback overhearing · TTS leakage

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/M4.01.5