Speaking to a dialogue system hands the content to everyone in earshot
Aliases: spoken disclosure · overhearing · public dialogue leak
What it is
Opening a dialogue speaks the content into every ear in the room. In a rideshare, “navigate to 14 Maple, apartment 3B, then text Dr. Chen I am on my way” gives the driver — and the next pooled passenger — not merely the fact of voice use, but an address, a unit number, and a social tie. That is overheard dialogue content. Server-side encryption and on-device recognition do not stop the sentence in the air. The claim is about dialogue as a public performance, not about whether a capture light is on, and not about whether secrets should be designed as spoken input.
Why it happens
Speech has no view cone: it falls off with distance and room reverberation, not with “this is only for the device.” Dialogue also forces full referring expressions — a human passenger can nod at “the usual place”; a recognizer needs a noun phrase it can match. The slots least fit for broadcast (place, person, time, body) are exactly the stable words in the command. Harm tracks content class: a next-turn instruction costs almost nothing; a home address and a recipient hand a life-structure to the cabin. Users can drop volume, but task success still needs intelligibility above the recognition floor, and that floor is usually enough for the person in the next seat.
Studying it
An in-situ diary of public voice use: at each intended wake, log the place, who is within an arm’s length, the content class about to be spoken (place / person / calendar / non-sensitive), and whether the speaker hedged, switched to a deictic, or typed instead. Dependent measures are identifiable slots actually spoken and who the speaker later thought had heard — not laboratory word error.
A quieter cut is concealed intelligibility in a parked cabin or matched soundscape: scripted commands with known slots, listeners at varied distances repeating them, scored by content class. Anechoic recognition rates will not show this.
Where it stops holding
A solo drive with the windows up, an empty elevator, or a trip whose address has already been given to the driver: spoken exposure is near zero or already spent. Counter service and classroom call-outs are meant to be heard. Whispering lowers intelligibility; it does not cancel it, especially for digits and names. Labeling all voice as unfit for public use would ban “turn right at the next light,” which is already public. Harm follows content class and who is present, not the mere fact of speaking.
Applying it
- Tag skills by content class. Those that carry home addresses, contacts, calendar detail, or health slots should not be wakeable as public skills by default, or should require a screen or headphone confirm before they run.
- Allow vague references to bind to on-screen focus or to an already confirmed object, so people are not forced to re-broadcast the full name to be understood.
- Do not put “please say your full home address” in example prompts; public-scene examples should use non-identifying content.
- How to check: walk the target skills in a real rideshare or an equally crowded cabin. If finishing the task requires speaking a unit number, a name, or a medical condition clearly, the skill is not yet a public dialogue.
Related
- Same group: M4.01.2 Addressing a device in public is still socially marked · M4.01.3 Social cost drives people to abandon voice in public · M4.01.4 Bystanders pay a cost for being forced to overhear · M4.01.5 Spoken replies leak more than spoken commands · M4.01.6 Headphones privatize the reply, not the request
- Nearby: C7.07 Privacy visibility of voice input · M1.01 When voice-first is appropriate · M4.07 Felt privacy of always-on microphones
- Search terms:
overheard dialogue content·spoken disclosure·bystander