C7.07.2Bystander speech must not be treated as the userdesignresearch

Bystanders must not be treated as the user

Aliases: bystander · voiceprint gate · unauthorized capture

What it is

In a living room, an office, or at a counter, the microphone hears more people than the owner. Treating a bystander’s utterance as the current user’s input is both a functional error and a way of dragging someone who did not consent into recognition and possible cloud processing. Bystanders must not be treated as the user: by default they are not enrolled, not executed, and their audio is not a lawful session. That is a privacy boundary, not only a source-separation accuracy contest.

Why it happens

A far-field capture volume is larger than the social act of “talking to the machine.” A wake word said by a television or a guest, a guest interrupting while the session window is still open, or “whoever is loudest is the user” when no voiceprint is on, all send bystanders into ASR. Once decoded, text may land in history, feed personalization, or execute (change a calendar, send a message). The bystander did no entry action and still takes on the data consequences of being the user. Voiceprints, direction, and “facing the device” lower probability; they do not replace consent in a legal sense. Compared with attribution under overlap, the emphasis here is that an unauthorized person must not enter the user identity, not which of two streams should run.

Studying it

Stage guest speech, television voices, and people outside the window. Measure whether a session opens, whether a transcript is produced, whether anything executes, and whether audio leaves the device. Compare speaker verification on and off. Interview bystanders about whether they knew they were heard. Inviting only the owner to finish tasks measures the bystander path away.

Where it stops holding

A counter, a classroom clicker, or a car in which every passenger may issue navigation commands intentionally treats several people as legitimate users. That has to be stated in the scene, and people present must see capture. Children’s voices often fail adult-trained voiceprints and are either shut out as bystanders or accepted as executable users—both failures need their own tests. A wake word inside a public broadcast is a content problem and a bystander problem.

Applying it

  • Bind commands by default to an enrolled speaker or to the direction that just said the wake word; sources that fail the gate stay on local keyword spotting and do not enter full recognition.
  • Mark uncertain-source items in session history and let the owner delete them; do not write guest speech into a personalization profile.
  • In acceptance tests, have an unenrolled person talk beside the device and confirm no executable command and no syncable transcript.

Related

  • Same group: C7.07.1 Capture state must be visible and the indicator must not be disableable · C7.07.3 Speaking the content in public is itself a disclosure
  • Adjacent: C7.05 Noisy Environments · C7.08 False Wakes
  • Search: bystander · speaker verification · unauthorized capture

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/C7.07.2