M3.09.3Lombard escalation after ignored barge-indesignresearch

Ignored barge-in makes people raise their voice

Aliases: shouting over TTS · Lombard loop · ignored interrupt

What it is

Nav is still saying “continue on this road for eight hundred metres.” The user said “stop” and it did not stop; the second attempt is louder, the third close to a shout. That is not temper. It is the Lombard effect: when the other party keeps making sound, speakers treat themselves as unheard and raise level, pitch, and duration. Lombard escalation after ignored barge-in is that loop on a dialogue system — the system’s own playback is the noise that will not yield.

Why it happens

Lombard adjustment is automatic: if reception is not heard, the production end adds gain. Continuing TTS is both a masker and a social signal that the floor has not been given up. Once level rises, near-end punches further into the loudspeaker’s echo, AEC has a harder time separating voice from playback, detection keeps failing, and another step of gain is added — the loop is acoustic, not merely attitudinal. Over-level also clips the recognition front end, so even a successful stop leaves “stop” harder to recognise.

Ignored barge-in also changes the wording: short commands become whole-sentence repeats, words become shouted names. Those look like new intents; they are the same intent sent at a higher setting. Treating the upgrade as a new request or as affect sends repair down the wrong path.

Studying it

Run ignored-barge-in trials: deliberately fail to stop on the first insert, then measure SPL, F0, duration, and spectral tilt on the second attempt against the first. Dependent measures: the Lombard feature jump, whether a third attempt appears, and whether AEC hit rate falls on the second insert. Independent variables: playback SPL, whether the first attempt got any stop at all (even 200 ms of silence before resume).

In production, look at the energy gap between two near-end events in the same session. A clearly louder second with no stop in between is escalation, not two unrelated commands. A lab that replaces the second attempt with a button will not see Lombard. Calibrate on within-session relative jump, not absolute dB across people.

Where it stops holding

Users who are already loud (shop floor, windows down) start in the Lombard region; a further ignore may not trace the same rise, but clipping risk remains. An earpiece takes playback out of the room; ignore shows up as saying it twice, not as shouting. Social cost in public can suppress the level jump; people abandon or reach for a button — it looks like “no Lombard,” but the channel was socially closed. Treating the raised voice as rudeness sends work into scolding the user and misses that one stop would cut the loop.

Applying it

  • On suspected near-end, stop playback first, even before recognition finishes. A couple of hundred milliseconds of yield is enough to block most second-attempt escalations.
  • When second-attempt energy in the same session is clearly above the first, treat it as a Lombard resend of the same command, not a new intent, and do not answer with “please speak more slowly” while still holding the speaker.
  • On loudspeaker skills, test escalation: miss the first “stop” on purpose, watch second-attempt SPL and clipping. Where the jump is large, fix stopping before fixing recognition.
  • How to check: record SPL on attempt one and two. No stop in between and a clear rise means the loop is on. After “stop then recognise,” the jump should fall; if it does not, the stop is too late or playback resumes immediately.

Related

  • Same group: M3.09.1 Barge-in detection is capped by echo cancellation · M3.09.2 After barge-in, tell correction from a new request
  • Nearby: M3.04 Barge-in · C7.05 Noisy Environments · C7.12 Recognition Degradation in Noise
  • Search terms: Lombard escalation after ignored barge-in · Lombard effect · shouting over TTS

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/M3.09.3