Too narrow a window misreads normal sequential expression as separate commands
Aliases: window too short · split intention · sequential multimodal input
What it is
When the window closes too early, natural orders—speak then point, look then confirm, select then modify by voice—are split into separate inputs. The first part may fail for lack of an object or parameter; the second lacks the first part's context. The result can be two apparently valid but incomplete commands.
Why it happens
People often emit a primary signal before a supporting signal completes. Speech deixis needs gesture to resolve its object; gesture needs speech to select an action. A narrow window mistakes “supporting input not finished” for “not part of this event” and closes interpretation immediately after the primary signal. Premature closure also leaves residue: the first part is staged or rejected, then the second triggers an unrelated action. Users experience this as incomprehension, not delay.
Studying it
Ask participants to speak then point, gaze then select, or coarsely select then revise by voice at their natural pace; delay the auxiliary input until it just crosses candidate windows and compare natural and forced segmentation. Variables include order, utterance length, recognizer completion time, and practice. Outcomes include the number of split intentions, error rate of the second command, repetitions, and immediate corrections. For tactile or button tasks, measure the distribution from first-part completion to confirmation rather than only the mean.
Where it stops holding
Not every sequence should merge. A pause after finishing one action, a task switch, or “cancel” followed by a new command is deliberately separate. The narrow-window problem applies when both parts concern the same object or parameter and the interval lies within normal expression. Safety-critical commands still require explicit confirmation; if waiting keeps the wrong object highlighted, show “waiting for completion” rather than implying execution.
Applying it
- Keep a brief open state for speak-then-point and similar pairings, and show the candidate object or action while awaiting completion.
- Set lower bounds from users' actual interval distributions and include recognizer time in the window instead of starting at message arrival.
- Turn an unfinished first part into a cancelable prompt at closure rather than executing or silently discarding it.
- Verification: have participants operate naturally without rehearsal, then locate splits, wrong executions, and repeated inputs.
Related
- Within the group: D5.07.1 Multimodal input must fall within one temporal window to count as one intention · D5.07.3 Too wide a window fuses unrelated input
- Adjacent: D5.02.3 Modality combinations need explicit fusion-failure handling · D3.08.1 Haptic cues need temporal alignment with visual focus
- Search terms:
premature fusion·split intention·sequential multimodal input