D5.07.1Temporal integration windowdesignresearch

Multimodal input must fall within one temporal window to count as one intention

Aliases: multimodal fusion window · input binding · temporal contiguity

What it is

Multimodal inputs are interpreted as one intention only when they are temporally close enough. The constraint is often called a temporal integration or fusion window. “Put that there” plus a pointing gesture, or a spoken command plus a confirmation press, has a temporal shape. The window is not millisecond simultaneity; it accommodates the normal order in which people express one intention.

Why it happens

Temporal contiguity is a strong cue for deciding whether signals belong to the same event. Speech, gesture, gaze, touch, and buttons each have an onset, peak, and completion. If a fusion layer sees only queue-arrival times, asynchronous evidence becomes unrelated. A useful window spans the natural interval from intention onset to the completion of the supporting signal while excluding residue from the previous task. It therefore aligns recognition latency, sustained input, and human pacing, not merely arrival timestamps.

Studying it

Use controlled-delay paradigms to locate the boundary: present speech and pointing, tap and button, or gaze and gesture in one task while systematically varying the interval between onsets or completions. Record whether evidence is fused, split into two intentions, or uninterpretable. Variables include order, signal duration, recognition latency, and task context; outcomes include binding accuracy, command-completion time, repeated input, and selection of the wrong object. Field logs can supplement this, but deliberate consecutive input must be distinguished from device-induced delay.

Where it stops holding

One value cannot span all pairings. An utterance may last seconds; a haptic click lasts tens of milliseconds. Gaze as a prior cue relates to confirmation differently than two consecutive keys do. Motor-control differences, speech hesitation, and assistive-device latency further shift the natural span. Nor can a window replace semantic checking: even closely timed signals should trigger clarification when the pointed object and spoken content conflict.

Applying it

  • Record onset, completion, and recognizer-return times for each modality, and fuse over these states rather than queue-arrival time alone.
  • Set onset offsets, maximum span, and reset conditions per pairing; reuse defaults only for the same task and comparable latency.
  • Show waiting, expired, conflicting, or partially missing evidence so users know which input the system still expects.
  • Verification: instrument real input streams, measure one intention split into two and two intentions merged, then retest boundaries under simulated delay.

Related

  • Within the group: D5.07.2 Too narrow a window misreads normal sequential expression as separate commands · D5.07.4 Window length should be tuned per modality pairing, not fixed universally
  • Adjacent: D5.02.2 Fusion needs temporal alignment · D4.07.2 Output latency breaks perceived causality between events
  • Search terms: multimodal fusion window · temporal contiguity · input binding

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/D5.07.1