A3.15.4Onset asynchronyresearchdesign

Simultaneous sources need onset-time differences to be separated

Aliases: common onset · concurrent sound segregation

What it is

The previous entries deal with sequential sounds and which stream they join. This one deals with simultaneous sound: if several sound components start at exactly the same instant, the auditory system readily treats them as one whole and hears a single fused sound. But if just one component's onset comes even a few tens of milliseconds earlier or later than the rest — a difference small enough to be nearly imperceptible on its own — that component gets "pulled out" of the whole and heard as a separate, individual sound. This difference in start time is called onset asynchrony, and it is one of the strongest cues the auditory system uses to separate simultaneously sounding sources.

This is distinct from how sequential sounds get bound into one stream, covered earlier. Those entries deal with attribution across a sequence in time; this one deals with pulling apart sounds that are stacked on top of each other at the same moment.

Why it happens

In the real world, all the sound components produced by a single physical event almost always start vibrating at the same instant — pluck a string, and its fundamental and every harmonic begin together. The auditory system treats perfectly synchronous onset as default evidence that these components share one source, and so tends to merge synchronously starting components and process them as a single unified timbre.

Once one component's onset falls out of sync with the rest, that default evidence is broken: an asynchronous onset suggests this component likely comes from a separate, independent vibrating process. Even if its frequency happens to form a perfectly clean harmonic relationship with the rest — normally the strongest possible evidence for a shared source — a mismatched onset time still tends to make the auditory system pull it out on its own. This shows that onset-timing evidence carries considerable weight, capable of overriding the pull toward fusion that harmonic frequency relations would otherwise produce.

Studying it

The classic method builds a harmonic complex tone (a fundamental plus several harmonics presented together) and shifts the onset of one harmonic component earlier or later relative to the rest, manipulating the size of this onset asynchrony. The measure is whether listeners can hear that component out as a distinct timbral element (report it as separate), or how strongly it affects overall timbre identification as the asynchrony grows.

In interface research this logic is commonly used to assess whether, when multiple sounds are triggered at once, their onsets need to be deliberately staggered so that each is heard as a distinct, identifiable event rather than blurring into one mass.

Where it stops holding

  • Onset asynchrony is only one of several simultaneous-grouping cues. When other cues are already strong — spatial location, or a timbre clearly outside any harmonic relationship — the contribution of onset asynchrony is weakened or masked, and it should not be considered in isolation.
  • If the asynchrony grows too large, the percept changes into something else entirely: no longer "one component pulled out of a fused whole," but two fully separate, sequential events — a different perceptual experience from the one described here.
  • Reverberant environments smear the onset itself (the attack edge of a sound gets stretched out by reflections), so this cue becomes markedly less usable in strongly reverberant settings and should not be assumed to work equally well in every acoustic environment.
  • This entry concerns whether a component can be heard as a distinct sound, not which of several simultaneous sounds overpowers or masks the others in loudness — that is a separate, level-based question with different criteria, and the two should not be conflated.

Applying it

  • When multiple cue sounds may be triggered at once (overlapping notifications, stacked alerts), staggering their onsets by a few tens of milliseconds does more to keep each one heard as a distinct, identifiable event than changing timbre alone.
  • When designing a set of sounds meant to play together while each remains individually recognizable, avoid making all components start in strict synchrony — even if their frequencies and timbres are already well differentiated, synchronous onset still reduces how separable they sound.
  • How to check: play candidate sounds to listeners both perfectly synchronized and staggered by a chosen number of milliseconds, and ask how many distinct sound components they hear; use the reported count, not the designer's own judgment, to confirm whether the stagger is large enough.

Related

  • Same group: A3.15.1 Frequency-close, temporally continuous sounds fuse into one auditory stream · A3.15.2 Fast alternation between two pitches splits into two independent melodic lines · A3.15.3 Stream formation depends on temporal regularity; irregular gaps hinder it
  • Nearby: A3.04 Auditory masking (a loudness-level question of whether something can be heard at all, distinct from whether it can be told apart)
  • Search terms: onset asynchrony · common onset · concurrent sound segregation · harmonic mistuning

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/A3.15.4