Frequency-close, temporally continuous sounds fuse into one auditory stream
Aliases: primitive stream integration · sequential grouping
What it is
A sequence of sounds played one after another, if their pitches are close together and there is no long silent gap between them, gets automatically welded into a single auditory stream — a perceptual unit treated as "one ongoing source," rather than a string of unrelated sound events. This is the most basic organizing rule in auditory scene analysis.
It is easy to misread as "one physical source equals one stream." The opposite holds: streaming is governed by perceptual rules, not physical source count. A single real source producing sounds with large, abrupt pitch jumps can still be heard as two streams; conversely, two different sources can fuse into one stream if their pitch and timing happen to line up.
Why it happens
This is the auditory analogue of Gestalt grouping. Frequency proximity (small pitch steps between successive elements) plus temporal continuity (short gaps, no long silences) together provide evidence that the elements likely come from one ongoing sound-producing process, so the auditory system binds and tracks them as a single perceptual object instead of processing each in isolation.
This is not passive similarity-matching — it is inference. A real vibrating source (a string, a speaker's vocal folds) tends to change pitch smoothly and gradually rather than jumping wildly within tens of milliseconds; the auditory system exploits this regularity of the physical world to infer "smooth change implies same source." The smaller the frequency step and the shorter the gap, the stronger this evidence, and the more tightly the stream binds; once the step grows or the gap lengthens, the evidence weakens and the sequence tends to split into multiple streams instead.
Studying it
The standard paradigm presents a sequence of alternating or gradually shifting tones, manipulating the frequency separation between successive tones and the inter-tone gap, then asking listeners whether they hear "one melody" or "two separate lines." An indirect, objective measure is often used instead of subjective report: listeners judge a rhythmic pattern embedded across the sequence — if the sequence has already split into two streams, cross-stream rhythmic relations become hard to judge, and a drop in accuracy is taken as evidence that segregation occurred.
Typical independent variables: frequency separation between adjacent elements, element duration and gap (which together set presentation rate), and total sequence duration (streaming builds up gradually rather than occurring instantly). Typical dependent variables: subjective stream count, accuracy on rhythm/order judgments, reaction time.
In interface research this paradigm is commonly used to test whether a sequence of cue tones, synthesized speech segments, or layered background audio will be heard by users as one coherent whole or as disconnected fragments.
Where it stops holding
- Frequency separation and presentation rate act jointly: the same frequency gap may stay fused at a slow rate and split apart at a faster one, so a frequency-separation number alone is meaningless without specifying rate (detailed in the next entry).
- Streaming needs time to build up: the first few hundred milliseconds of a sequence are often still heard as one stream, and segregation only stabilizes after the sequence has continued for a while — a couple of isolated tones cannot reveal the pattern.
- This describes the default, bottom-up outcome. When a listener actively tries to track a particular sub-sequence, the reported stream boundaries can shift, showing this organization is not entirely immune to attention.
Applying it
- To make a sequence of cue tones, spoken prompts, or layered background audio sound like one continuous thing, keep the pitch steps between elements small and the gaps short, and avoid an obvious silent break partway through.
- Conversely, when two channels of information are meant to be perceived as separate and unrelated (e.g., background media audio versus a system notification), widening the pitch gap or loosening the timing does more work than changing timbre alone.
- How to check: play the finished audio sequence to naive listeners and simply ask whether they hear one continuous sound or several alternating ones; use the reported stream count as the criterion, not the designer's own impression.
Related
- Same group: A3.15.2 Fast alternation between two pitches splits into two independent melodic lines · A3.15.3 Stream formation depends on temporal regularity; irregular gaps hinder it · A3.15.4 Simultaneous sources need onset-time differences to be separated
- Nearby: A3.04 Auditory masking (answers "can it be heard," a different question from "which stream does it belong to") · A3.06 The cocktail party effect (top-down selective attention layered on top of already-segregated streams) · A3.13 Temporal resolution and rhythm perception
- Search terms:
auditory stream segregation·auditory scene analysis·primitive grouping·frequency proximity