Localization relies on interaural time and level differences
Aliases: ITD · ILD · interaural cues
What it is
Without looking, people can tell roughly whether a sound comes from the left or the right, from the front or behind, using two kinds of difference between what each ear receives: the interaural time difference (ITD), how much later or earlier the sound arrives at one ear versus the other, and the interaural level difference (ILD), how much louder or quieter it is at one ear versus the other. Together these form the binaural localization cues, the main physical basis for horizontal sound-source localization.
It is easy to oversimplify this as "whichever ear hears it first is that side" — that only captures the time-difference half. Level difference is an independent second cue, and the two carry the load in different frequency ranges.
Why it happens
If a source is not directly ahead or behind, the path length to each ear differs, producing a tiny difference in arrival time — this is ITD, and people use this difference to judge which side a sound leans toward. But this cue only works well at low frequencies: for long-wavelength low-frequency sound, the phase difference between the waveforms at the two ears can be tracked reliably by the nervous system and converted into an arrival-time difference. Once the wavelength gets shorter than head size, the phase difference becomes periodically ambiguous — the same phase difference could correspond to several different actual time differences — and the time-difference cue breaks down.
Level difference works the opposite way, becoming prominent at high frequencies: the head itself obstructs sound, and short-wavelength high-frequency sound is more easily blocked by the head, casting an acoustic shadow that makes the far ear receive noticeably less energy. Long-wavelength low-frequency sound diffracts around the head easily, so the two ears receive nearly identical levels, and the level-difference cue does little at low frequencies.
The two cues therefore divide labor by frequency: low frequencies rely on time difference, high frequencies on level difference — together known as the duplex theory of localization.
Studying it
The typical method delivers separately controlled signals to each ear over headphones, imposing a precise, artificial time difference or level difference (decoupled from the way both naturally co-occur in real space), and asks listeners to report which side the sound seems to come from and by how much. This isolates how much perceived shift each cue produces on its own, and how the weighting of the two cues shifts with frequency.
Typical independent variables: the imposed size of time difference or level difference, and the test frequency. Typical dependent variables: subjectively reported direction and magnitude of lateral shift (sometimes measured as minimum audible angle).
Where it stops holding
- These two cues resolve left versus right; for front-back and up-down directions, the same combination of time and level difference can correspond to several different positions, and binaural cues alone cannot tell them apart — expanded in the next entry.
- In the mid-frequency range (near the wavelength matching head size), neither cue is very reliable — time difference is starting to become ambiguous while level difference is not yet large enough — making this a transitional band where binaural localization performs relatively poorly.
- This mechanism assumes both ears receive sound normally. Listening with only one ear leaves neither difference able to form at all, detailed in another entry in this group.
Applying it
- When encoding direction for stereo or spatial audio, relying purely on inter-channel volume difference (traditional level-panning) works well for high-frequency content but produces almost no perceived directional shift for low-frequency content — which is also why a subwoofer's physical placement usually does not need to be strictly centered without noticeably hurting spatial accuracy.
- For full-spectrum, higher-precision directional playback (e.g., needing to clearly convey which side a low-frequency warning tone comes from), level-panning alone is not enough — real interaural time-difference encoding is needed (binaural rendering, head-related-transfer-function-based processing) rather than making do with level difference alone.
- How to check: present the same low- and high-frequency material using pure level-panning versus panning with added time-difference encoding, and have listeners report directional accuracy for each, confirming whether low-frequency content is indeed poorly localized under the level-only approach.
Related
- Same group: A3.05.2 Front-back and up-down localization are the least precise · A3.05.3 Mono output loses all localization information
- Nearby: A3.06 The cocktail party effect (directional difference is used there as a cue for boosting selective attention; the underlying mechanism belongs here)
- Search terms:
interaural time difference·interaural level difference·duplex theory·binaural localization