Timbre is set by the harmonic structure above the fundamental, and is the primary cue for telling sources apart
Aliases: spectral envelope · harmonic structure · tone color
What it is
A piano and a violin playing the same pitch at the same loudness are still instantly distinguishable — not by pitch, not by loudness, but by timbre. Timbre is carried by the relative strength of the harmonics (overtones) sitting above the fundamental frequency; the shape of that distribution across frequency is called the spectral envelope. Once fundamental and overall energy are held equal between two sounds, whatever still lets a listener tell them apart lives entirely in this harmonic structure.
Timbre is often reduced to "how good a sound quality is." That's a downstream application (see related entries) — its primary role is as a classification cue: the main thing the ear uses to sort sounds into different source categories.
Why it happens
A vibrating object rarely produces a single frequency; it vibrates simultaneously at the fundamental and a series of integer multiples of it (harmonics), and the relative strength of each harmonic is shaped by the material, geometry, and excitation method of the vibrating body — a wooden resonant box and a metal reed will boost or suppress different harmonics in entirely different proportions. This distribution of harmonic strengths is effectively a spectral fingerprint of the source: the fundamental sets pitch, the fingerprint carries source identity, and the two are orthogonal, non-interfering channels of information.
The auditory system is highly sensitive to the shape of this spectral envelope, and can separate sources by harmonic-strength distribution alone even when pitch and loudness are both held fixed. On top of the steady-state envelope, the rate at which harmonics change during the sound's onset (attack) typically carries more identity information than the steady-state portion — real instruments and voices show fast harmonic evolution at onset, which is why stripping the attack and looping only the steady-state harmonics makes timbre identification noticeably harder.
Studying it
A common approach synthesizes a set of tones matched in fundamental and overall energy but varying only in harmonic amplitude distribution, has listeners make similarity judgments or classification responses, then projects the subjective similarity data onto a small number of continuous dimensions using multidimensional scaling, checking how those dimensions relate to physical measures like spectral centroid or harmonic decay rate — a method generally called timbre space modeling.
Common independent variables: relative amplitude of each harmonic, spectral centroid (the center of mass of the energy distribution), rate of harmonic change during onset. Common dependent variables: similarity ratings, classification accuracy, coordinates within the derived timbre space.
Where it stops holding
- Steady-state harmonic structure is not the whole story: cut the attack off a real instrument recording and loop only the steady-state harmonics, and identification accuracy drops noticeably — the onset transient alone carries a nontrivial share of identity information, so analyzing steady-state spectral envelope alone underestimates real-world identification difficulty.
- Pure tones (fundamental only, no harmonics) have no timbre difference to speak of; this mechanism only applies to complex tones that actually contain harmonic structure.
- Timbre-space dimensions depend on the specific set of sounds used in testing; swap in a set of sources that are more or less similar to each other and the number and meaning of the extracted dimensions will shift — there is no universal, one-size-fits-all list of "timbre dimensions."
Applying it
- When designing a family of sounds meant to read as "the same kind of event" (e.g., every confirmation action within one app), keep the overall harmonic-structure profile consistent even while varying pitch, to preserve category recognizability; conversely, to make two event types read as coming from different sources, changing harmonic structure is more reliable than changing loudness alone.
- Don't trim or compress only the steady-state portion of a sound to save resources — the onset's harmonic evolution is a substantial part of identity recognition, and cutting it makes a sound hard to categorize even if the steady-state spectral envelope is preserved intact.
- Verification: produce several candidate sound effects that share fundamental and loudness but differ in harmonic structure, and have listeners classify or rank them for similarity under blind conditions — confirm that sources meant to sound "the same" or "different" are actually sorted that way by listeners, rather than assuming a timbre change is perceptible just because it was intended.
Related
- Same group: A3.12.2 Lossy compression that discards harmonic detail changes timbre, not just loudness · A3.12.3 Realistic sound effects imply physical material by matching its harmonic signature · A3.12.4 Voice quality shapes listeners' subjective judgment of content credibility
- Nearby: A3.11 Pitch discrimination · A3.07 Distinguishability of alarm sounds
- Search terms:
timbre·spectral envelope·harmonic structure·multidimensional scaling of timbre