D2.01.3Auditory vocabulary loaddesignresearch

The number of notification sounds is constrained by memory

Aliases: sound vocabulary · earcon memory · auditory overload · notification sound count

What it is

Auditory vocabulary load is the set of sound-to-meaning associations people must remember, distinguish, and retrieve at the right moment to interpret system notifications. It has no fixed maximum for every product. Usable size depends on family structure, frequency of use, task context, availability of supporting evidence, and confusability among nearby sounds. The practical limit is reached when people cannot reliably turn a just-vanished sound into actionable meaning: adding another sound then adds uncertainty rather than information.

Why it happens

Brief sounds compete with the current task for limited attention and working memory. Associations for infrequent events decay; similar sounds interfere; sounds from other applications, devices, and surroundings draw on the same recognition resource. Each additional independent mapping requires not only remembering itself but also its boundaries against existing sounds. As a library grows by function, misrecognition often rises nonlinearly in adjacent categories, rare states, and stressful contexts. Structure, repetition, and multimodal pairing reduce load but cannot make arbitrary codes infinitely memorable.

Studying it

Test a set that represents an actual use period, rather than only playing a few new designs. After a distracting task, present a notification and ask participants to identify its event, priority, and next action; repeat after delay to simulate forgetting of low-frequency features. Build a sound–meaning confusion matrix and segment it by experience, event frequency, and environment. Add sounds incrementally and find which class begins to meaningfully lower recognition of existing ones. That produces a product-specific vocabulary budget rather than a universal “maximum number of sounds.”

Where it stops holding

Fewer is not inherently better. Collapsing events that demand different responses merely to reduce sound count can hide consequential distinctions; high-stakes events need sufficiently clear differentiation and another channel. Conversely, users of a consistent daily professional tool can sustain more code than occasional users of a consumer service. The goal is not to hear every state, but to decide which states must be routed immediately by sound alone and which can use a general attention cue plus visual or inspectable evidence.

Applying it

  • Maintain a sound vocabulary inventory; review each sound for event frequency, consequence, required response, and whether another channel can confirm it.
  • Prefer rule-governed family reuse: a base sound plus a few stable variations can express class or state more efficiently than an independent clip for each feature.
  • Use direct text, speech, object state, or an openable record for rare, complex, or high-risk events; let sound primarily attract attention.
  • Run regression confusion tests before adding a sound. If it weakens recognition of an existing critical one, merge meanings, revise hierarchy, or remove the sound rather than asking users to memorize more.

Related

  • Within the group: D2.01.1 Abstract notification sounds need learning before they acquire meaning · D2.01.2 A notification sound family needs discriminable structural relationships
  • Adjacent: D2.03.2 Urgency levels should stay within discriminable range · D5.02.2 Redundant cues should reduce, not add, cognitive load
  • Search terms: auditory vocabulary load · sound vocabulary · earcon memory

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/D2.01.3