M3.01.3naturalness inflates competence expectationsdesignresearch

Too much naturalness raises expectations of understanding

Aliases: over-natural TTS · competence from voice quality · human-like oversell

What it is

When synthetic speech is too human, listeners budget understanding as if a person were on the other end. A hotel-lobby smart display greets with a highly natural concierge voice — “Hi, I can help” — and guests ask which room has a tub, is quiet, and can check in early: staff talk, not three slots. If the backend is a bounded intent set, the miss is scored as “it talks this well and still doesn’t get it,” which burns trust harder than an obviously synthetic voice. What inflates is the expectation of dialogue competence, not the intelligibility score.

Why it happens

Listeners use voice quality as a proxy for social and cognitive competence — computers as social actors. High naturalness plus fillers (“sure, let me see”) issues an open-domain licence: long questions, stacked constraints, common-sense follow-ups all feel allowed. When a slot-filler cannot take them, the fault is not acoustic; it sits in the expectation channel — the other party was heard as a reasoning agent. A machine voice lowers coverage expectations and users collapse earlier to short commands; over-naturalness makes the same backend sound dumber. The inflation is sharpest on first contact, before failure samples exist to recalibrate.

Studying it

Hold the NLU fixed and swap two voices: a high-MOS neural voice versus a clearly synthetic one. The independent variable is naturalness; do not also change wording or capability copy. Dependent measures: length and openness of the user’s next turns, out-of-coverage rate, and post-failure blame (“it should have understood” versus “I used the wrong command”). Right after the first prompt, probe “what do you think it can do,” and code whether the answer is in-skill or front-desk work.

Wizard-of-Oz can change only the voice, before NLU is frozen, to see whether the voice alone lifts expectation. If the lab first hands out a card that says “this unit only books rooms,” the measure is no longer what the voice did.

Where it stops holding

A mandated celebrity voice is a marketing constraint: bound it with a first-sentence capability fence rather than pretending the backend is human. Repeat users recalibrate; the inflation is a first-visit effect. When the command set is on the wall and people are naming from it, naturalness barely moves expectation. Children may not treat “sounds human” as a competence cue the way adults do. Collapsing every line into a cold telegram will suppress expectation and also suppress persona — a different ledger.

Applying it

  • Match persona warmth to actual coverage. Three intents do not get “ask me anything.”
  • The first prompt of a high-natural voice should bound the domain: “I can look up room types and check you in,” not a full concierge.
  • On failure, avoid human-talker wording such as “I didn’t catch your meaning”; name what is still possible.
  • How to check: sample first-visit sessions and count how many first user turns would need a human concierge. A high share means the voice is overselling.

Related

  • Same group: M3.01.1 Naturalness and intelligibility are two different measures · M3.01.2 In noise, intelligibility comes first · M3.01.4 Proper names and digit strings fail first on intelligibility · M3.01.5 Prosodic error hurts comprehension more than a dull voice · M3.01.6 Long listening accumulates auditory fatigue
  • Nearby: M2.10 Persona and consistency · M2.06 Discoverability and help · M2.05 Persona
  • Search terms: naturalness inflates competence expectations · computers as social actors · TTS persona

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/M3.01.3