Persona must stay consistent across every reply
Aliases: persona drift · stable persona · CASA consistency
What it is
The moment a voice system speaks, it is heard as a persona: formal or casual, brisk or lingering, jokey or strictly on task. That impression is not the one-line self-introduction at the open. It is stacked by every reply. A municipal tax help line that banters through FAQ and then switches to circular-prose for a late-fee calculation is not heard as “this function is more serious.” It is heard as two different interlocutors. Consistency means the persona is predictable in every reply, including lookup, refusal, reading numbers, and handing the person off.
Why it happens
People apply social rules for interlocutors to machines — the starting claim of the Computers Are Social Actors (CASA) line. The first turns form a model of the other (this voice is a brisk clerk); later turns are used to check the model. Drift is prediction error: humour drops out, address suddenly cools, rhythm jumps from short lines to a routine, and the listener reopens “who is this,” moving attention from the task to identity. Inconsistency is also heard as several backends fighting for one mouth. Trust is priced off the least predictable turn.
Persona consistency here is predictability across turns, not a scorecard that splits wording, length, and terms of address into separate parts — those are ingredients. The requirement is: once the ingredients are chosen, lookup turns, number-reading turns, and handoff turns use the same set, rather than each being authored on its own.
Studying it
CASA-style experiments move a social-psychology manipulation onto a device: two stable personas (for example restrained versus warm) plus a mid-session switch. Dependent measures: whether the other is judged to be the same actor, trust, and whether task performance drops on the switch. The point is not which persona is “better.” It is whether the switch itself has a cost.
On a product, tag replies by turn for register, humour, address, length band, and look for band-jumps inside a session. If jumps cluster on a skill type (late fees, reading a statute), those modules were written separately. Raters can hear shuffled single turns, but that understates the jolt of a mid-dialogue change of person — whole sessions still have to be heard.
Where it stops holding
Legally required scripts (rights notice, recording announcement) are not a chat register, and people expect “this is a reading.” A perceptible boundary into and out of the script is not a persona break. When several languages or brands share an engine, consistency is per product, not per model; changing product is changing person, and should change the opening, not the middle of a session. Slowing rate for access is not a new persona; it is the same persona made reachable. Flattening every skill into one cold template is consistent and unusable — consistency used to dodge design.
Applying it
- Write one page of persona bounds: tone for what this voice will and will not do, whether jokes are on, how people are addressed. Review every skill’s replies against that page; do not let each skill set its own key.
- Put refusal, number reading, handoff, and empty results on a required-review list — those turns are most often written by someone else, and drift first.
- Inside one session, do not change brand name, joke policy, or honorific band. To change, close the previous segment and open a new one.
- How to check: listen to whole sessions and mark the jumps where “this is no longer the same other.” Whichever skill type the jumps land on is the copy to rewrite, not the welcome line.
Related
- Same group: M2.05.2 Persona must not mask capability limits · M2.05.3 Too much personality raises information-density cost
- Nearby: M2.10 Persona and consistency · M4.11 Gender and stereotypes in voice assistants · M3.01 Naturalness and intelligibility of synthetic speech
- Search terms:
cross-turn persona consistency·CASA·persona drift