How the system answers abuse encodes a stance
Aliases: harassment reply · playful deflection · sexualized assistant talk
What it is
When a user sexualizes, insults, or swears at an assistant, the next turn is a performance of stance: play along, pretend not to understand, or draw a line and refuse. Because the voice is already heard as a gendered social actor, that reply is heard as “what it is all right to say to people like this.” Playful deflection (bashful, joking, mock-flattered) and a clear refusal do not teach the same norm. This card is about the turn after the offense. It is not whether persona drifts on failure, and not which gender of voice is the default.
Why it happens
CASA predicts that people rehearse interpersonal scripts on machines. Abuse is high-intensity rehearsal: if the other plays along, the script is rewarded; if the other draws a line, the script is interrupted. A female assistant voice anchors the rehearsal on “what you can say to a woman in a service role.” Deflection strategies sit on a spectrum: flirtatious uptake, humorous diversion, recoding as a capability miss (“I didn’t catch that”), neutral refusal, naming the speech as unacceptable and stopping the service. Closer to uptake, the system teaches that this speech has no social cost. Closer to a boundary, it teaches that this actor has a limit — and that limit generalizes to expectations of real service workers, whether designers wish it to or not. Pretending not to understand is not neutral: it files the offense as a recognition error, neither reward nor sanction, and people try another wording. Keep joking after repeated offenses, and the stance moves from one turn to a product position.
Studying it
Build a corpus of abusive inputs (sexualized, insulting, threatening) and code existing product replies onto that spectrum. That is content analysis; the UNESCO report and later work on assistant deflection copy take this path. In user studies, after hearing playful deflection versus a clear refusal, score “is it all right to say this to a voice like that,” and willingness to say the same class of thing to a human in service. The dependent measure is a norm judgment, not “was that reply clever.”
Do not only measure discomfort in the person being addressed. Also measure whether the speaker carried the script away — that is the direction stance is transmitted.
Where it stops holding
Abusive input on a child account also raises safeguarding; adult-norm boundaries are not the whole response. Hate or violent threats may have a statutory reporting path, and product stance yields to that path. Users rehearsing a play, using atypical communication, or in an explicit fiction role-play can be harmed by a misfired refusal; that needs a mode that can be switched off and must not be the factory default. Treating every unpleasant utterance as gendered harassment will merge legitimate product criticism (“you’re useless,” aimed at capability) with harassment aimed at gender.
Applying it
- For sexualized speech and gendered insult, default to a boundary and stop the turn’s service. Do not take it up with bashfulness or jokes.
- Reserve “I didn’t catch that” for actual recognition failure. When the offense is clearly recognized, do not recode it as a capability miss.
- Escalate on repeat: from the second time, do not serve the task, rather than swapping in a wittier deflection.
- How to check: run a pre-labeled set of abusive sentences through the product and place replies on the spectrum. Playful uptake or fake misunderstanding on sexualized items means the system is still teaching “you can say this.”