Signals of Aggression: Modelling Multimodal Cues and Perceptual Effects in Virtual Agents

Affective Human-Computer DialogueSocial Robot InteractionEmpathy & Emotional DesignPhysicians, Nurses & CliniciansPsychiatrists & PsychotherapistsHCI Researchers

Paper Title

Signals of Aggression: Modelling Multimodal Cues and Perceptual Effects in Virtual Agents

Publication Info

  • Topic area: Multimodal aggression modeling in Intelligent Virtual Agents (IVAs) for training and simulation.
  • Keywords: Intelligent Virtual Agents, multimodal aggression, emotion modeling, virtual reality, customer service training, facial expressions, body movement, voice synthesis, aggression perception, user studies.

Background and Problem

  • Problem / challenge: Prior work on Intelligent Virtual Agents (IVAs) has focused on basic emotions and unimodal cues, leaving a gap in understanding how aggression can be modeled multimodally and scaled systematically by intensity.
  • Significance: Modeling aggression in IVAs is critical for training frontline staff in customer service and conflict de-escalation, as it prepares them for real-world hostile interactions.
  • Motivation and related work: Existing IVA-based training systems prioritize realism but lack clarity on how aggression is modeled across modalities or how cues interact. Previous studies have not comprehensively examined the perception of aggression in IVAs across language, voice, and body cues.

Solution

  • Proposed approach: A psychologically grounded multimodal aggression model for IVAs that parametrizes language, voice, body movement, and facial expressions across four aggression levels.
  • Novelty:
    1. Development of a multimodal aggression model validated across language, voice, and body cues.
    2. Two-stage experimental evaluation to assess unimodal and multimodal aggression perception.
    3. Design guidelines for creating IVAs that convey graded aggression effectively.
  • Procedure and key techniques:
    1. Language aggression modeled using message types (e.g., teasing, threats, character attacks) based on communication taxonomies.
    2. Voice aggression modulated using loudness and spectral roughness via Azure Text-to-Speech.
    3. Body aggression designed using Laban Movement Analysis (LMA) and Facial Action Coding System (FACS) to adjust body movement and facial expressions.
    4. Two user studies with flight attendants to evaluate unimodal and multimodal aggression perception.

Results

  • Concrete findings:
    • Participants reliably perceived aggression across modalities, with body cues being the most effective.
    • Language cues plateaued at higher aggression levels, and subtle language cues were often ambiguous.
    • Multimodal combinations produced three distinct aggression levels (Low, Mid, High), with congruent cues yielding the clearest distinctions.
  • Advantage over baselines:
    • The model provides a systematic, graded aggression framework, surpassing prior binary or unimodal approaches.
    • Body and facial cues demonstrated the strongest predictive power for perceived aggression.
  • Experiments / evaluation:
    • Experiment 1: Unimodal perception study with 19 flight attendants, showing that body and voice cues were more reliable than language for conveying aggression.
    • Experiment 2: Multimodal perception study with 19 flight attendants, revealing that aligned multimodal cues produced clear aggression levels, while incongruent cues led to perceptual averaging.
  • Limitations and future work:
    • The model is based on Western norms and may not generalize to other cultures.
    • Focuses on reactive aggression and excludes instrumental aggression.
    • Limited to non-violent aggression in customer service contexts; future work should explore high-stakes environments and adaptive IVA behaviors.

Summary

This paper introduces a multimodal aggression model for Intelligent Virtual Agents (IVAs), integrating language, voice, and body cues to simulate graded aggression. Two user studies with flight attendants validated the model, showing that participants reliably perceived aggression, particularly through body and facial cues. The findings support the use of multimodal, congruent cues for clear aggression signaling, with applications in training frontline staff for conflict de-escalation. Future work should address cultural generalizability, adaptive behaviors, and high-stakes scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222976/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791031
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Affective Human-Computer Dialogue, Social Robot Interaction, Empathy & Emotional Design
work
Professions
Physicians, Nurses & Clinicians, Psychiatrists & Psychotherapists, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers