The Personalization Paradox: Trade-offs Between Social Presence and Task Efficiency in Embodied AR Instructors
Authors
Paper Title
The Personalization Paradox: Trade-offs Between Social Presence and Task Efficiency in Embodied AR Instructors
Publication Info
- Topic area: Personalization in embodied virtual instructors for Mixed Reality (MR) environments.
- Keywords: Mixed Reality, virtual instructors, personalization, social presence, task efficiency, user-agent similarity, voice cloning, gender matching, Big Five personality, large language models.
Background and Problem
- Problem / challenge: While virtual instructors in MR environments have advanced significantly, the impact of aligning their attributes (e.g., personality, voice, gender) with individual users on engagement, social presence, and task performance remains underexplored.
- Significance: Understanding these effects is critical for designing effective virtual instructors that balance social presence and task efficiency in fields like education and training.
- Motivation and related work: Prior studies suggest that user-agent similarity enhances trust and engagement, but these findings are limited to simpler systems or non-immersive settings. This study investigates whether these benefits extend to MR environments with lifelike avatars and advanced conversational AI.
Solution
- Proposed approach: A user study comparing four virtual instructor conditions with varying degrees of personalization: fully matched (VA), gender-matched (VG), gender and voice-matched (VGV), and non-matched (VX).
- Novelty:
- Systematic evaluation of personality, voice, and gender matching in MR-based virtual instructors.
- Integration of advanced voice cloning and Big Five personality modeling into virtual instructor design.
- Identification of the "Personalization Paradox," where subjective gains do not always translate into task efficiency.
- Procedure and key techniques:
- Developed a Mixed Reality setup with a transparent screen for face-to-face interaction with virtual instructors.
- Used GPT-4o for conversational AI, ElevenLabs for voice cloning, and MetaHuman for avatar creation.
- Conducted a within-subjects experiment with 25 participants performing assembly tasks under four instructor conditions.
- Measured task performance (e.g., duration, errors) and subjective feedback (e.g., social presence, user experience).
Results
- Concrete findings:
- Fully matched instructors (VA) were overwhelmingly preferred, scoring highest in user experience and social presence.
- Gender-matched instructors (VG) led to faster task completion compared to non-matched instructors (VX).
- Voice matching (VGV) did not significantly improve outcomes over gender matching alone, possibly due to uncanny valley effects.
- Non-matched instructors (VX) performed worst in both subjective and objective measures.
- Advantage over baselines:
- VA condition significantly outperformed VX in subjective measures like hedonic quality and co-presence.
- VG condition showed modest improvements in task efficiency over VX.
- Experiments / evaluation:
- Participants (N=25) completed four assembly tasks with randomized instructor conditions.
- Objective metrics: task duration, instructor errors, clarifications, and corrections.
- Subjective metrics: User Experience Questionnaire (UEQ), Harms-Biocca Social Presence questionnaire, and instructor rankings.
- Limitations and future work:
- Conducted in a controlled lab environment with a single, simple task, limiting external validity.
- Future research should explore more complex tasks, diverse user demographics, and longitudinal interactions.
- Further investigation needed into improving voice naturalness and emotional expressiveness.
Summary
This study examines the effects of aligning virtual instructor attributes (personality, voice, gender) with individual users in MR environments. Fully matched instructors significantly enhanced social presence and user experience but did not improve task efficiency, highlighting a "Personalization Paradox." Gender matching alone modestly improved task performance, while voice matching showed limited benefits due to potential uncanny valley effects. These findings suggest that effective virtual instructor design requires coherent, multi-layered personalization strategies tailored to task context and user needs.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
FAIR: Framing AI’s Role in Programming Competitions — Understanding How LLMs Are Changing the Game in Competitive Programming
CHI '26· Human-LLM Collaboration +2
- 63%
Investigating the Effects of LLM Use on Critical Thinking Under Time Constraints: Access Timing and Time Availability
CHI '26· Human-LLM Collaboration +2
- 63%
Modelling Experts' Sampling Strategy to Balance Multiple Objectives During Scientific Explorations
HRI '24· Human-LLM Collaboration +2
- 63%
A Multimodal Investigation of Controllability and Cognitive Load in Interactive Machine Learning
IUI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)