Skin-Deep Bias: How Avatar Appearances Shape Perceptions of AI Hiring
Honorable MentionAuthors
Paper Title
Skin-Deep Bias: How Avatar Appearances Shape Perceptions of AI Hiring
Publication Info
- Topic area: Human–AI interaction in hiring contexts
- Keywords: AI hiring, fairness, bias, avatars, embodied conversational agents, trust, social identity theory, multimodal measures, generative AI, eye tracking
Background and Problem
- Problem / challenge: Existing research on algorithmic fairness in AI hiring systems has focused primarily on technical bias mitigation, neglecting how visual identity cues in AI avatars influence perceptions of fairness and bias during interactions.
- Significance: Understanding these perceptions is critical for designing equitable AI systems, as identity cues in avatars can reproduce stereotypes and amplify perceived unfairness, especially in high-stakes contexts like hiring.
- Motivation and related work: Prior studies have shown that people apply social scripts to AI systems (CASA paradigm) and that fairness perceptions are shaped by group-level bias and individual outcomes. However, research has not explored real-time adaptive AI avatars in hiring contexts or the intersectional effects of race and gender identity cues.
Solution
- Proposed approach: A crowdsourcing study using a real-time generative AI-based embodied conversational agent (ECA) platform to simulate AI hiring interviews, systematically varying avatar phenotypic traits (race and sex) and measuring perceptions of trust, fairness, and bias.
- Novelty:
- Demonstrates how intersectional identity cues (race × gender) influence fairness attributions in adaptive AI interviews.
- Combines self-reports, sentiment analysis, and webcam-based eye tracking to capture explicit and implicit responses.
- Introduces a scalable platform for studying fairness-sensitive ECA interactions in real-time.
- Procedure and key techniques:
- Participants completed a simulated AI-based job interview with avatars varying in race (black/white) and sex (male/female).
- A 2×2 experimental design assigned participants to conditions of identity match/mismatch (no match, gender match, racial match, full match).
- Measures included trust, perceived fairness, perceived bias, sentiment analysis, and gaze data collected via eye tracking.
- All participants received a standardized rejection message to control for outcome favorability bias.
Results
- Concrete findings:
- Racial mismatch increased perceived ethnic bias (+0.37 on a 5-point scale).
- Partial matches (race-only or gender-only) reduced perceived fairness compared to both-match or neither-match conditions (~0.28–0.30 points on a 5-point scale).
- Black avatars were rated higher on cognitive trust (+0.36 on a 7-point scale), future interview intention (+0.37 on a 5-point scale), and comfort (+0.46 on a 5-point scale) than white avatars.
- Gaze data showed higher focal attention to mismatching avatars’ faces (+23% relative to matched).
- Advantage over baselines:
- Extends prior work on asynchronous video interviews by introducing real-time adaptive ECAs and multimodal measures.
- Highlights intersectional identity effects that challenge single-category fairness assumptions.
- Experiments / evaluation:
- Crowdsourced study with 215 participants from the UK, Germany, and the US.
- Experimental design manipulated avatar race and sex to match/mismatch participants’ self-identified identities.
- Measures included Likert-scale self-reports, sentiment analysis, and webcam-based gaze tracking.
- Limitations and future work:
- Limited to binary race (black/white) and sex (male/female) categories, excluding non-binary and multi-ethnic identities.
- Simulated hiring context lacked real employment stakes, reducing ecological validity.
- Technical issues (e.g., ASR failures, eye-tracking accuracy) affected data quality.
- Future work should expand identity representations, test in non-WEIRD contexts, and evaluate real-world hiring scenarios.
Summary
This study investigates how avatar race and sex influence perceptions of fairness, trust, and bias in AI-mediated hiring interviews. Results show that racial mismatches heighten perceived bias, while partial matches reduce perceived fairness compared to both full and no match. Trust remained high across conditions, but gaze data revealed increased attention to mismatching avatars. By combining self-reports, sentiment analysis, and eye tracking, the study provides empirical evidence and methodological tools for examining fairness in adaptive AI systems. Findings highlight the need for careful design of AI interview avatars to mitigate perceived fairness concerns and support equitable interactions in high-stakes domains.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Funding AI for Good: A Call for Meaningful Engagement
CHI '26· AI Ethics, Fairness & Accountability +2
- 86%
To Live in Their Utopia: Why Algorithmic Systems Create Absurd Outcomes
CHI '21· AI Ethics, Fairness & Accountability +2
- 86%
Participatory AI Justice in HCI: A Scoping Review
CHI '26· Participatory Design +3
- 71%
Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AI
CHI '20· AI Ethics, Fairness & Accountability +2
- 71%
Conceptualizing Algorithmic Stigmatization
CHI '23· AI Ethics, Fairness & Accountability +2
- 71%
Can AI Be a Moral Victim? The Role of Moral Patiency and Ownership Perceptions in Ethical Judgments of Using AI-Generated Content
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
“Are Compliments Bad Now?”: Comparing LLMs and Human Interpretations of Gender Microaggressions in the Workplace
CHI '26· AI Ethics, Fairness & Accountability +2
- 71%
What is Safety? Corporate Discourse, Power, and the Politics of Generative AI Safety
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 71%
Beyond Microsoft and Monsanto: Denaturing the Monoculture Metaphor in Computing
CHI '26· AI Ethics, Fairness & Accountability +2
- 67%
Understanding the Boundaries between Policymaking and HCI
CHI '19· Algorithmic Fairness & Bias +1
Based on Jaccard similarity of research subtopics & professions (≥60%)