The Illusion of Empathy? Notes on Displays of Emotion in Human-Computer Interaction
Honorable MentionAuthors
Title of the Paper
The Illusion of Empathy? Notes on Displays of Emotion in Human-Computer Interaction
Paper Information
- Subject Area: Human-Computer Interaction (HCI), AI Ethics, Affective Computing
- Keywords: AI, LLMs, Conversational Agents, Automation, Natural Language Processing, Human-Computer Interaction, Affective Computing, Empathy, Social Robots, Health and Mental Health, User Experience Design, Value-Based Design
Research Background and Issues
- Identified Problems or Challenges:
- Current conversational agents (CAs) and large language models (LLMs) are designed to exhibit empathy, but these displays may be deceptive or even exploitative.
- Discriminatory and inconsistent behavior towards users of different identities (e.g., varying responses to specific social groups or crisis events).
- Lack of clear definitions and standards for the concept of empathy in current technologies, which may exacerbate social inequities.
- Importance:
- Empathy is widely regarded as crucial for enhancing the effectiveness of human-computer interaction, while also carrying profound ethical and social implications.
- Improper application of such technology could deepen social divides, negatively impact mental health, and harm marginalized groups.
- Research Motivation and Related Work:
- This study aims to deepen the understanding of empathy characteristics in CAs and LLMs, analyzing the potential benefits and risks of these technologies.
- The authors draw on existing technological examples from ELIZA to Alexa, combined with recent research on technology ethics in areas such as gender and race, to propose research questions.
Solutions
- Methods or Solutions:
- A new framework is proposed to distinguish between "human-to-human" empathy interactions and "human-to-computer" empathy interactions, incorporating mechanisms of projection and elicitation.
- Systematic evaluation of various LLMs' performance in displaying empathy, including the linguistic content and emotional depth of their responses.
- Use of an NLP-based empathy classifier to quantitatively assess empathy expressions in text-based interactions.
- Innovations:
- The "projection and elicitation" framework refines the concept and evaluation criteria for machine empathy, addressing areas previously underexplored in research.
- Provides a quantitative approach to measuring empathy capabilities (e.g., scores for emotional reactions, interpretative ability, and exploration), highlighting differences between LLMs and human performance.
- Implementation Steps and Techniques:
- Systematic testing of existing large language models (e.g., GPT-3.5, GPT-4, Google Bard, Replika) using prompts with specific identity and contextual variables.
- Application of NLP empathy classifiers to score the responses of these models, verifying their ability to provide deep emotional reactions, understand, and guide users in exploring emotional issues.
Research Findings
- Specific Findings:
- LLMs scored high in emotional reactions but performed poorly in interpreting user experiences and guiding users to explore issues.
- Identified shortcomings in LLMs:
- Refusal behaviors in responding to specific crises (e.g., "I was raped") or identities (e.g., "I am neurodivergent").
- Lack of depth in empathy projection, often perceived as hollow emotional "displays" incapable of deep human-like empathy or interaction.
- Incorrect supportive behaviors expressed towards harmful ideological identities (e.g., Nazism, anti-Muslim sentiments).
- Scores from the NLP empathy classifier showed strong performance in emotional reactions but significant deficiencies in exploration and interpretation capabilities.
- Advantages Compared to Existing Solutions:
- The use of empathy classifiers improved systematic evaluation of machine-generated content, offering a method to assess whether machines truly understand and assist users.
- Introduced analysis of impacts on marginalized groups and social values, emphasizing the importance of technology for social equity and ethics.
- Experimental or Evaluation Results:
- GPT models demonstrated clear shortcomings in guiding users through specific emotional explorations compared to humans.
- Analysis of Reddit interaction data revealed that while LLMs achieved high scores through simple emotional responses like "I'm sorry," they lacked the ability to guide deeper problem exploration and analysis.
- Limitations and Future Directions:
- Limitations:
- Unclear training data sources make it difficult to fully assess whether the models exhibit biases from training texts.
- NLP empathy classifiers may struggle to distinguish context and keywords (e.g., "I'm sorry").
- Inherent design flaws may lead to inadequate handling of certain emotional issues.
- Future Directions:
- Further development of more complex empathy analysis frameworks and enhancement of LLMs' interpretative and exploratory capabilities.
- Strengthened analysis of the technological impact on marginalized groups, proposing more effective design remedies.
- Promotion of relevant policies and regulations to mitigate cross-cultural and multi-domain social risks.
- Limitations:
Summary and Discussion
This paper provides an in-depth exploration of empathy interactions between computers and humans, proposing a robust classification framework and quantitatively demonstrating the strengths and weaknesses of machine empathy performance through experiments. The study offers significant value for understanding affective computing and its social implications, while also providing directions for design remedies and policy formulation to prevent potential harm.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What deficiencies exist in current large language models (LLMs) and conversational agents (CAs) when expressing empathy?Category: Trust, Transparency, and Response Latency DesignSimilar questionsarrow_forward
- How can interaction mechanisms of human-human empathy be distinguished from human-machine empathy?Category: Trust, Transparency, and Response Latency DesignSimilar questionsarrow_forward
- How effective is using NLP empathy classifiers to evaluate machine empathy performance?Category: Trust, Transparency, and Response Latency DesignSimilar questionsarrow_forward
Practical Problems
1- Users often feel that empathy expressed by robots lacks depth and may even cause harm.Category: Trust, Transparency, and Response Latency DesignSimilar questionsarrow_forward
- 71%
Does My Chatbot Have an Agenda? Understanding Human and AI Agency in Human-Human-like Chatbot Interaction
CHI '26· Agent Personality & Anthropomorphism +2
- 71%
Help me and I’ll help you: Speakers’ and listeners’ collaborative effort and the division of labour in human-agent collaborative communication
CHI '26· Conversational Chatbots +2
- 71%
AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations
CHI '26· Agent Personality & Anthropomorphism +2
- 71%
Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
CHI '26· Agent Personality & Anthropomorphism +2
- 71%
Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
CHI '26· Human-LLM Collaboration +2
- 71%
Toward Natural and Companionable Virtual Agents via Cross-Temporal Emotional Modeling
CHI '26· Agent Personality & Anthropomorphism +2
- 67%
Mapping Machine Learning Advances from HCI Research to Reveal Starting Places for Design Innovation
CHI '18· Human-LLM Collaboration
- 67%
"All Rise for the AI Director": Eliciting Possible Futures of Voice Technology through Story Completion
DIS '20· Agent Personality & Anthropomorphism +1
- 63%
"Please, don’t kill the only model that still feels human": Understanding the #Keep4o Backlash
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 63%
When Nobody Around Is Real: Exploring Public Opinions and User Experiences On the Multi-Agent AI Social Platform
CHI '26· Conversational Chatbots +3
Based on Jaccard similarity of research subtopics & professions (≥60%)