Emotion Recognition in Conversations using Brain and Physiological Signals
Authors
Title of the Paper
Emotion Recognition in Conversations Using Brain and Physiological Signals
Paper Information
- Research Area: Human-Computer Interaction, Emotion Recognition
- Keywords: Human-Computer Interaction, Multimodal Emotion Recognition, Conversational Emotion Recognition, Electroencephalogram (EEG), Galvanic Skin Response (GSR), Photoplethysmography (PPG), Multimodal Fusion
Research Background and Problem
-
Problem or Challenge:
- Emotion recognition is crucial for interpersonal interactions and human-computer interaction. However, traditional methods based on facial expressions, voice, and body movements have limitations, such as the inability to detect hidden or masked emotions.
- Most studies rely on passive tasks to induce emotions (e.g., watching videos or images), which fail to encompass complex emotional states such as those occurring in everyday conversations.
-
Significance of the Research:
- Applying emotion recognition to fields such as healthcare, education, and social interaction can enhance user experience.
- Brain activity and physiological signals have demonstrated potential in emotion recognition, showing uncontrollability and high reliability, which are significant for developing emotion-aware interactive systems.
-
Motivation and Related Work:
- There is a lack of research on real-time conversational emotion recognition using EEG and physiological signals in non-acted scenarios.
- The authors aim to create a passive conversational environment to collect emotional response data, analyze it, and explore the possibilities and limitations of this approach.
Solution
-
Method or Solution:
- A novel experimental setup is proposed, inducing emotions through face-to-face dynamic conversations and collecting multimodal data, including EEG, GSR, and PPG signals.
- Two emotion classification strategies are designed and tested: Random Forest Classifier (RFC) and Long Short-Term Memory network (LSTM).
- A multimodal data fusion strategy is implemented, using weighted decision fusion to enhance emotion recognition performance.
-
Innovations:
- An attempt to recognize emotions during conversations in non-acted scenarios, simulating emotional responses in real-life situations.
- Creation of a multimodal dataset, PEGCONV, including EEG, GSR, and PPG data.
- Exploration of the performance differences between single-modal and fused multimodal approaches, with suggestions for future improvements.
-
Implementation Steps and Key Techniques:
- Experimental Design: Establish a face-to-face dynamic conversational environment, guided by experienced psychologists to induce five target emotions (happiness, sadness, anger, fear, neutral).
- Data Collection: Record brain activity using OpenBCI EEG devices, collect GSR and PPG data using Shimmer GSR+ modules, and simultaneously record audio and video.
- Data Preprocessing: Filter and standardize EEG, GSR, and PPG signals to remove noise and motion artifacts.
- Feature Extraction: Extract features from EEG band power, GSR, and PPG signals using statistical methods and spectral analysis.
- Classification and Fusion: Classify single-modal data using Random Forest and LSTM classifiers, and enhance accuracy through weighted decision fusion.
Research Findings
-
Specific Findings:
- During the induction of five target emotions (happiness, sadness, anger, fear, neutral), the experiment confirmed that most participants could experience the target emotions, validating the effectiveness of the emotion induction strategy.
- The LSTM classifier achieved approximately 68% accuracy for arousal (high/low) and 64% for emotional valence (positive/negative) on independent participant data, significantly outperforming the Random Forest classifier.
-
Advantages Compared to Existing Solutions:
- Emotion recognition in non-acted scenarios significantly enhances the realism of the research context, making it more applicable to everyday life compared to traditional methods that induce emotions using videos or images.
- Multimodal fusion improves the robustness of emotion recognition, with specific emotion classification performance surpassing single-modal approaches.
-
Experimental or Evaluation Results:
- In emotion evaluation based on individual data, the average F1 scores for arousal level and emotional valence recognition reached 77.6% and 80%, respectively.
- Using deep learning (LSTM), acceptable emotion recognition performance was achieved even for independent participants.
-
Limitations and Future Directions:
- Sensor Limitations: The use of gel-based EEG devices imposes time constraints that affect signal quality, highlighting the need for more comfortable hardware.
- Emotion Labeling Methods: The current use of post-experiment self-reported emotion labels may lack precision, necessitating the exploration of real-time, multimodal labeling methods.
- Personality and Emotional Diversity: Variations in personality traits and emotional responses lead to data diversity, requiring further exploration to balance such influences.
- Optimization of Fusion Strategies: The full potential of multimodal fusion has yet to be realized, and more sophisticated deep learning methods are needed to improve performance.
- Data Collection Scale: The current sample size is small, necessitating an expansion of the dataset to support training of higher-performance models.
The authors plan to further integrate video and audio data and apply the findings to human-computer interaction systems (e.g., digital assistants and emotional robots) to enhance the emotional interaction capabilities of intelligent interfaces.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can EEG and physiological signals (e.g., GSR and PPG) be combined to recognize hidden or complex emotions in human conversation?Category: Biometric Authentication and Secure IdentificationSimilar questionsarrow_forward
- Can multimodal data fusion improve the accuracy of real-time conversational emotion recognition?Category: Biometric Authentication and Secure IdentificationSimilar questionsarrow_forward
- In non-performative settings, how can experiments be designed to effectively elicit and collect emotional responses in authentic conversation?Category: Biometric Authentication and Secure IdentificationSimilar questionsarrow_forward
Practical Problems
1- Hidden or complex emotions are difficult to capture in user communication, limiting affective interaction in intelligent systems.Category: Biometric Authentication and Secure IdentificationSimilar questionsarrow_forward
- 100%
Workload Alerts - Using Physiological Measures of Mental Workload to Provide Feedback during Tasks
CHI '18· Brain-Computer Interface (BCI) & Neurofeedback +1
- 100%
Cyberoception: Finding A Painlessly-Measurable New Sense In The Cyberworld Towards Emotion-awareness In Computing
CHI '25· Brain-Computer Interface (BCI) & Neurofeedback +1
- 100%
TiiS: A Classification Model for Sensing Human Trust in Machines Using EEG and GSR
IUI '19· Brain-Computer Interface (BCI) & Neurofeedback +1
Based on Jaccard similarity of research subtopics & professions (≥60%)