SAE: A Multimodal Sentiment Analysis Large Language Model
Authors
The effective capture of subtle emotional changes and long-term affective trends in cross-modal information, particularly within speech modes, is found to be challenging by the current multimodal emotion analysis models. To address these challenges, an end-to-end multimodal Sentiment Analysis model, designated as the Sentiment Analysis Engine (SAE), has been proposed. Video, audio, and text information are integrated by SAE, with speech being converted into a text vector through the utilization of DeepSpeech and LSTM networks. Emotional features from speech are extracted via the EmoVoiceAnalyzer (EVA) module. Visual features are extracted through the application of ResNet-50 networks for timing modeling, while sensitivity to microexpression details is enhanced by the Multi-scale Efficient Channel Spatio Attention (MECS) mechanism. For the purpose of achieving efficient multi-modal fusion, a self- attention mechanism is employed, leading to the generation of descriptive text and the execution of in-depth emotion analysis.The experimental results show that SAE ranks first in NExT-QA testing and achieves performance indicators ahead of other SOTA in Class 2-7 emotion recognition tasks and IEMOCAP tests using the CMU-MOSI dataset, with F1 scores of 91.14 and 86.5, respectively. Furthermore, in ablation experiments, the removal of the EVA module resulted in an average 5% decrease in classification accuracy, thereby confirming the critical role played by voice components in enhancing the precision of emotion analysis.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
Patterns for Representing Knowledge Graphs to Communicate Situational Knowledge of Service Robots
CHI '21· Context-Aware Computing +1
- 67%
IM Receptivity and Presentation-type Preferences among Users of a Mobile App with Automated Receptivity-status Adjustment
CHI '21· Context-Aware Computing +1
- 67%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 67%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)