SAE: A Multimodal Sentiment Analysis Large Language Model

Generative AI (Text, Image, Music, Video)Context-Aware ComputingSoftware Engineers & DevelopersUI/UX DesignersAI/ML Researchers & Engineers

The effective capture of subtle emotional changes and long-term affective trends in cross-modal information, particularly within speech modes, is found to be challenging by the current multimodal emotion analysis models. To address these challenges, an end-to-end multimodal Sentiment Analysis model, designated as the Sentiment Analysis Engine (SAE), has been proposed. Video, audio, and text information are integrated by SAE, with speech being converted into a text vector through the utilization of DeepSpeech and LSTM networks. Emotional features from speech are extracted via the EmoVoiceAnalyzer (EVA) module. Visual features are extracted through the application of ResNet-50 networks for timing modeling, while sensitivity to microexpression details is enhanced by the Multi-scale Efficient Channel Spatio Attention (MECS) mechanism. For the purpose of achieving efficient multi-modal fusion, a self- attention mechanism is employed, leading to the generation of descriptive text and the execution of in-depth emotion analysis.The experimental results show that SAE ranks first in NExT-QA testing and achieves performance indicators ahead of other SOTA in Class 2-7 emotion recognition tasks and IEMOCAP tests using the CMU-MOSI dataset, with F1 scores of 91.14 and 86.5, respectively. Furthermore, in ablation experiments, the removal of the EVA module resulted in an average 5% decrease in classification accuracy, thereby confirming the critical role played by voice components in enhancing the precision of emotion analysis.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/195852/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3708359.3712106
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Context-Aware Computing
work
Professions
Software Engineers & Developers, UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Abstract only
hub
Related Papers
4 related papers