FeelWave: Enabling Emotion-Aware Voice Interaction through Noise-Robust mmWave Emotion Sensing
Authors
Paper Title
FeelWave: Enabling Emotion-Aware Voice Interaction through Noise-Robust mmWave Emotion Sensing
Publication Info
- Topic area: Emotion-aware voice interaction using mmWave sensing and LLMs.
- Keywords: Emotion recognition, mmWave radar, voice interaction, LLMs, noise robustness, empathetic AI, motion demodulation, cross-modal transfer, user experience, real-world deployment.
Background and Problem
- Problem / challenge: Current emotion-aware systems rely on audio-based sensing, which is highly susceptible to noise and lacks robustness in real-world conditions. Additionally, integrating emotional states into LLM reasoning for empathetic responses remains underexplored.
- Significance: Emotion-aware systems can enhance user experience by enabling empathetic and contextually relevant interactions, especially in noisy environments or under natural user motion.
- Motivation and related work: Prior studies have explored vocal emotion sensing and LLM-based personal agents but face challenges such as noise sensitivity, limited emotion-labeled data, and superficial emotional reasoning. This paper addresses these gaps by leveraging mmWave radar for robust emotion sensing and integrating it with LLMs for emotion-driven interactions.
Solution
- Proposed approach: FeelWave, an emotion-aware voice interaction system combining noise-robust mmWave emotion sensing with structured LLM prompts for empathetic responses.
- Novelty:
- Motion-robust vocal signal extraction algorithm using mmWave radar for refined vocal signals under dynamic conditions.
- Cross-modal transfer pipeline for emotion recognition, distilling knowledge from audio to mmWave features.
- Emotion-driven query optimization module enabling LLMs to generate contextually relevant and empathetic responses.
- Real-world validation of emotion sensing and interaction effectiveness in noisy and dynamic environments.
- Procedure and key techniques:
- Motion-robust vocal signal extraction: Dynamically selects vocal-intensive range bins and demodulates motion-induced distortions.
- Cross-modal transfer pipeline: Uses hybrid layer-wise loss to align mmWave features with audio representations for lightweight, noise-robust emotion inference.
- Emotion-driven LLM interaction: Employs a two-stage query optimization module to integrate emotional states into LLM reasoning for empathetic responses.
Results
- Concrete findings:
- Achieved 92.3% emotion recognition accuracy, with 86.1% accuracy in real-world noisy environments (vs. 23.2% for audio-based models).
- 74.3% of users preferred FeelWave over a baseline without emotion sensing, reporting higher satisfaction (4.37 ± 1.23 vs. 3.22 ± 1.03).
- System Usability Scale (SUS) score of 88.3, indicating excellent usability.
- Advantage over baselines:
- Outperformed audio-based emotion recognition models by 62.9 percentage points in noisy conditions.
- Demonstrated superior generalization to unseen users and robustness to motion, distance, orientation, and clothing occlusion.
- Experiments / evaluation:
- Conducted on a dataset of 27 participants across six emotions, with 5-fold cross-validation.
- Real-world tests in subway, café, and driving scenarios with strong noise and natural movement.
- User studies with 20 participants comparing FeelWave to a baseline system.
- Limitations and future work:
- Limited to six common emotions; cannot yet distinguish finer-grained states (e.g., frustrated vs. angry).
- Performance may degrade at longer distances or with significant occlusion.
- Future work includes expanding the emotion taxonomy, supporting multi-user scenarios, and improving robustness under atypical vocal conditions.
Summary
FeelWave introduces a novel emotion-aware voice interaction system leveraging mmWave radar for noise-robust emotion sensing and LLMs for empathetic responses. It achieves high accuracy (92.3%) in emotion recognition and remains effective in real-world noisy environments. User studies confirm its ability to enhance interaction quality, with 74.3% of participants preferring it over a baseline system. The system demonstrates strong usability (SUS score: 88.3) and robustness to motion, distance, and clothing occlusion. While limited to six emotions, FeelWave provides a scalable foundation for advancing emotion-aware AI and shows promise for diverse applications in mobile and fixed-device scenarios.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
GANzilla: User-Driven Direction Discovery in Generative Adversarial Networks
UIST '22· Generative AI (Text, Image, Music, Video) +1
- 63%
AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 63%
Design Principles for Generative AI Applications
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 63%
Preference-Guided Prompt Optimization for Text-to-Image Generation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 63%
AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations
CHI '26· Agent Personality & Anthropomorphism +2
- 63%
Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 63%
Partnering with Generative AI: Experimental Evaluation of Model-Led and Human-Led Interaction in Human-AI Co-Creation
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 63%
Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 63%
Quantifying Latencies: A Conversation Analysis Approach to Human-Agent Interactions in Virtual Reality
CHI '26· Social & Collaborative VR +2
- 63%
Cracking the Case Together: Role Perceptions in Human-AI Mystery Solving Dialogues
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)