UEQManager: A Non-intrusive and Real-time System for Recognizing and Managing UEQ in Multilingual Voice Assistants

Intelligent Voice Assistants (Alexa, Siri, etc.)Multilingual & Cross-Cultural Voice InteractionHuman-LLM CollaborationExplainable AI (XAI)AI/ML Researchers & EngineersUI/UX DesignersHCI Researchers

Paper Title

UEQManager: A Non-intrusive and Real-time System for Recognizing and Managing UEQ in Multilingual Voice Assistants

Publication Info

  • Topic area: Human-Computer Interaction (HCI) and adaptive voice assistant systems.
  • Keywords: UEQ, multilingual voice assistants, gaze tracking, deep learning, adaptive interfaces, real-time systems, user experience, non-intrusive sensing, SHAP analysis, proactive adaptation.

Background and Problem

  • Problem / challenge: Current methods for assessing User Experience Quality (UEQ) in voice assistants rely heavily on post-hoc questionnaires, which are retrospective, intrusive, and fail to capture real-time fluctuations in user experience.
  • Significance: Real-time UEQ recognition and management are critical for improving user satisfaction, engagement, and trust in multilingual voice assistants, especially in diverse linguistic contexts.
  • Motivation and related work: Prior research has explored UEQ dimensions and multilingual capabilities but lacks dynamic, non-intrusive systems for real-time UEQ assessment and adaptive interaction. Behavioral signals like gaze have shown promise for inferring user states, but their integration into practical systems remains underexplored.

Solution

  • Proposed approach: UEQManager—a real-time, non-intrusive system that uses webcam-based gaze tracking to predict UEQ dimensions and adapt voice assistant behavior proactively.
  • Novelty:
    1. Development of interpretable deep learning models to predict UEQ and its seven subdimensions using gaze features.
    2. Integration of gaze-based UEQ predictions into adaptive interaction strategies using large language models (LLMs).
    3. Implementation of a real-time, privacy-preserving system for continuous UEQ monitoring and management.
    4. Empirical validation through user testing, demonstrating significant improvements in UEQ metrics compared to baselines.
  • Procedure and key techniques:
    • Module 1: Gaze-based UEQ prediction using deep learning architectures (BiLSTM-A, Hybrid-P, Hybrid-A) trained on eye-tracking data.
    • Module 2: Adaptive interface design using LLMs to generate context-specific micro-interventions based on predicted UEQ deficits.
    • Module 3: System integration with real-time webcam-based gaze sensing and privacy-preserving processing.
    • User testing with 18 participants to evaluate system effectiveness.

Results

  • Concrete findings:
    • UEQManager improved overall UEQ by 27.29% compared to baseline systems.
    • Significant improvements were observed across all UEQ subdimensions, with Trustworthiness of Content showing the highest improvement (+7.22 points).
    • Predictive models achieved moderate R² values (e.g., 0.19–0.35 for UEQKPI) and consistently outperformed linear, CNN, and Transformer baselines.
    • Real-time system latency was measured at 39.2 ± 3.1 ms, meeting interactive thresholds (<50 ms).
  • Advantage over baselines:
    • UEQManager consistently outperformed both non-adaptive and random-strategy baselines in user experience metrics and behavioral speech features (e.g., speech rate, jitter).
  • Experiments / evaluation:
    • Within-subjects user testing (N=18) with three simulated VAs (baseline, random strategy, UEQManager).
    • Metrics included standardized UEQ+ questionnaires and behavioral speech features (e.g., speech rate, jitter).
    • Statistical analyses (paired t-tests, ANOVA) confirmed significant improvements in UEQ metrics and speech fluency under UEQManager conditions.
  • Limitations and future work:
    • Prediction accuracy remains moderate; future work could enhance model robustness with larger datasets.
    • Correlational evidence from interpretability analysis does not directly validate cognitive mechanisms.
    • System robustness under varied operational conditions (e.g., lighting, background noise) was not tested.
    • Ethical and privacy concerns around camera-based sensing require further safeguards and participatory design.

Summary

UEQManager is a proof-of-concept system that integrates webcam-based gaze sensing, interpretable deep learning models, and adaptive interaction strategies to recognize and manage UEQ in multilingual voice assistants. The system demonstrated significant improvements in user experience metrics, outperforming baseline approaches in both subjective and behavioral evaluations. By leveraging gaze cues for real-time UEQ prediction and proactive adaptation, UEQManager offers a practical pathway for enhancing user satisfaction, trust, and engagement in voice assistant interactions. Future work should address prediction accuracy, operational robustness, and ethical considerations to support broader deployment.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222617/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790990
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Multilingual & Cross-Cultural Voice Interaction, Human-LLM Collaboration, Explainable AI (XAI)
work
Professions
AI/ML Researchers & Engineers, UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers