Beyond Words: Measuring User Experience through Speech Analysis in Voice User Interfaces
Honorable MentionAuthors
Paper Title
Beyond Words: Measuring User Experience through Speech Analysis in Voice User Interfaces
Publication Info
- Topic area: User experience evaluation in voice user interfaces using speech analysis.
- Keywords: Voice user interfaces, user experience, speech analysis, paralinguistic cues, machine learning, acoustic features, real-time UX sensing, adaptive systems, interaction quality, conversational agents.
Background and Problem
- Problem / challenge: Existing methods for evaluating voice user interfaces (VUIs) rely heavily on post-hoc questionnaires and task performance metrics, which fail to capture in-the-moment dynamics and are prone to biases such as recency effects and social desirability.
- Significance: Real-time, non-intrusive evaluation of user experience (UX) is critical for improving VUI design and responsiveness, enabling systems to adapt dynamically to user needs.
- Motivation and related work: Prior research has shown that speech carries paralinguistic cues indicative of affect and cognitive effort, but few studies have systematically linked these features to UX during real-time VUI interactions. Existing studies often focus on retrospective measures or controlled setups, leaving a gap in understanding how speech reflects UX during naturalistic, task-oriented interactions.
Solution
- Proposed approach: A speech-based framework for real-time UX assessment in VUIs, leveraging acoustic, prosodic, and linguistic features to infer interaction quality.
- Novelty:
- Development of a within-subjects testbed with three distinct VUI personas and scenarios to systematically elicit diverse UX.
- Introduction of a turn-level speech-feature pipeline for real-time and offline UX sensing.
- Empirical evidence linking specific speech features to UX dimensions and demonstrating their use in machine learning models for UX classification.
- Procedure and key techniques:
- Conducted a 3 (persona) × 3 (scenario) within-subjects study with 49 participants.
- Extracted acoustic, prosodic, and linguistic features from user speech during interactions.
- Collected standardized UX ratings (e.g., attractiveness, trust, satisfaction) and analyzed correlations with speech features.
- Trained machine learning models (e.g., SVM, 1D-CNN) to classify UX levels (positive, neutral, negative) based on speech features.
Results
- Concrete findings:
- Speech features such as spectral centroid, jitter, and engagement rate varied systematically across UX levels.
- Negative UX was associated with brighter, noisier spectra, greater vocal instability, and reduced engagement.
- Machine learning models achieved up to 76.47% accuracy in classifying UX levels using speech features.
- Advantage over baselines:
- Demonstrated that speech features can reliably infer UX in real-time, outperforming traditional post-hoc methods in capturing interaction dynamics.
- Provided interpretable acoustic markers of UX, enabling privacy-preserving, non-lexical analysis.
- Experiments / evaluation:
- Conducted on a custom web-based platform with three VUI personas varying in response latency, error handling, and voice quality.
- Evaluated using a combination of subjective UX ratings, acoustic analysis, and machine learning classification.
- Limitations and future work:
- Heterogeneity in recording conditions and limited demographic diversity may affect generalizability.
- Need for controlled experiments to isolate the effects of specific design factors.
- Future work should explore advanced deep learning models, cross-cultural validation, and multimodal approaches integrating acoustic, linguistic, and interactional features.
Summary
This study demonstrates that speech features can serve as reliable, real-time indicators of user experience (UX) in voice user interfaces (VUIs). By analyzing acoustic, prosodic, and linguistic characteristics, the authors identified systematic patterns linking vocal behavior to UX dimensions such as trust, satisfaction, and attractiveness. Machine learning models achieved high accuracy in classifying UX levels, highlighting the potential for adaptive VUIs that respond dynamically to user state. These findings pave the way for privacy-preserving, in-situ UX assessment and inform the design of more responsive and user-centered conversational agents. Future research should address generalizability, explore advanced modeling techniques, and integrate multimodal signals for enhanced UX prediction.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)