Feasibility and Utility of Multimodal Micro Ecological Momentary Assessment on a Smartwatch

Smartwatches & Fitness BandsBiosensors & Physiological MonitoringContext-Aware ComputingUniversity Professors & ResearchersHCI Researchers

Research Background and Problems

  • What problems or challenges did the authors identify?
    Current human activity recognition (HAR) systems require high-quality labeled data to train and validate models. Traditional data collection methods (e.g., studies conducted in laboratory environments) have limitations, such as failing to fully reflect the complex and diverse behaviors of people in real-life scenarios. Additionally, using video recordings captured via wearable cameras for manual annotation not only increases privacy risks but is also highly time- and resource-intensive.
    Micro-ecological Momentary Assessment (𝜇EMA), while capable of delivering quick surveys via smartwatches, may suffer from response bias due to the reliance on a single input modality (e.g., touch or voice), which can be influenced by environmental contexts.

  • Why is this problem important?
    High-quality real-time behavior labels can not only optimize the performance of HAR models but also reduce errors in real-world applications, enhancing their potential for broader use cases such as health monitoring or real-time intervention systems.

  • Research Motivation and Related Work
    The authors propose a novel multimodal 𝜇EMA system that combines touch and voice inputs to overcome the contextual bias of traditional unimodal methods and maximize user data collection. Through a review of related literature, the authors found that multimodal interaction can improve user experience and generate more high-quality data by increasing response rates.

Solution

  • What methods or solutions did the authors propose?
    The authors designed and implemented a multimodal 𝜇EMA data collection system that allows participants to report real-time posture and activity data via touch or voice input on a smartwatch.

  • What is innovative about this solution?

    1. Offers both touch and voice input modes, increasing user flexibility.
    2. Enables high-frequency data collection via smartwatches (every 5 minutes), generating time-dense labels.
    3. Combines heart rate, wrist motion data, and historical feedback to optimize touch options, improving labeling efficiency.
  • What are the implementation steps and key technologies used?

    1. Prompt Design: Uses haptic vibration to prompt users to report, avoiding sound-based notifications to minimize environmental disturbance. The screen displays four activity suggestions based on heart rate data and historical reports.
    2. Multimodal Interaction: Users can select activities via touch or report behavior labels through voice input. Audio recordings from voice input are transcribed into text in real time.
    3. System Implementation: The system is built on the Android Wear platform, leveraging automatic speech recognition (ASR) models and open large language models (LLMs) for real-time label extraction.

Research Findings

  • What specific results were achieved?

    1. In a seven-day real-world experiment, the system achieved a high response rate of 72.4%, with participants responding to an average of 84 prompts per day.
    2. Data showed that participants exhibited different preferences for touch and voice input modes depending on the context, such as favoring voice input at home or during high activity levels.
    3. The automatic speech recognition model achieved an 85.9% transcription accuracy for user voice input, significantly outperforming related studies using unoptimized models.
  • What advantages does it have compared to existing solutions?
    This study demonstrated the advantages of multimodal interaction, addressing the environmental bias issues associated with single-input methods while collecting richer behavioral labels (e.g., combining activity and posture). Additionally, the system shows potential for real-time label extraction.

  • What were the experimental or evaluation results?

    1. Quantitative analysis revealed that factors such as heart rate, wrist activity levels, and environmental audio significantly influenced response rates and input mode preferences.
    2. Voice input was found to be more effective in quiet environments or at participants' homes, while touch input was preferred in busier or public settings.
    3. The data quality and label diversity were confirmed to meet the requirements of high-accuracy HAR models, with LLMs achieving 65%-78% accuracy in automatic label mapping tasks.
  • Limitations and Future Directions

    1. Limitations:
      • The study sample was relatively young and tech-savvy, which may limit the generalizability of the findings.
      • The short battery life of smartwatches and reliance on network connectivity constrained the system's continuous functionality.
    2. Future Directions:
      • Explore longer-term or longitudinal studies to evaluate the system's adaptability and changes in participant behavior over time.
      • Add system feedback features to help reduce users' cognitive load and improve label quality.
      • Investigate ways to further optimize LLM and ASR models to achieve truly real-time label extraction.

Conclusion

This study proposed an innovative multimodal 𝜇EMA system and validated its feasibility and advantages in high-frequency data collection. By integrating touch and voice input modalities, the solution effectively reduced response bias associated with unimodal input methods, providing robust data support for real-time human activity recognition systems. However, future work is needed to address hardware limitations and user diversity issues in real-world deployments, as well as to explore ways to enhance the system's adaptability to user needs and feedback.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189651/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714086
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Smartwatches & Fitness Bands, Biosensors & Physiological Monitoring, Context-Aware Computing
work
Professions
University Professors & Researchers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers