EchoScriptor: Automatic Lifelogging Narratives via Activity-Based Audio–Language Model

Health Self-TrackingBehavior Change & Reflection TechnologyAffective Human-Computer DialoguePsychiatrists & PsychotherapistsCommunity Health WorkersElderly Care Workers

Paper Title

EchoScriptor: Automatic Lifelogging Narratives via Activity-Based Audio–Language Model

Publication Info

  • Topic area: Audio-based lifelogging and narrative generation
  • Keywords: Lifelogging, audio-language models, human activity recognition, narrative summaries, memory support, assistive technologies, personal informatics, privacy, real-time deployment, multimodal sensing

Background and Problem

  • Problem / challenge: Existing lifelogging systems often rely on isolated event labels or fragmented data, lacking the narrative coherence and contextual richness needed for effective memory support and personal informatics. Audio-based systems face challenges such as noise, privacy concerns, and limited expressiveness.
  • Significance: Lifelogging has practical applications in memory rehabilitation, personal informatics, and assistive technologies, especially for individuals with memory impairments or aging-related conditions. A narrative-based, unobtrusive system could significantly enhance usability and effectiveness.
  • Motivation and related work: Prior work in audio-based human activity recognition (HAR) and lifelogging has focused on classification or multimodal approaches, often requiring manual input or producing disconnected outputs. No prior system has generated coherent, audio-only natural-language narratives optimized for lifelogging, leaving a gap in user-centered, context-aware solutions.

Solution

  • Proposed approach: EchoScriptor, an audio-language system that transforms raw household audio into natural-language narratives capturing human activities and acoustic contexts.
  • Novelty:
    1. Development of EchoLLM, a large audio-language model for generating moment-level activity descriptions.
    2. Introduction of a narrative construction pipeline for producing coherent, human-readable summaries.
    3. Creation of the first large-scale dataset for home-activity audio descriptions, enabling robust training and evaluation.
    4. Demonstration of real-world applicability through user studies and a web-based deployment.
  • Procedure and key techniques:
    1. Stage 1: EchoLLM processes raw audio into moment-level descriptions using an Audio Spectrogram Transformer (AST), modality adapters, and a fine-tuned LLaMA language model.
    2. Stage 2: A narrative construction pipeline refines these descriptions through semantic filtering, temporal aggregation, and natural-language rewriting.
    3. Evaluation on a synthesized dataset of 199,800 audio clips and real-world recordings, with metrics including accuracy, F1-score, and semantic similarity.

Results

  • Concrete findings:
    • Achieved 94.15% activity recognition and 89.25% background recognition on the synthesized dataset.
    • Reached an F1 score of 0.92 for narrative summaries, outperforming the AST+LLM+GPT baseline.
    • Demonstrated 88.53% activity recognition accuracy in real-world recordings.
  • Advantage over baselines:
    • EchoScriptor outperformed the AST+LLM+GPT baseline in activity accuracy (+7.43%), background accuracy (+5.32%), and overall semantic quality (+20.56%).
    • Generated summaries were rated as clearer, more trustworthy, and more useful than baseline outputs, approaching human-written descriptions.
  • Experiments / evaluation:
    • Synthesized dataset: 199,800 audio clips covering 24 activities and 8 background contexts.
    • Real-world validation: 15–20 minute recordings from 3 participants in natural home environments.
    • User study: 20 participants evaluated summaries for 10 recordings, comparing EchoScriptor, baseline, and human-written outputs.
  • Limitations and future work:
    • Residual summarization errors and limited handling of overlapping events.
    • Fixed 10-second audio segmentation restricts temporal continuity.
    • Dataset synthesized from online sources may not fully capture real-world acoustic diversity.
    • Future directions include improving robustness, expanding to multimodal sensing, and addressing privacy concerns.

Summary

EchoScriptor is an audio-language system that generates coherent, context-aware narratives from raw household audio, addressing gaps in existing lifelogging technologies. It combines the EchoLLM model for moment-level descriptions with a narrative construction pipeline for human-readable summaries. Evaluations on a large synthesized dataset and real-world recordings demonstrated high accuracy and user-rated utility, outperforming baseline systems. While limitations remain in handling complex acoustic scenarios and ensuring privacy, EchoScriptor represents a significant step toward practical lifelogging applications, with potential for memory support, personal informatics, and assistive technologies.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223137/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791528
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Health Self-Tracking, Behavior Change & Reflection Technology, Affective Human-Computer Dialogue
work
Professions
Psychiatrists & Psychotherapists, Community Health Workers, Elderly Care Workers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers