EchoScriptor: Automatic Lifelogging Narratives via Activity-Based Audio–Language Model
Authors
Paper Title
EchoScriptor: Automatic Lifelogging Narratives via Activity-Based Audio–Language Model
Publication Info
- Topic area: Audio-based lifelogging and narrative generation
- Keywords: Lifelogging, audio-language models, human activity recognition, narrative summaries, memory support, assistive technologies, personal informatics, privacy, real-time deployment, multimodal sensing
Background and Problem
- Problem / challenge: Existing lifelogging systems often rely on isolated event labels or fragmented data, lacking the narrative coherence and contextual richness needed for effective memory support and personal informatics. Audio-based systems face challenges such as noise, privacy concerns, and limited expressiveness.
- Significance: Lifelogging has practical applications in memory rehabilitation, personal informatics, and assistive technologies, especially for individuals with memory impairments or aging-related conditions. A narrative-based, unobtrusive system could significantly enhance usability and effectiveness.
- Motivation and related work: Prior work in audio-based human activity recognition (HAR) and lifelogging has focused on classification or multimodal approaches, often requiring manual input or producing disconnected outputs. No prior system has generated coherent, audio-only natural-language narratives optimized for lifelogging, leaving a gap in user-centered, context-aware solutions.
Solution
- Proposed approach: EchoScriptor, an audio-language system that transforms raw household audio into natural-language narratives capturing human activities and acoustic contexts.
- Novelty:
- Development of EchoLLM, a large audio-language model for generating moment-level activity descriptions.
- Introduction of a narrative construction pipeline for producing coherent, human-readable summaries.
- Creation of the first large-scale dataset for home-activity audio descriptions, enabling robust training and evaluation.
- Demonstration of real-world applicability through user studies and a web-based deployment.
- Procedure and key techniques:
- Stage 1: EchoLLM processes raw audio into moment-level descriptions using an Audio Spectrogram Transformer (AST), modality adapters, and a fine-tuned LLaMA language model.
- Stage 2: A narrative construction pipeline refines these descriptions through semantic filtering, temporal aggregation, and natural-language rewriting.
- Evaluation on a synthesized dataset of 199,800 audio clips and real-world recordings, with metrics including accuracy, F1-score, and semantic similarity.
Results
- Concrete findings:
- Achieved 94.15% activity recognition and 89.25% background recognition on the synthesized dataset.
- Reached an F1 score of 0.92 for narrative summaries, outperforming the AST+LLM+GPT baseline.
- Demonstrated 88.53% activity recognition accuracy in real-world recordings.
- Advantage over baselines:
- EchoScriptor outperformed the AST+LLM+GPT baseline in activity accuracy (+7.43%), background accuracy (+5.32%), and overall semantic quality (+20.56%).
- Generated summaries were rated as clearer, more trustworthy, and more useful than baseline outputs, approaching human-written descriptions.
- Experiments / evaluation:
- Synthesized dataset: 199,800 audio clips covering 24 activities and 8 background contexts.
- Real-world validation: 15–20 minute recordings from 3 participants in natural home environments.
- User study: 20 participants evaluated summaries for 10 recordings, comparing EchoScriptor, baseline, and human-written outputs.
- Limitations and future work:
- Residual summarization errors and limited handling of overlapping events.
- Fixed 10-second audio segmentation restricts temporal continuity.
- Dataset synthesized from online sources may not fully capture real-world acoustic diversity.
- Future directions include improving robustness, expanding to multimodal sensing, and addressing privacy concerns.
Summary
EchoScriptor is an audio-language system that generates coherent, context-aware narratives from raw household audio, addressing gaps in existing lifelogging technologies. It combines the EchoLLM model for moment-level descriptions with a narrative construction pipeline for human-readable summaries. Evaluations on a large synthesized dataset and real-world recordings demonstrated high accuracy and user-rated utility, outperforming baseline systems. While limitations remain in handling complex acoustic scenarios and ensuring privacy, EchoScriptor represents a significant step toward practical lifelogging applications, with potential for memory support, personal informatics, and assistive technologies.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)