OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question Answering
Authors
Research Background and Issues
-
What problems or challenges did the authors identify?
The authors pointed out that while existing AI tools can query personal memories recorded in photos, videos, and other formats through natural language, these tools are primarily designed to retrieve individual pieces of information (e.g., specific objects in a photo). They struggle to answer more complex questions requiring contextual connections. This limitation arises because such "captured memories" often lack explicit contextual information, which is typically implicitly dispersed across multiple interconnected memories. -
Why is this issue important?
Developing the ability to query complex personal memories can help users better reflect on past experiences, enabling them to make more informed decisions efficiently. Moreover, creating a system capable of synthesizing these captured memories for question answering could significantly enhance the practicality and interaction level of intelligent personal assistants. -
Research Motivation and Related Work
Existing solutions, such as Retrieval-Augmented Generation (RAG) systems, can answer questions by leveraging external databases. However, these methods rely on explicitly linked external data and are ineffective at handling dispersed contextual information in personal memories. To address this gap, the authors proposed a novel method, OmniQuery.
Solution
-
What methods or solutions did the authors propose?
The authors introduced the OmniQuery system, which integrates semantic contextual information from multimodal personal captured memories (e.g., photos and videos) to answer more complex, contextually demanding personal memory queries. The core mechanisms of OmniQuery include context-based information enhancement and the full-process design of the question-answering system. -
What are the innovative aspects of this solution?
- A novel context taxonomy was proposed, categorizing three key types of context:
- Atomic Context: Context directly accessible within a single memory instance.
- Composite Context: Events or behaviors derived from synthesizing multiple related memories.
- Semantic Knowledge: General knowledge or behavioral patterns inferred from multiple memories.
- The sliding window method was utilized to identify composite contexts based on the temporal sequence of captured memories.
- A question-answering system was integrated, supporting natural language input, multimodal memory retrieval, and chain-of-thought reasoning to generate answers.
- A novel context taxonomy was proposed, categorizing three key types of context:
-
What are the implementation steps and key technologies used?
- Memory Structuring: Processing raw memory content (e.g., generating titles, identifying geographic/time metadata) and annotating atomic contexts.
- Composite Context Identification: Automatically identifying event contexts within temporal sequences using large language models (LLMs) and the sliding window method.
- Semantic Knowledge Derivation: Analyzing extracted contexts to generate higher-level semantic information, such as behavioral patterns and habits.
- Question-Answering System: Leveraging a Retrieval-Augmented Generation (RAG) architecture to combine contextual information and retrieved memories for answering questions.
Research Outcomes
-
What specific outcomes were achieved?
- A context taxonomy for captured memories was proposed, summarizing real user queries about personal memories through a one-month diary study.
- A memory enhancement pipeline based on this taxonomy was introduced, enabling the integration of previously scattered personal memories through contextual synthesis.
- An end-to-end question-answering system was developed, outperforming traditional RAG baselines in user evaluations.
-
What advantages does it have compared to existing solutions?
OmniQuery can handle more complex, multi-step reasoning tasks. Compared to traditional RAG systems, its accuracy improved by 28.4% (71.5% vs. 43.1%), with a win-or-tie rate of 74.5% across all tests. -
What were the experimental or evaluation results?
- The user experiment included 137 complex queries, covering direct content retrieval, context filtering, and mixed queries from real-world scenarios of user-captured memories.
- The average User Perceived Accuracy (UPA) was 3.98 (out of 5), compared to 3.06 for the baseline system.
- User case studies demonstrated OmniQuery's superior performance in handling mixed queries requiring comprehensive reasoning and contextual integration.
-
Limitations and Future Directions
The authors identified several limitations:- Challenges remain in processing complex relationships, such as identifying personal interactions and social connections.
- The study is currently based on specific user data and lacks standardized public datasets for general testing.
- Precision in active querying for multi-layered nested questions (e.g., "best selfie") is subject to subjective variability.
Future work includes:
- Developing fixed benchmark datasets to further optimize system parameters.
- Expanding support for real-time multimodal input/output (e.g., supplementing queries with voice or images).
- Integrating enhanced privacy protection technologies, such as differential privacy and enabling more on-device computation.
Through its rigorous taxonomy and generative question-answering capabilities, OmniQuery demonstrates significant potential in multimodal personal memory management and complex question answering, paving the way for scalable intelligent personal assistants.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can complex question answering over multimodal personal memories be achieved?Category: Personal Multimodal Memory RetrievalSimilar questionsarrow_forward
- How can contextual information be integrated between single memory instances and multiple related memory instances?Category: Personal Multimodal Memory RetrievalSimilar questionsarrow_forward
- How can higher-order semantic knowledge be generated from personal-memory context?Category: Personal Multimodal Memory RetrievalSimilar questionsarrow_forward
Practical Problems
1- Users cannot use smart assistants to answer complex questions spanning personal memories.Category: Personal Multimodal Memory RetrievalSimilar questionsarrow_forward
- 83%
SQLucid: Grounding Natural Language Database Queries with Interactive Explanations
UIST '24· Explainable AI (XAI) +1
- 71%
Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
CHI '24· Human-LLM Collaboration +2
- 71%
Understanding and Supporting Peer Review Using AI-reframed Positive Summary
CHI '25· Human-LLM Collaboration +2
- 71%
Campus AI vs. Commercial AI: Comparing How Students and Employees Perceive their University’s LLM Chatbot vs. ChatGPT
CHI '26· Human-LLM Collaboration +2
- 71%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
- 71%
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
CHI '26· Human-LLM Collaboration +2
- 71%
DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows
CHI '26· Human-LLM Collaboration +2
- 71%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 71%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 71%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)