OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs
Authors
Human-LLM CollaborationContext-Aware Computing
Document Title
OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs
Document Information
- Topic Area: Integration of multimodal perception and large language models, prediction of digital actions in human-computer interaction
- Keywords: digital follow-up actions, predictive interface, large language models, dataset, ubiquitous augmented reality, diary study
Research Background and Problem
- Identified Issues or Challenges:
- Existing systems typically focus on single-modal inputs (text, images, or audio) and can only provide predefined hard-coded actions.
- Lack of predictive capability for digital follow-up actions based on multimodal perception.
- Limited understanding of users' "follow-up action needs" in real-world scenarios involving multimodal information.
- Significance:
- Action prediction capabilities can help alleviate physical, cognitive, or social interaction barriers for users.
- Supporting smarter user interface design is crucial for the future realization of "ubiquitous augmented reality."
- Research Motivation and Related Work:
- Previous studies have explored mobile information needs and multimodal interaction technologies, but there has been limited focus on action requirements in multimodal user interactions.
- The reasoning capabilities of LLMs (Large Language Models) offer new possibilities for enriching multimodal human-computer interaction.
Solution
- Proposed Methods and Solutions:
- Designed and developed a novel pipeline system OmniActions, which processes multimodal sensory data using large language models to predict users' potential digital follow-up actions.
- OmniActions includes the following modules:
- Transforming multimodal information into structured text.
- Explicit reasoning based on the "Chain-of-Thoughts (CoT)" mechanism.
- Predicting target information and follow-up actions.
- Innovations:
- Proposed a comprehensive "action design space" covering seven general categories and 17 specific categories of digital follow-up actions.
- Applied the "Chain-of-Thoughts" reasoning capability of LLMs to multimodal data.
- Developed a mobile interaction prototype to collect feedback based on real user needs.
- Implementation Steps:
- Conducted a diary study to collect user interaction data with multimodal information in real-world scenarios.
- Developed the OmniActions pipeline: used image description models (e.g., InstructBLIP), object detection models (e.g., Detectron2), and audio classifiers (e.g., YAMNet) to convert multimodal information into structured text.
- Leveraged large language models (e.g., GPT-4) for context-based learning and fine-tuning.
- Designed and evaluated a mobile interaction prototype.
Research Outcomes
- Specific Results:
- Proposed a design space encompassing seven general categories—"sharing, storing, reminding, querying, digital extraction, media manipulation, and complex actions"—further refined into 17 specific categories.
- Experimental results demonstrated high performance of GPT-4's "Chain-of-Thoughts" reasoning in predicting digital follow-up actions (achieving 94.3% accuracy in predicting the top-3 general actions).
- Developed a smartphone interaction prototype capable of predicting users' target information and corresponding actions in advance.
- Advantages:
- OmniActions is not limited to single-modal inputs and can process combinations of image, audio, and other multimodal data.
- Innovations in model transparency and interpretability, leveraging CoT reasoning to make the decision-making process of predictions more apparent.
- Experiments or Evaluation Results:
- 48% of participants found the system easy to operate, and 78% expressed optimism about the system's future potential.
- Users considered the system's predictions generally accurate, and when predictions did not meet expectations, the feature offering "more" action options effectively compensated for prediction errors.
- Limitations and Future Directions:
- The action types in the dataset are biased toward smartphone ecosystems, lacking coverage for new demands that may arise on future AR platforms.
- Prediction capabilities in cross-platform and multimodal augmented reality environments require further optimization.
- User feedback indicated that excessive action options may increase cognitive load. Future work could focus on reducing complexity through hierarchical menus and optimized option presentation.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can LLMs process multimodal sensory data to predict users' digital follow-up actions?Category: Multimodal Perception and Cross-Channel UnderstandingSimilar questionsarrow_forward
- How does chain-of-thought reasoning on multimodal data improve precision in predicting digital actions?Category: Multimodal Perception and Cross-Channel UnderstandingSimilar questionsarrow_forward
- What common design categories exist for users' follow-up action needs in multimodal environments?Category: Multimodal Perception and Cross-Channel UnderstandingSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Users struggle to obtain needed digital follow-up action suggestions in multimodal interaction scenarios.Category: Multimodal Perception and Cross-Channel UnderstandingSimilar questionsarrow_forward
- 67%
Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User Modeling
CHI '25· AR Navigation & Context Awareness +2
- 67%
From Goals to Actions: Designing Context-aware LLM Chatbots for New Year's Resolutions
CUI '25· Conversational Chatbots +2
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642068
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Context-Aware Computing
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
2 related papers