OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs

Human-LLM CollaborationContext-Aware Computing

Document Title

OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs

Document Information

  • Topic Area: Integration of multimodal perception and large language models, prediction of digital actions in human-computer interaction
  • Keywords: digital follow-up actions, predictive interface, large language models, dataset, ubiquitous augmented reality, diary study

Research Background and Problem

  • Identified Issues or Challenges:
    • Existing systems typically focus on single-modal inputs (text, images, or audio) and can only provide predefined hard-coded actions.
    • Lack of predictive capability for digital follow-up actions based on multimodal perception.
    • Limited understanding of users' "follow-up action needs" in real-world scenarios involving multimodal information.
  • Significance:
    • Action prediction capabilities can help alleviate physical, cognitive, or social interaction barriers for users.
    • Supporting smarter user interface design is crucial for the future realization of "ubiquitous augmented reality."
  • Research Motivation and Related Work:
    • Previous studies have explored mobile information needs and multimodal interaction technologies, but there has been limited focus on action requirements in multimodal user interactions.
    • The reasoning capabilities of LLMs (Large Language Models) offer new possibilities for enriching multimodal human-computer interaction.

Solution

  • Proposed Methods and Solutions:
    • Designed and developed a novel pipeline system OmniActions, which processes multimodal sensory data using large language models to predict users' potential digital follow-up actions.
    • OmniActions includes the following modules:
      1. Transforming multimodal information into structured text.
      2. Explicit reasoning based on the "Chain-of-Thoughts (CoT)" mechanism.
      3. Predicting target information and follow-up actions.
  • Innovations:
    • Proposed a comprehensive "action design space" covering seven general categories and 17 specific categories of digital follow-up actions.
    • Applied the "Chain-of-Thoughts" reasoning capability of LLMs to multimodal data.
    • Developed a mobile interaction prototype to collect feedback based on real user needs.
  • Implementation Steps:
    1. Conducted a diary study to collect user interaction data with multimodal information in real-world scenarios.
    2. Developed the OmniActions pipeline: used image description models (e.g., InstructBLIP), object detection models (e.g., Detectron2), and audio classifiers (e.g., YAMNet) to convert multimodal information into structured text.
    3. Leveraged large language models (e.g., GPT-4) for context-based learning and fine-tuning.
    4. Designed and evaluated a mobile interaction prototype.

Research Outcomes

  • Specific Results:
    • Proposed a design space encompassing seven general categories—"sharing, storing, reminding, querying, digital extraction, media manipulation, and complex actions"—further refined into 17 specific categories.
    • Experimental results demonstrated high performance of GPT-4's "Chain-of-Thoughts" reasoning in predicting digital follow-up actions (achieving 94.3% accuracy in predicting the top-3 general actions).
    • Developed a smartphone interaction prototype capable of predicting users' target information and corresponding actions in advance.
  • Advantages:
    • OmniActions is not limited to single-modal inputs and can process combinations of image, audio, and other multimodal data.
    • Innovations in model transparency and interpretability, leveraging CoT reasoning to make the decision-making process of predictions more apparent.
  • Experiments or Evaluation Results:
    • 48% of participants found the system easy to operate, and 78% expressed optimism about the system's future potential.
    • Users considered the system's predictions generally accurate, and when predictions did not meet expectations, the feature offering "more" action options effectively compensated for prediction errors.
  • Limitations and Future Directions:
    • The action types in the dataset are biased toward smartphone ecosystems, lacking coverage for new demands that may arise on future AR platforms.
    • Prediction capabilities in cross-platform and multimodal augmented reality environments require further optimization.
    • User feedback indicated that excessive action options may increase cognitive load. Future work could focus on reducing complexity through hierarchical menus and optimized option presentation.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147206/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642068
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Context-Aware Computing
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
2 related papers