PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences

Mixed Reality WorkspacesContext-Aware ComputingUI/UX Designers

Paper Title

PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences

Paper Information

  • Domain: Mixed Reality (MR), Document Processing, and User Experience Design
  • Keywords: Mixed Reality, Instruction Documents, Tutorials and Help Systems, Context-Aware MR, Interface Optimization, Task Segmentation

Research Background and Problem Statement

  • Challenges:

    • Current paper-based instruction documents are disconnected from users' physical environments, making the process of referencing documents while performing tasks in real-world scenarios more complex.
    • Users often experience frequent context switching when reading instructions scattered across documents, which increases cognitive load during task learning and execution.
    • Persistently placing virtual instructional content may obstruct critical interaction areas, thereby reducing user experience.
  • Significance:

    • Paper-based instruction documents are vital knowledge-sharing mediums widely used in device operation, facility maintenance, and educational contexts. When tasks and documents are not directly linked, users may need to repeatedly shift their attention, negatively impacting work efficiency and task outcomes.
  • Motivation and Related Work:

    • While existing MR systems have demonstrated the potential of MR technology in spatial task guidance, such as enhancing document experiences via AR, these methods often require lengthy manual annotation and setup processes and lack adaptability to dynamic environments.

Solution

  • Proposed Approach:

    • Introduce PaperToPlace, a system comprising two core workflows: a creation pipeline for rapidly converting paper-based instruction documents into MR experiences, and a consumption pipeline for optimizing the display of instructional content based on user context.
    • The creation pipeline uses OCR and pre-trained language models to parse instruction documents into step-by-step text, allowing users to further edit and associate key objects.
    • The consumption pipeline employs optimization algorithms to determine the best layout for instructional text, considering user gaze, gesture tracking, and critical interaction areas.
  • Innovations:

    • The creation phase leverages machine learning models to automatically extract associations between instructions and environmental anchors, significantly improving document processing efficiency.
    • During the consumption phase, the system uses probabilistic optimization methods to dynamically adjust the placement of instructional content while providing accessible interaction options (e.g., gesture control, automatic layout recommendations).
    • Supports step-by-step interaction and interruption recovery design, forming a modular information delivery format that optimizes the learning curve.
  • Implementation Steps:

    • Creation Pipeline: Users first capture images of paper documents, and the system extracts text using OCR. The model then assists users in associating each step of the instructions with physical objects in the target scene (e.g., a microwave).
    • Consumption Pipeline: After wearing MR devices, users can view each step of the instructions through dynamically optimized layout positions and gesture interactions. The system uses an environmental model (semantically annotated spatial anchors) and real-time user state to optimize the placement of instructions.

Research Outcomes

  • Key Results:

    • Both the creation and consumption pipelines demonstrated strong applicability and user experience in real-world scenarios.
    • PaperToPlace successfully provides spatialized MR instructions without obstructing interaction areas and effectively reduces the frequency of user context switching, leading to more efficient task execution.
  • Advantages Over Existing Solutions:

    • More efficient creation process: Compared to traditional MR content creation, machine learning-assisted methods reduce the time and cognitive load required to generate spatialized instruction documents.
    • More natural consumption experience: Instruction placement is dynamically optimized, eliminating the need for manual adjustments and reducing hand and visual strain.
    • Supports tasks in complex environments: Applicable not only to static environments but also capable of dynamically adjusting virtual object layouts based on real-time user behavior data.
  • Experimental or Evaluation Results:

    • Creation Pipeline: Both modes (manual and ML-supported) exhibited high system usability (SUS scores >70), with ML assistance significantly reducing task completion time (TCT decreased by approximately 30% on average).
    • Consumption Pipeline: Compared to baseline models, PaperToPlace successfully reduced user context switching time, head movement paths, and angle changes.
    • Feedback: Most users reported that PaperToPlace significantly reduced reading barriers and information retrieval costs while improving task completion convenience and efficiency.
  • Limitations and Future Directions:

    • Limitations:

      • Currently focused on static key objects, lacking support for dynamic moving objects (e.g., handheld devices).
      • Does not fully support instruction documents with complex metadata (e.g., warning information and multi-step associations).
      • The system relies on high-quality spatial anchor data, while existing MR devices have limited environmental modeling accuracy.
    • Future Directions:

      • Enhance the ability to parse documents with complex formats and dynamic content.
      • Explore cross-domain applicability, extending to more types of instruction documents (e.g., industrial maintenance tasks, medical operation guides).
      • Optimize user interaction by reducing manual input and enhancing full-process automation.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/126790/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3586183.3606832
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Mixed Reality Workspaces, Context-Aware Computing
work
Professions
UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers