PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences
Authors
Paper Title
PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences
Paper Information
- Domain: Mixed Reality (MR), Document Processing, and User Experience Design
- Keywords: Mixed Reality, Instruction Documents, Tutorials and Help Systems, Context-Aware MR, Interface Optimization, Task Segmentation
Research Background and Problem Statement
-
Challenges:
- Current paper-based instruction documents are disconnected from users' physical environments, making the process of referencing documents while performing tasks in real-world scenarios more complex.
- Users often experience frequent context switching when reading instructions scattered across documents, which increases cognitive load during task learning and execution.
- Persistently placing virtual instructional content may obstruct critical interaction areas, thereby reducing user experience.
-
Significance:
- Paper-based instruction documents are vital knowledge-sharing mediums widely used in device operation, facility maintenance, and educational contexts. When tasks and documents are not directly linked, users may need to repeatedly shift their attention, negatively impacting work efficiency and task outcomes.
-
Motivation and Related Work:
- While existing MR systems have demonstrated the potential of MR technology in spatial task guidance, such as enhancing document experiences via AR, these methods often require lengthy manual annotation and setup processes and lack adaptability to dynamic environments.
Solution
-
Proposed Approach:
- Introduce PaperToPlace, a system comprising two core workflows: a creation pipeline for rapidly converting paper-based instruction documents into MR experiences, and a consumption pipeline for optimizing the display of instructional content based on user context.
- The creation pipeline uses OCR and pre-trained language models to parse instruction documents into step-by-step text, allowing users to further edit and associate key objects.
- The consumption pipeline employs optimization algorithms to determine the best layout for instructional text, considering user gaze, gesture tracking, and critical interaction areas.
-
Innovations:
- The creation phase leverages machine learning models to automatically extract associations between instructions and environmental anchors, significantly improving document processing efficiency.
- During the consumption phase, the system uses probabilistic optimization methods to dynamically adjust the placement of instructional content while providing accessible interaction options (e.g., gesture control, automatic layout recommendations).
- Supports step-by-step interaction and interruption recovery design, forming a modular information delivery format that optimizes the learning curve.
-
Implementation Steps:
- Creation Pipeline: Users first capture images of paper documents, and the system extracts text using OCR. The model then assists users in associating each step of the instructions with physical objects in the target scene (e.g., a microwave).
- Consumption Pipeline: After wearing MR devices, users can view each step of the instructions through dynamically optimized layout positions and gesture interactions. The system uses an environmental model (semantically annotated spatial anchors) and real-time user state to optimize the placement of instructions.
Research Outcomes
-
Key Results:
- Both the creation and consumption pipelines demonstrated strong applicability and user experience in real-world scenarios.
- PaperToPlace successfully provides spatialized MR instructions without obstructing interaction areas and effectively reduces the frequency of user context switching, leading to more efficient task execution.
-
Advantages Over Existing Solutions:
- More efficient creation process: Compared to traditional MR content creation, machine learning-assisted methods reduce the time and cognitive load required to generate spatialized instruction documents.
- More natural consumption experience: Instruction placement is dynamically optimized, eliminating the need for manual adjustments and reducing hand and visual strain.
- Supports tasks in complex environments: Applicable not only to static environments but also capable of dynamically adjusting virtual object layouts based on real-time user behavior data.
-
Experimental or Evaluation Results:
- Creation Pipeline: Both modes (manual and ML-supported) exhibited high system usability (SUS scores >70), with ML assistance significantly reducing task completion time (TCT decreased by approximately 30% on average).
- Consumption Pipeline: Compared to baseline models, PaperToPlace successfully reduced user context switching time, head movement paths, and angle changes.
- Feedback: Most users reported that PaperToPlace significantly reduced reading barriers and information retrieval costs while improving task completion convenience and efficiency.
-
Limitations and Future Directions:
-
Limitations:
- Currently focused on static key objects, lacking support for dynamic moving objects (e.g., handheld devices).
- Does not fully support instruction documents with complex metadata (e.g., warning information and multi-step associations).
- The system relies on high-quality spatial anchor data, while existing MR devices have limited environmental modeling accuracy.
-
Future Directions:
- Enhance the ability to parse documents with complex formats and dynamic content.
- Explore cross-domain applicability, extending to more types of instruction documents (e.g., industrial maintenance tasks, medical operation guides).
- Optimize user interaction by reducing manual input and enhancing full-process automation.
-
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can paper instruction documents be rapidly converted into spatialized and context-aware mixed reality (MR) experiences?Category: Mixed Reality Instruction Guidance and Task SupportSimilar questionsarrow_forward
- In mixed reality, how can instruction text positioning be dynamically optimized to reduce users' frequent context switching?Category: Mixed Reality Instruction Guidance and Task SupportSimilar questionsarrow_forward
- Can machine learning models improve the efficiency of MR instruction content generation and enhance user experience?Category: Mixed Reality Instruction Guidance and Task SupportSimilar questionsarrow_forward
Practical Problems
1- When performing tasks, users frequently switch between documents and the physical environment, resulting in low efficiency and high error rates.Category: Mixed Reality Instruction Guidance and Task SupportSimilar questionsarrow_forward
- 75%
Causality-preserving Asynchronous Reality
CHI '22· Mixed Reality Workspaces +2
- 75%
Collaborating Across Realities: Analytical Lenses for Understanding Dyadic Collaboration in Transitional Interfaces
CHI '23· Mixed Reality Workspaces +1
- 75%
Spatial Heterogeneity in Distributed Mixed Reality Collaboration
CHI '25· Mixed Reality Workspaces +1
- 75%
MR Object Identification and Interaction: Fusing Object Situation Information from Heterogeneous Sources
UbiComp '23· Mixed Reality Workspaces +1
- 67%
Paracentral and near-peripheral visualizations: Towards attention-maintaining secondary information presentation on OHMDs during in-person social interactions
CHI '22· Mixed Reality Workspaces +1
- 67%
On the Benefits of Image-Schematic Metaphors when Designing Mixed Reality Systems
CHI '24· Mixed Reality Workspaces
- 67%
SwitchSpace: Understanding Context-Aware Peeking Between VR and Desktop Interfaces
CHI '24· Mixed Reality Workspaces +1
- 67%
VXSlate: Exploring Combination of Head Movement and Mobile Touch for Large Virtual Display Interaction
DIS '21· Mixed Reality Workspaces
- 60%
Affinity Lens: Data-Assisted Affinity Diagramming with Augmented Reality
CHI '19· Mixed Reality Workspaces +2
- 60%
Getting out of Out of Sight: Evaluation of AR Mechanisms for Awareness and Orientation Support in Occluded Multi-Room Settings
CHI '20· AR Navigation & Context Awareness +2
Based on Jaccard similarity of research subtopics & professions (≥60%)