Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human–AI Collaboration
Authors
Paper Title
Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human–AI Collaboration
Publication Info
- Topic area: Human–AI collaboration using first-person perspective for cognitive alignment.
- Keywords: Human–AI collaboration, first-person perspective, cognitive alignment, augmented reality, joint attention, common ground, multimodal feedback, wearable AI, egocentric vision.
Background and Problem
- Problem / challenge: Current vision-based AI assistants struggle with collaborative tasks due to communication and understanding gulfs. These include inefficiencies in translating human intentions into AI commands and interpreting embodied cues.
- Significance: Addressing these gulfs can improve task performance, reduce interaction friction, and enhance trust in human–AI partnerships, especially in dynamic, context-dependent scenarios.
- Motivation and related work: Prior research on wearable AI assistants and egocentric vision has focused on one-way input processing and predefined knowledge bases, lacking mechanisms for dynamic adaptation to user-specific rules and implicit cues. This paper builds on these foundations by conceptualizing egocentric vision as a bidirectional shared perception space.
Solution
- Proposed approach: Eye2Eye framework, leveraging first-person perspective as a shared channel for cognitive alignment in human–AI collaboration.
- Novelty:
- Conceptualization of first-person perspective as a shared perception channel for achieving cognitive alignment.
- Implementation of a real-time AR-based prototype integrating multimodal signals (gaze, gestures, speech).
- Controlled user study across three task types demonstrating reduced grounding costs and enhanced collaboration trust.
- Procedure and key techniques:
- Joint Attention Coordination: Aligning human and AI focus using gaze, gestures, and AR highlights.
- Accumulated Common Ground: Maintaining dynamic memory units that evolve with user interactions.
- Reflective Situated Feedback: Delivering context-aware multimodal guidance and refining AI understanding based on user feedback.
Results
- Concrete findings:
- Eye2Eye reduced error rates by ~58% and clarification costs by ~50% compared to baseline systems.
- Improved task completion time (average reduction of 15 seconds overall) and interaction efficiency across procedural, classification, and inspection tasks.
- Enhanced user trust, fluency, and shared awareness, with significant reductions in cognitive workload (NASA-TLX scores).
- Advantage over baselines:
- Outperformed baseline systems in accuracy, usefulness, and temporal robustness during ablation studies.
- Demonstrated strong synergy between framework components, with significant degradation in performance when components were removed.
- Experiments / evaluation:
- User study with 60 participants across three tasks (coffee machine operation, book classification, circuit board troubleshooting).
- Mixed experimental design comparing Eye2Eye to baseline systems using objective metrics (error rate, interaction turns) and subjective ratings (trust, workload).
- Post-hoc pipeline evaluation with ablated variants to assess component contributions.
- Limitations and future work:
- Processing latency (~4–5 seconds) limits responsiveness in time-sensitive scenarios.
- Study sample skewed toward younger participants with high AI acceptance; broader demographic studies needed.
- Offline ablation studies decoupled from real-time user reactions; future work should explore online evaluations and disentangle interdependent components.
Summary
The Eye2Eye framework transforms first-person perspective into a shared channel for cognitive alignment in human–AI collaboration. By integrating joint attention, dynamic memory, and multimodal feedback, it reduces interaction friction, improves task performance, and fosters trust and copresence. Controlled studies and ablation experiments confirm its effectiveness and synergy, though challenges such as latency and task-specific interruptions remain. The framework holds promise for diverse applications, including industrial assembly and assistive technologies, while raising critical design and ethical considerations for wearable AI systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Overcoming Translation Delays: Towards Better Subtitle Design for Foreign Language Conversations in Extended Reality
CHI '26· Multilingual & Cross-Cultural Voice Interaction +2
- 67%
The Effect of the Vergence-Accommodation Conflict on Virtual Hand Pointing in Immersive Displays
CHI '22· AR Navigation & Context Awareness +1
- 67%
User-Aware Rendering: Merging the Strengths of Device- and User-Perspective Rendering in Handheld AR
MobileHCI '23· AR Navigation & Context Awareness +1
- 63%
Virtual Minds, Real Work: LLM-Powered Preference-Based Planning through Spatial Multi-Agent-Human Collaboration
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)