Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human–AI Collaboration

Human-LLM CollaborationAR Navigation & Context AwarenessImmersion & Presence ResearchAI/ML Researchers & EngineersUI/UX DesignersUniversity Professors & Researchers

Paper Title

Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human–AI Collaboration

Publication Info

  • Topic area: Human–AI collaboration using first-person perspective for cognitive alignment.
  • Keywords: Human–AI collaboration, first-person perspective, cognitive alignment, augmented reality, joint attention, common ground, multimodal feedback, wearable AI, egocentric vision.

Background and Problem

  • Problem / challenge: Current vision-based AI assistants struggle with collaborative tasks due to communication and understanding gulfs. These include inefficiencies in translating human intentions into AI commands and interpreting embodied cues.
  • Significance: Addressing these gulfs can improve task performance, reduce interaction friction, and enhance trust in human–AI partnerships, especially in dynamic, context-dependent scenarios.
  • Motivation and related work: Prior research on wearable AI assistants and egocentric vision has focused on one-way input processing and predefined knowledge bases, lacking mechanisms for dynamic adaptation to user-specific rules and implicit cues. This paper builds on these foundations by conceptualizing egocentric vision as a bidirectional shared perception space.

Solution

  • Proposed approach: Eye2Eye framework, leveraging first-person perspective as a shared channel for cognitive alignment in human–AI collaboration.
  • Novelty:
    1. Conceptualization of first-person perspective as a shared perception channel for achieving cognitive alignment.
    2. Implementation of a real-time AR-based prototype integrating multimodal signals (gaze, gestures, speech).
    3. Controlled user study across three task types demonstrating reduced grounding costs and enhanced collaboration trust.
  • Procedure and key techniques:
    • Joint Attention Coordination: Aligning human and AI focus using gaze, gestures, and AR highlights.
    • Accumulated Common Ground: Maintaining dynamic memory units that evolve with user interactions.
    • Reflective Situated Feedback: Delivering context-aware multimodal guidance and refining AI understanding based on user feedback.

Results

  • Concrete findings:
    • Eye2Eye reduced error rates by ~58% and clarification costs by ~50% compared to baseline systems.
    • Improved task completion time (average reduction of 15 seconds overall) and interaction efficiency across procedural, classification, and inspection tasks.
    • Enhanced user trust, fluency, and shared awareness, with significant reductions in cognitive workload (NASA-TLX scores).
  • Advantage over baselines:
    • Outperformed baseline systems in accuracy, usefulness, and temporal robustness during ablation studies.
    • Demonstrated strong synergy between framework components, with significant degradation in performance when components were removed.
  • Experiments / evaluation:
    • User study with 60 participants across three tasks (coffee machine operation, book classification, circuit board troubleshooting).
    • Mixed experimental design comparing Eye2Eye to baseline systems using objective metrics (error rate, interaction turns) and subjective ratings (trust, workload).
    • Post-hoc pipeline evaluation with ablated variants to assess component contributions.
  • Limitations and future work:
    • Processing latency (~4–5 seconds) limits responsiveness in time-sensitive scenarios.
    • Study sample skewed toward younger participants with high AI acceptance; broader demographic studies needed.
    • Offline ablation studies decoupled from real-time user reactions; future work should explore online evaluations and disentangle interdependent components.

Summary

The Eye2Eye framework transforms first-person perspective into a shared channel for cognitive alignment in human–AI collaboration. By integrating joint attention, dynamic memory, and multimodal feedback, it reduces interaction friction, improves task performance, and fosters trust and copresence. Controlled studies and ablation experiments confirm its effectiveness and synergy, though challenges such as latency and task-specific interruptions remain. The framework holds promise for diverse applications, including industrial assembly and assistive technologies, while raising critical design and ethical considerations for wearable AI systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222072/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791059
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Human-LLM Collaboration, AR Navigation & Context Awareness, Immersion & Presence Research
work
Professions
AI/ML Researchers & Engineers, UI/UX Designers, University Professors & Researchers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers