GazeCoT: Unleashing Social Intelligence in Multimodal LLMs With Gaze-Informed Chain-of-Thought Reasoning

Eye Tracking & Gaze InteractionHuman-LLM CollaborationAI-Assisted Decision-Making & AutomationAffective Human-Computer DialogueUniversity Professors & ResearchersHCI ResearchersSociologists & Anthropologists

Paper Title

GazeCoT: Unleashing Social Intelligence in Multimodal LLMs With Gaze-Informed Chain-of-Thought Reasoning

Publication Info

  • Topic area: Enhancing multimodal large language models (MLLMs) with gaze-based social intelligence for human-AI interaction.
  • Keywords: Artificial social intelligence, gaze estimation, multimodal LLMs, chain-of-thought reasoning, human-AI interaction, social perception, theory-of-mind reasoning, joint attention, visual prompting, trustworthiness.

Background and Problem

  • Problem / challenge: Multimodal LLMs (MLLMs) struggle to understand non-verbal social cues, particularly human gaze, due to limitations in fine-grained visual perception and spatial reasoning. This restricts their ability to engage in socially intelligent reasoning in multimodal scenarios.
  • Significance: Accurate gaze perception is crucial for understanding human attention and intent, which are foundational for applications like human-robot interaction, smart spaces, and joint attention analysis.
  • Motivation and related work: While prior research has explored integrating egocentric gaze into MLLMs for AR/VR applications, third-person gaze integration remains unexplored due to the lack of ground truth gaze data in real-world scenarios. Advances in gaze estimation models provide an opportunity to address this gap.

Solution

  • Proposed approach: GazeCoT (Gaze-Informed Chain-of-Thought), a pipeline that integrates third-person gaze information into MLLMs using gaze estimation models and hybrid visual-text prompting strategies.
  • Novelty:
    1. Introduces third-person gaze into MLLM reasoning contexts using gaze estimation models.
    2. Develops GazeLLE-v3, a state-of-the-art gaze estimator leveraging DINOv3 for improved accuracy.
    3. Implements hybrid visual-text prompting to reduce hallucination and enhance social reasoning.
    4. Provides a plug-and-play pipeline adaptable to diverse human-computer interaction (HCI) applications.
  • Procedure and key techniques:
    • GazeLLE-v3 generates gaze heatmaps using DINOv3 features.
    • Visual prompting overlays images with bounding boxes, gaze lines, and fixation points.
    • Text prompting describes the region of interest (ROI) around the fixation point.
    • Structured reasoning context reduces hallucination and organizes prompts for efficient MLLM processing.
    • Parallelized pipeline design minimizes inference delay.

Results

  • Concrete findings:
    • GazeCoT improved gaze target recognition accuracy by 25 percentage points over baseline MLLMs.
    • Achieved a 10pp gain on the Gaze-grounded Social Intelligence (GSI) benchmark, demonstrating enhanced social intelligence.
    • User study showed significant improvements in accuracy, comprehensiveness, usefulness, explainability, and trustworthiness of MLLM-generated outputs.
  • Advantage over baselines:
    • GazeCoT consistently outperformed standalone MLLMs across benchmarks and user studies.
    • Reduced hallucination and aligned MLLM reasoning with human norms for social perception and reasoning.
  • Experiments / evaluation:
    • Benchmarks: Gaze target recognition (907 image-question pairs) and GSI (320 social intelligence questions).
    • User study: Parent-child joint media engagement analysis with 10 experts rating field notes on 6 metrics.
    • Metrics: Accuracy, comprehensiveness, usefulness, explainability, trustworthiness, overall preference.
  • Limitations and future work:
    • Latency and token cost are higher than baseline MLLMs, requiring optimization for real-time applications.
    • Current reliance on face detection limits performance in certain camera angles; head detection models could address this.
    • User studies focused on a single scenario; broader evaluations in diverse HCI applications are needed.

Summary

GazeCoT is a plug-and-play pipeline that integrates third-person gaze information into MLLMs, significantly enhancing their social intelligence through gaze-informed chain-of-thought reasoning. By leveraging the state-of-the-art GazeLLE-v3 gaze estimator and hybrid visual-text prompting, GazeCoT improves gaze perception, social reasoning, and human-AI interaction. Experiments on benchmarks and user studies demonstrate substantial gains in accuracy, explainability, and trustworthiness. Future work includes optimizing latency, expanding gaze estimation capabilities, and adapting GazeCoT to diverse HCI applications while addressing ethical considerations.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222165/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790922
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Affective Human-Computer Dialogue
work
Professions
University Professors & Researchers, HCI Researchers, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers