SeekUI: Predicting Visual Search Behavior on Graphical User Interfaces with a Reward-Augmented Vision Language Model

Eye Tracking & Gaze InteractionExplainable AI (XAI)User Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingUI/UX DesignersSoftware Engineers & DevelopersHCI Researchers

Paper Title

SeekUI: Predicting Visual Search Behavior on Graphical User Interfaces with a Reward-Augmented Vision Language Model

Publication Info

  • Topic area: Visual search behavior modeling on graphical user interfaces (GUIs).
  • Keywords: Visual search, scanpath prediction, graphical user interfaces, vision language models, reinforcement learning, human-computer interaction, usability testing, multimodal modeling, layout optimization, eye-tracking.

Background and Problem

  • Problem / challenge: Existing models for visual search, particularly those trained on natural images or free-viewing tasks, fail to accurately predict human scanpaths on GUIs due to their inability to account for GUI-specific structures, semantics, and task-driven behaviors.
  • Significance: Accurate prediction of visual search behavior on GUIs is crucial for improving usability, reducing user frustration, and supporting interface design and optimization without requiring extensive user testing.
  • Motivation and related work: Prior research has identified key human search behaviors on GUIs, such as the Guess–Scan–Confirm strategy, but no computational model has successfully reproduced these patterns across diverse interface types. Existing models lack the ability to integrate GUI layout understanding and task-driven cues into scanpath predictions.

Solution

  • Proposed approach: SeekUI, a reward-augmented Vision Language Model (VLM) designed to predict human-like scanpaths on GUIs by integrating visual and textual cues and leveraging reinforcement learning (RL) for sequence-level optimization.
  • Novelty:
    1. Introduction of explanation modeling to predict not only "where" users look but also "what and why" drives their gaze shifts.
    2. Use of a two-stage training process: instruction tuning for explanation–scanpath alignment and RL fine-tuning to improve scanpath coherence.
    3. Demonstration of human-like behavior reproduction, including the Guess–Scan–Confirm strategy and sensitivity to visual clutter.
    4. Practical applications for GUI evaluation and layout optimization without requiring eye-tracking data.
  • Procedure and key techniques:
    1. Instruction Tuning: Fine-tune a VLM to generate explanations and scanpaths from GUI screenshots and textual target cues, using a curated dataset augmented with synthesized explanations.
    2. Reinforcement Learning: Optimize scanpath predictions using a ScanMatch-based reward function to enhance spatial and temporal alignment with human behavior.
    3. Evaluation: Compare SeekUI against baseline models using metrics such as ScanMatch, MultiMatch, and success rates, and assess its ability to replicate human search characteristics.

Results

  • Concrete findings:
    • SeekUI achieves a 62% success rate in locating targets, approaching the human benchmark of 70% and outperforming baselines (≤18%).
    • Significant improvements in metrics such as ScanMatch (0.347 vs. 0.236 for the best baseline) and NSS (1.394 vs. 0.810).
    • Accurate reproduction of human search patterns, including saccade direction biases, heavy-tailed saccade lengths, and clutter-driven search times.
  • Advantage over baselines:
    • Outperforms state-of-the-art models for free-viewing and natural-image search across all evaluated metrics and GUI types (mobile, desktop, web).
    • Better alignment with human scanpaths and higher success rates in task-driven search scenarios.
  • Experiments / evaluation:
    • Dataset: VSGUI10K, with 1,616 human scanpaths across 730 GUI screenshots.
    • Metrics: ScanMatch, MultiMatch, SED, NSS, AUC, and others to assess spatial, temporal, and distributional alignment.
    • Ablation studies confirm the importance of explanation modeling and RL for performance.
  • Limitations and future work:
    • Struggles with target ambiguity and highly cluttered GUIs.
    • Limited to text-based target cues; future work could incorporate multimodal signals (e.g., icons, colors).
    • Does not model dynamic interactions like scrolling or multi-page navigation.
    • Lacks explicit cognitive mechanisms, relying instead on statistical learning.

Summary

SeekUI is a reward-augmented Vision Language Model designed to predict human-like visual search behavior on GUIs. By integrating visual and textual cues and employing explanation modeling and reinforcement learning, SeekUI achieves state-of-the-art performance in scanpath prediction, surpassing baselines in accuracy and human-likeness. It reproduces key behavioral patterns, such as the Guess–Scan–Confirm strategy, and supports practical applications in GUI evaluation and layout optimization. Future work could address target ambiguity, dynamic interactions, and richer multimodal cues to further enhance its applicability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222994/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791178
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Explainable AI (XAI), User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
UI/UX Designers, Software Engineers & Developers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
7 related papers