SalChartQA: Question-driven Saliency on Information Visualisations

Explainable AI (XAI)Interactive Data VisualizationVisualization Perception & CognitionSoftware Engineers & DevelopersHCI ResearchersStatisticians & Data Scientists

Title of the Paper

SalChartQA: Question-driven Saliency on Information Visualisations

Bibliographic Information

  • Domain: Information Visualization and Visual Attention Modeling
  • Keywords: Information Visualization, Eye-tracking Studies, Gaze Behavior, Visual Saliency, Deep Learning

Research Background and Problem

  • Identified Problems or Challenges:
    • Current research lacks large-scale, diverse datasets to explore the relationship between visual attention and user information needs.
    • User attention behavior in information visualization has not been thoroughly linked to specific information needs (e.g., question-driven scenarios).
    • Existing visual saliency models, particularly those tailored for information visualization, fail to accurately predict task (question)-driven saliency maps.
  • Significance of Research:
    • Analyzing visual attention behavior can enhance the clarity, memorability, and comprehensibility of information visualization design.
    • Question-driven visual saliency research contributes to applications such as explainable artificial intelligence (XAI) and task-optimized visualization.
  • Motivation and Related Work:
    • Previous studies have primarily focused on saliency modeling for natural images and attention behavior under "free-viewing" conditions.
    • Task-driven visual attention has been studied in contexts such as web browsing or gaming, but targeted methods and datasets for information visualization remain extremely limited.

Solution

  • Method or Solution:

    1. SalChartQA Dataset: Developed a large-scale, question-driven saliency dataset through crowdsourcing, comprising 3,000 visualizations, 6,000 questions, 74,340 answers, and corresponding saliency maps.
    2. Impact Analysis: Data analysis demonstrated the significant influence of questions on visual saliency.
    3. VisSalFormer Model: Proposed a Transformer-based model capable of predicting question-driven saliency maps for information visualization.
  • Innovations:

    • SalChartQA is the first large-scale, question-driven saliency dataset, significantly expanding data scale and diversity in information needs.
    • The VisSalFormer model integrates visualization and question semantics, achieving the first computational prediction of question-driven saliency.
  • Implementation Steps:

    1. Data Collection: Utilized the BubbleView interface on a crowdsourcing platform to track user click behavior, simulating visual attention and forming the SalChartQA dataset.
    2. Data Analysis: Analyzed the impact of question characteristics (e.g., type, length) on click count, saliency coverage, and attention consistency.
    3. Model Training: Trained VisSalFormer using SalChartQA to generate saliency maps by integrating visual and question features.
    4. Comparative Experiments: Compared VisSalFormer with five baseline methods and conducted ablation studies to evaluate component contributions.

Research Outcomes

  • Specific Results:

    1. Created a large-scale, question-driven dataset, SalChartQA (3,000 visualizations, 6,000 questions).
    2. The VisSalFormer model outperformed existing models across five saliency prediction metrics (e.g., NSS, CC, KL).
    3. Data analysis revealed the strong effect of information needs (questions) on visual saliency, further supporting the practical exploration of question-driven theories.
  • Comparison with Existing Solutions:

    • Compared to traditional image feature-based models (e.g., DVS, TranSalNet), VisSalFormer can generate saliency maps directly linked to user questions.
    • By incorporating question semantics, VisSalFormer demonstrated higher prediction accuracy and question-answer adaptability.
  • Experimental or Evaluation Results:

    • VisSalFormer significantly outperformed TranSalNet in NSS scores (1.782 vs. 0.794) and showed statistically significant advantages in metrics such as CC and KL (p<0.001).
    • Ablation studies indicated that the cross-modal feature fusion module and question embedding were key components for performance improvement.
  • Limitations and Future Directions:

    • Limitations:

      • The BubbleView method cannot fully capture text-reading behavior, thus excluding text saliency analysis.
      • In certain scenarios (e.g., when colors are similar), the model's predictions lack precision.
    • Future Directions:

      1. Explore the interaction between image and text saliency to develop joint models.
      2. Integrate question-driven saliency into chart question-answering systems (CQA) to enhance model performance and interpretability.
      3. Optimize information visualization design based on saliency analysis, such as improving user interaction efficiency with complex graphics.
      4. Use question accuracy as a metric for evaluating visualization quality and further investigate how design can improve the correctness of user answers.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147419/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642942
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Explainable AI (XAI), Interactive Data Visualization, Visualization Perception & Cognition
work
Professions
Software Engineers & Developers, HCI Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
6 related papers