Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging System

Interactive Data VisualizationHuman-LLM CollaborationAI-Assisted Decision-Making & AutomationData Scientists & AnalystsAI/ML Researchers & EngineersHCI Researchers

Paper Title

Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging System

Publication Info

  • Topic area: Visualization debugging using human-AI collaboration.
  • Keywords: Visualization debugging, Vega-Lite, human-AI collaboration, large language models, Stack Overflow, mixed-initiative systems, code generation, empirical study, retrieval augmentation, hybrid feedback.

Background and Problem

  • Problem / challenge: Visualization debugging is complex due to its reliance on human visual perception, and users often struggle to resolve issues independently. Existing solutions, such as Q&A forums and Large Language Models (LLMs), have limitations in accuracy, speed, and alignment with user intent.
  • Significance: Effective debugging is essential for creating accurate and aesthetically pleasing visualizations, which are critical for data communication and decision-making.
  • Motivation and related work: Prior research has explored programmers’ help-seeking behaviors, visualization authoring, and debugging tools. However, there is limited understanding of the strengths and weaknesses of human and AI debugging approaches, and no systematic framework exists for combining these methods effectively.

Solution

  • Proposed approach: A mixed-initiative human-AI co-debugging system that integrates multimodal clarification, retrieval augmentation, code preview, and hybrid feedback to address visualization debugging challenges.
  • Novelty:
    1. Curated a dataset of 297 Vega-Lite debugging cases from Stack Overflow.
    2. Conducted an empirical study comparing human and AI debugging approaches, revealing their respective strengths and limitations.
    3. Designed and implemented a mixed-initiative system that combines human interpretive strengths with AI generative capabilities.
    4. Evaluated the system through a user study, demonstrating significant improvements in debugging accuracy and efficiency.
  • Procedure and key techniques:
    1. Dataset Curation: Collected and validated 297 Vega-Lite debugging cases from Stack Overflow.
    2. Empirical Study: Analyzed debugging questions, human responses, and AI performance across 297 cases.
    3. System Design: Developed modules for multimodal clarification, retrieval augmentation, code preview, and hybrid feedback.
    4. Evaluation: Conducted a user study with 12 participants using 36 unresolved debugging cases to compare the system with LLM and forum baselines.

Results

  • Concrete findings:
    • The hybrid system resolved 86% of cases, outperforming forum answers (42%) and LLM baselines (42%).
    • First-turn accuracy was 61% for the system, compared to 28% for the LLM baseline.
    • Participants required fewer interaction rounds with the system (1.53 rounds) than with the LLM baseline (2.33 rounds).
  • Advantage over baselines:
    • The system combined the accuracy of human responses with the speed of AI, achieving higher resolution rates and efficiency.
    • Retrieval augmentation and hybrid feedback significantly improved debugging accuracy and user satisfaction.
  • Experiments / evaluation:
    • Dataset: 36 unresolved Vega-Lite debugging cases from Stack Overflow.
    • Metrics: Accuracy (first-turn and final), interaction efficiency (rounds to solution), and user satisfaction (Likert scale ratings).
    • Participants: 12 users with varied debugging expertise.
  • Limitations and future work:
    • Dataset scope was limited to Vega-Lite, and findings may not generalize to other ecosystems.
    • Participant diversity was limited to those with computer science backgrounds.
    • Future work includes extending the system to other visualization grammars, improving dataset diversity, and integrating the framework into community platforms.

Summary

This paper presents a mixed-initiative human-AI co-debugging system for visualization debugging, addressing the limitations of standalone human and AI approaches. By combining multimodal clarification, retrieval augmentation, code preview, and hybrid feedback, the system achieved 86% accuracy in resolving debugging cases, significantly outperforming forum and LLM baselines. A user study demonstrated the system’s effectiveness in improving accuracy, efficiency, and user satisfaction. These findings highlight the complementary roles of humans and AI in visualization debugging and provide a foundation for future collaborative debugging platforms.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222524/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791441
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Interactive Data Visualization, Human-LLM Collaboration, AI-Assisted Decision-Making & Automation
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers