Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging System
Authors
Paper Title
Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging System
Publication Info
- Topic area: Visualization debugging using human-AI collaboration.
- Keywords: Visualization debugging, Vega-Lite, human-AI collaboration, large language models, Stack Overflow, mixed-initiative systems, code generation, empirical study, retrieval augmentation, hybrid feedback.
Background and Problem
- Problem / challenge: Visualization debugging is complex due to its reliance on human visual perception, and users often struggle to resolve issues independently. Existing solutions, such as Q&A forums and Large Language Models (LLMs), have limitations in accuracy, speed, and alignment with user intent.
- Significance: Effective debugging is essential for creating accurate and aesthetically pleasing visualizations, which are critical for data communication and decision-making.
- Motivation and related work: Prior research has explored programmers’ help-seeking behaviors, visualization authoring, and debugging tools. However, there is limited understanding of the strengths and weaknesses of human and AI debugging approaches, and no systematic framework exists for combining these methods effectively.
Solution
- Proposed approach: A mixed-initiative human-AI co-debugging system that integrates multimodal clarification, retrieval augmentation, code preview, and hybrid feedback to address visualization debugging challenges.
- Novelty:
- Curated a dataset of 297 Vega-Lite debugging cases from Stack Overflow.
- Conducted an empirical study comparing human and AI debugging approaches, revealing their respective strengths and limitations.
- Designed and implemented a mixed-initiative system that combines human interpretive strengths with AI generative capabilities.
- Evaluated the system through a user study, demonstrating significant improvements in debugging accuracy and efficiency.
- Procedure and key techniques:
- Dataset Curation: Collected and validated 297 Vega-Lite debugging cases from Stack Overflow.
- Empirical Study: Analyzed debugging questions, human responses, and AI performance across 297 cases.
- System Design: Developed modules for multimodal clarification, retrieval augmentation, code preview, and hybrid feedback.
- Evaluation: Conducted a user study with 12 participants using 36 unresolved debugging cases to compare the system with LLM and forum baselines.
Results
- Concrete findings:
- The hybrid system resolved 86% of cases, outperforming forum answers (42%) and LLM baselines (42%).
- First-turn accuracy was 61% for the system, compared to 28% for the LLM baseline.
- Participants required fewer interaction rounds with the system (1.53 rounds) than with the LLM baseline (2.33 rounds).
- Advantage over baselines:
- The system combined the accuracy of human responses with the speed of AI, achieving higher resolution rates and efficiency.
- Retrieval augmentation and hybrid feedback significantly improved debugging accuracy and user satisfaction.
- Experiments / evaluation:
- Dataset: 36 unresolved Vega-Lite debugging cases from Stack Overflow.
- Metrics: Accuracy (first-turn and final), interaction efficiency (rounds to solution), and user satisfaction (Likert scale ratings).
- Participants: 12 users with varied debugging expertise.
- Limitations and future work:
- Dataset scope was limited to Vega-Lite, and findings may not generalize to other ecosystems.
- Participant diversity was limited to those with computer science backgrounds.
- Future work includes extending the system to other visualization grammars, improving dataset diversity, and integrating the framework into community platforms.
Summary
This paper presents a mixed-initiative human-AI co-debugging system for visualization debugging, addressing the limitations of standalone human and AI approaches. By combining multimodal clarification, retrieval augmentation, code preview, and hybrid feedback, the system achieved 86% accuracy in resolving debugging cases, significantly outperforming forum and LLM baselines. A user study demonstrated the system’s effectiveness in improving accuracy, efficiency, and user satisfaction. These findings highlight the complementary roles of humans and AI in visualization debugging and provide a foundation for future collaborative debugging platforms.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences
UIST '25· Human-LLM Collaboration +1
- 75%
Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
IUI '26· Human-LLM Collaboration +3
- 71%
Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models
CHI '19· Explainable AI (XAI) +2
- 71%
Dango: A Mixed-Initiative Data Wrangling System using Large Language Model
CHI '25· Human-LLM Collaboration +2
- 71%
Finding Needles in Document Haystacks: Augmenting Serendipitous Claim Retrieval Workflows
CHI '25· Human-LLM Collaboration +2
- 71%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
- 71%
The Bots of Persuasion: Examining How Conversational Agents' Linguistic Expressions of Personality Affect User Perceptions and Decisions
CHI '26· Agent Personality & Anthropomorphism +2
- 71%
Belief Updating and Delegation in Multi-Task Human–AI Interaction: Evidence from Controlled Simulations
CHI '26· Human-LLM Collaboration +2
- 71%
Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
IUI '26· Human-LLM Collaboration +2
- 71%
Making Absence Visible in Intelligent Summarization Interfaces
IUI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)