Automatic Synthesis of Visualization Design Knowledge Bases
Authors
Paper Title
Automatic Synthesis of Visualization Design Knowledge Bases
Publication Info
- Topic area: Automated generation of visualization design knowledge bases for effective visualization recommendations.
- Keywords: Visualization design, knowledge base synthesis, feature selection, machine learning, Draco 2, genomics visualization, data-driven methods, logical reasoning, design recommendation, automated reasoning.
Background and Problem
- Problem / challenge: Existing visualization design knowledge bases rely on fixed, manually authored rules, which limit their adaptability to new domains and make it difficult to extend or generalize them.
- Significance: Automating the synthesis of knowledge bases can enable scalable, interpretable, and domain-specific visualization recommendations, reducing the manual effort required for creating and updating these systems.
- Motivation and related work: Prior systems like Draco and KG4Vis encode design knowledge as fixed features, which restricts their ability to learn new constraints or adapt to specialized domains like genomics. This paper addresses the need for a generic, automated method to construct knowledge bases from visualization corpora.
Solution
- Proposed approach: A data-driven method for synthesizing visualization design knowledge bases by extracting, selecting, and rendering design features from a corpus of visualization pairs.
- Novelty:
- Automated extraction and selection of meaningful and generalizable design features.
- Benchmark evaluation against Draco 2, demonstrating improved prediction accuracy and generalization.
- Application to genomics visualization, showcasing adaptability to specialized domains.
- Procedure and key techniques:
- Extraction: Enumerate candidate features from visualization corpora using abstract syntax trees and frequency vectors.
- Selection: Use forward and backward feature selection with pre-selection metrics (e.g., meaningfulness, generalizability, usefulness) to identify a subset of features.
- Rendering: Convert selected features into formal representations compatible with Draco for logical reasoning and automated recommendations.
Results
- Concrete findings:
- Benchmark study: Synthesized features achieved 1–15% higher prediction accuracy compared to Draco 2, with accuracy up to 94% on holdout sets.
- Genomics study: Synthesized knowledge base achieved up to 97% prediction accuracy, capturing domain-specific features like interactivity and multi-view layouts.
- Advantage over baselines:
- Smaller feature sets (34–96 features vs. Draco’s 147) with competitive or superior performance.
- Improved generalization to sparse datasets (e.g., Zeng+ corpus).
- Experiments / evaluation:
- Benchmark: Used 1,384 chart pairs from graphical perception studies, comparing synthesized features to Draco 2.
- Genomics: Curated 296 labeled chart pairs, demonstrating the method's applicability to domain-specific visualizations.
- Expert feedback: Genomics experts rated 75–80% of selected features as helpful or very helpful.
- Limitations and future work:
- Limited corpus size and scope may omit recent or specialized design cases.
- Future work could explore broader domains (e.g., uncertainty visualization) and integrate task-specific features.
Summary
This paper presents a method for automatically synthesizing visualization design knowledge bases from ranked visualization pairs. The approach extracts, selects, and renders design features into formal representations, achieving competitive or superior performance compared to manually authored systems like Draco 2. Benchmark evaluations and a case study in genomics visualization demonstrate the method's generalizability and domain adaptability. Future work could expand its application to other visualization domains and refine corpus preparation for more comprehensive feature learning.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
Tessera: Discretizing Data Analysis Workflows on a Task Level
CHI '21· Interactive Data Visualization +1
- 80%
Chartist: Task-driven Eye Movement Control for Chart Reading
CHI '25· Interactive Data Visualization +1
- 80%
B2: Bridging Code and Interactive Visualization in Computational Notebooks
UIST '20· Interactive Data Visualization +1
- 67%
CrossData: Leveraging Text-Data Connections for Authoring Data Documents
CHI '22· Interactive Data Visualization +1
- 67%
How Data Analysts Use a Visualization Grammar in Practice
CHI '23· Interactive Data Visualization +2
- 60%
The Story in the Notebook: Exploratory Data Science using a Literate Programming Tool
CHI '18· Interactive Data Visualization +1
- 60%
How Data Science Workers Work with Data: Discovery, Capture, Curation, Design, Creation
CHI '19· Interactive Data Visualization +1
- 60%
Falx: Synthesis-Powered Visualization Authoring
CHI '21· Interactive Data Visualization
- 60%
From Detectables to Inspectables: Understanding Qualitative Analysis of Audiovisual Data
CHI '21· Interactive Data Visualization +1
- 60%
Interaction Illustration Taxonomy: Classification of Styles and Techniques for Visually Representing Interaction Scenarios
CHI '21· Interactive Data Visualization +1
Based on Jaccard similarity of research subtopics & professions (≥60%)