Revealing the Gap: Visual Comparison of Large-Scale Datasets via Multi-Scale Density Difference Map

Interactive Data VisualizationVisualization Perception & CognitionData Scientists & AnalystsAI/ML Researchers & EngineersHCI Researchers

Paper Title

Revealing the Gap: Visual Comparison of Large-Scale Datasets via Multi-Scale Density Difference Map

Publication Info

  • Topic area: Visual comparison of high-dimensional datasets for machine learning applications.
  • Keywords: Visual comparison, density difference map, large-scale datasets, machine learning, dataset visualization, multi-scale analysis, interactive tool, synthetic data, real-world data, hierarchical grids.

Background and Problem

  • Problem / challenge: Existing density-based visualization methods struggle with balancing global and local comparisons, fail to reveal minority dataset distributions in cluttered regions, and are computationally inefficient for large datasets.
  • Significance: Understanding dataset distribution differences is critical for machine learning tasks like classification and object detection, ensuring representativeness in sample selection and mitigating biases in synthetic data usage.
  • Motivation and related work: Prior methods such as FID, KL divergence, and kernel density estimation (KDE) quantify differences but do not visually explain them. Scatterplot-based and density-based methods have limitations in handling large-scale datasets with overlapping distributions.

Solution

  • Proposed approach: DiffGrid, a grid-based tool for visualizing density differences between large datasets, featuring regularized layouts, efficient computation, and hierarchical detail exploration.
  • Novelty:
    1. A regularized density difference visualization method for balanced multi-level analysis.
    2. An interactive tool integrating hierarchical grids and image labels for progressive detail exploration.
    3. Enhanced efficiency via an optimized Least-Squares Density Difference (LSDD+) algorithm.
  • Procedure and key techniques:
    • Regularization: Integral image transformation spreads points evenly while preserving cluster density.
    • LSDD+: Efficient density difference estimation using grid-based kernel centers and GPU-optimized L-BFGS solver.
    • Multi-scale exploration: Hierarchical grids refine details progressively during zooming, with inset squares indicating minority dataset presence.
    • Image label placement: Quadratic assignment problem optimizes label positioning to avoid occlusion and enhance interpretability.

Results

  • Concrete findings:
    • LSDD+ achieves up to 62.19× speedup over conventional LSDD, scaling to datasets with 5,000,000 samples while maintaining near-second runtime at 1,000,000 samples.
    • Mean squared error (MSE) decreases with dataset size, ensuring accurate density difference estimation.
    • User study shows DiffGrid improves accuracy in global (96.9% vs. 78.1%) and local (90.6% vs. 25.0%) tasks compared to baseline, with comparable task completion times.
  • Advantage over baselines:
    • Superior accuracy in identifying global and local differences.
    • Enhanced scalability and efficiency for large datasets.
    • Better usability and satisfaction ratings in user experience evaluations.
  • Experiments / evaluation:
    • Quantitative experiment comparing LSDD+ with LSDD on synthetic datasets.
    • Controlled user study with 16 participants evaluating task accuracy, response time, and subjective experience.
    • Case studies analyzing coresets vs. full datasets and synthetic vs. real-world infographics.
  • Limitations and future work:
    • Limited support for multi-dataset comparisons; future work could address the mental load of pairwise comparisons.
    • Lack of uncertainty-aware comparison for synthetic data; future exploration of ensemble methods is suggested.

Summary

DiffGrid introduces a novel grid-based visualization method for comparing large-scale datasets, addressing limitations in existing density-based techniques. By combining regularized layouts, efficient LSDD+ computation, and hierarchical exploration, it enables accurate identification of global and local differences while scaling to millions of samples. Quantitative experiments, user studies, and case studies demonstrate its effectiveness in tasks like sample selection and synthetic data analysis. Future work aims to extend its capabilities to multi-dataset comparisons and uncertainty-aware visualizations, enhancing its applicability in diverse machine learning contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222763/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791171
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Interactive Data Visualization, Visualization Perception & Cognition
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers