Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models

Human-LLM CollaborationInteractive Data VisualizationTime-Series & Network Graph VisualizationSoftware Engineers & DevelopersUI/UX DesignersData Scientists & AnalystsAI/ML Researchers & Engineers

Title of the Paper

Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models

Paper Information

  • Subject Area: Data visualization and natural language interfaces (NLIs), combined with applications of large language models (LLMs)
  • Keywords: Vega-Lite, natural language dataset, large language models, framework, natural language interface, data visualization

Research Background and Problem

  • What problems or challenges did the authors identify?

    • The scarcity of high-quality paired datasets (charts and natural language descriptions) hinders the development of natural language interfaces.
    • Existing chart datasets are often synthetic, lacking diversity in types and complexity in content, or are limited in scale.
    • The dataset creation process typically relies on crowdsourcing, which is costly, time-consuming, and struggles to capture linguistic diversity.
  • Why is this problem important?

    • NLIs can significantly simplify the use of data visualization tools and enhance the intuitiveness of human-computer interaction. However, effective NLIs rely on high-quality datasets, and the lack of such datasets limits technological progress.
  • Research Motivation and Related Work

    • Previous studies have constructed some chart-NL pairs through synthetic datasets and crowdsourcing methods, but most approaches are difficult to scale and fail to cover complex chart requirements.
    • Specific tasks like chart question answering and chart generation further fragment datasets, with each requiring different types of data.

Solution

  • What methods or solutions did the authors propose?

    • The authors proposed the VL2NL framework, which is based on the Vega-Lite specification and utilizes large language models to generate rich and diverse natural language datasets.
    • The new dataset includes 1,981 real-world Vega-Lite charts collected from GitHub, featuring high complexity and diversity.
    • A three-stage method was employed:
      1. Dataset Preprocessing: Optimizing Vega-Lite specifications to simplify inputs.
      2. Guided Discovery-Based Generation: Using "scaffolding" and key questions to guide the model in self-driven generation of more accurate NL datasets.
      3. Score-Based Rewriting: Enhancing linguistic diversity through four language axes (clarity, formality, subjectivity, professionalism).
  • What are the innovative aspects of this solution?

    • Enabling the model to generate datasets through self-guided discovery mechanisms, improving accuracy and linguistic diversity.
    • Introducing a novel score-based rewriting method that allows controlled generation of grammatically varied but semantically consistent diverse data.
    • The dataset is not only large in scale but also covers complex and richly interactive charts, which were rarely seen in previous work.
  • What are the implementation steps and key technologies used?

    • Using the GitHub API to collect and filter Vega-Lite specifications that meet the criteria.
    • Generating chart semantics and multi-layered natural language descriptions through LLMs, combined with Chain-of-Thought prompting strategies for guidance.
    • Applying rewriting techniques to generate synonymous sentences with varied grammatical styles based on the four-axis scoring.

Research Outcomes

  • What specific results were achieved?

    • The dataset includes 1,981 Vega-Lite charts, with complexity increased fourfold and interactivity increased fivefold compared to baseline datasets.
    • VL2NL-generated L1 and L2 descriptions achieved 89.4% and 76.0% accuracy, respectively.
    • The generated NL dataset demonstrated higher diversity across multiple statistical metrics.
  • What advantages does it have compared to existing solutions?

    • Significant improvements in data diversity and complexity, providing richer data support for applications involving complex charts and interactivity.
    • The generated natural language sentences are not only grammatically diverse but also expand the coverage of training data for NLIs.
  • What were the experimental or evaluation results?

    • Tests on charts of varying complexity showed that the VL2NL framework can robustly generate high-quality descriptions, questions, and sentences.
    • Refined experiments (including comparisons with baseline datasets and generation of rewritten sentences) demonstrated superiority in evaluation metrics.
  • Limitations and Future Directions

    • Limitations: Reliance on Vega-Lite specifications cannot capture implicitly expressed information; the generated NL dataset is limited to the current language axes and data formats used.
    • Future Directions:
      • Enhancing framework capabilities, such as supplementing common sense and domain knowledge through external resource access.
      • Expanding the scope of natural language dataset generation to cover interactive dialogues and multi-document references.
      • Developing enhanced methods for Vega-Lite specifications to increase the quantity and diversity of data specifications.

This paper provides a foundational and comprehensive framework for generating natural language datasets, which can significantly advance research in the field of data visualization and NLIs, showcasing the powerful application potential of LLMs.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147000/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642943
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Human-LLM Collaboration, Interactive Data Visualization, Time-Series & Network Graph Visualization
work
Professions
Software Engineers & Developers, UI/UX Designers, Data Scientists & Analysts, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers