Collecting and Characterizing Natural Language Utterances for Specifying Data Visualizations

Voice User Interface (VUI) DesignInteractive Data VisualizationSoftware Engineers & DevelopersUI/UX Designers

Title of the Paper

Collecting and Characterizing Natural Language Utterances for Specifying Data Visualizations

Paper Information

  • Subject Area: Data Visualization and Natural Language Interfaces
  • Keywords: Data Visualization, Natural Language Interfaces, Natural Language Processing, Natural Language Dataset, Visualization Specification

Research Background and Problem

  • Problem and Challenges: Natural language interfaces (NLIs) for data visualization are becoming increasingly popular, but there is currently a lack of empirical understanding of how users describe visualizations in natural language. For example, what types of natural language expressions do users commonly use? What information (e.g., chart types, data aggregation) do their inputs contain? How do existing systems support these habits and content?
  • Significance: Understanding users' natural language inputs can improve the design of existing visualization systems and enhance user experience, especially for novice users and those with specific needs (e.g., individuals with disabilities).
  • Research Motivation: This study aims to fill the gap in systematic understanding of users' natural language inputs and to facilitate the development, evaluation, and optimization of NLIs through the creation of a dataset.

Solution

  • Research Methods:

    1. Online User Study: Researchers conducted an online experiment with 102 participants, asking them to describe specific visualization charts using natural language.
    2. Data Collection and Cleaning: They compiled 893 natural language input descriptions (referred to as "utterances") from participants' inputs and performed data filtering and cleaning.
    3. Data Classification and Feature Analysis: The collected natural language inputs were categorized and analyzed, focusing on their syntactic structures and the visualization information they contained.
  • Innovations:

    1. Proposed a dedicated research method for collecting users' natural language inputs, employing iterative design optimization to ensure consistency between research content and user behavior.
    2. Created the first publicly available natural language dataset for visualization specification, covering common chart types and related linguistic data.
  • Implementation Steps and Techniques:

    • Used 10 common chart types (e.g., bar charts, scatter plots) and 3 real-world datasets containing categorical, numerical, and temporal data (e.g., automobile datasets).
    • Participants completed tasks in randomized order, and the study meticulously recorded and excluded invalid entries.
    • Open-source analysis tools (e.g., NL4DV) were used to evaluate how well the sample dataset supports existing NLIs.

Research Outcomes

  • Specific Results:

    1. Collected and cleaned a standardized dataset of 893 natural language inputs, encompassing common chart types and data attributes.
    2. Conducted a twofold analysis of natural language inputs:
      • Syntactic Structure: Natural language inputs were primarily categorized as imperative inputs (45%), query-based inputs (32%), question-based inputs (14%), and a small proportion of other descriptive forms.
      • Input Information: User inputs contained various information tags, such as explicitly specified attributes, implicit or explicit aggregation methods, chart types, and detailed encodings.
  • Comparison with Existing Solutions:

    • The data revealed that existing NLIs (e.g., the NL4DV tool) have limitations in parsing users' natural language inputs, such as weak support for mapping semantically related attributes.
    • The dataset demonstrated potential for improving NLI development when used to train machine learning models (e.g., TF-IDF classifiers).
  • Experimental or Evaluation Results:

    • Benchmark testing of the dataset with NL4DV showed that the system could produce partial or complete matches for approximately 75% of input tests.
    • When the dataset was used for classifier development, simple machine learning models (e.g., logistic regression) achieved up to 88% accuracy.
  • Limitations and Future Directions:

    1. Input forms were limited to text and did not include other interaction modes such as speech.
    2. The dataset size is relatively small and needs to be expanded through broader crowdsourcing or targeted user groups (e.g., visualization novices).
    3. The current study only addresses natural language inputs for single charts; future research could explore multi-chart or dashboard input scenarios.
    4. The data cleaning and classification process currently relies on manual operations, necessitating the development of more automated cleaning tools and rule libraries in the future.

Conclusion

  • This study represents the first systematic investigation into natural language descriptions for data visualization, collecting and analyzing users' natural language input data and providing a publicly available dataset for future research.
  • The dataset and research findings offer valuable references for improving current NLI designs and open new directions for applying natural language processing technologies in data visualization.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47729/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445400
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Voice User Interface (VUI) Design, Interactive Data Visualization
work
Professions
Software Engineers & Developers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers