Collecting and Characterizing Natural Language Utterances for Specifying Data Visualizations
Authors
Title of the Paper
Collecting and Characterizing Natural Language Utterances for Specifying Data Visualizations
Paper Information
- Subject Area: Data Visualization and Natural Language Interfaces
- Keywords: Data Visualization, Natural Language Interfaces, Natural Language Processing, Natural Language Dataset, Visualization Specification
Research Background and Problem
- Problem and Challenges: Natural language interfaces (NLIs) for data visualization are becoming increasingly popular, but there is currently a lack of empirical understanding of how users describe visualizations in natural language. For example, what types of natural language expressions do users commonly use? What information (e.g., chart types, data aggregation) do their inputs contain? How do existing systems support these habits and content?
- Significance: Understanding users' natural language inputs can improve the design of existing visualization systems and enhance user experience, especially for novice users and those with specific needs (e.g., individuals with disabilities).
- Research Motivation: This study aims to fill the gap in systematic understanding of users' natural language inputs and to facilitate the development, evaluation, and optimization of NLIs through the creation of a dataset.
Solution
-
Research Methods:
- Online User Study: Researchers conducted an online experiment with 102 participants, asking them to describe specific visualization charts using natural language.
- Data Collection and Cleaning: They compiled 893 natural language input descriptions (referred to as "utterances") from participants' inputs and performed data filtering and cleaning.
- Data Classification and Feature Analysis: The collected natural language inputs were categorized and analyzed, focusing on their syntactic structures and the visualization information they contained.
-
Innovations:
- Proposed a dedicated research method for collecting users' natural language inputs, employing iterative design optimization to ensure consistency between research content and user behavior.
- Created the first publicly available natural language dataset for visualization specification, covering common chart types and related linguistic data.
-
Implementation Steps and Techniques:
- Used 10 common chart types (e.g., bar charts, scatter plots) and 3 real-world datasets containing categorical, numerical, and temporal data (e.g., automobile datasets).
- Participants completed tasks in randomized order, and the study meticulously recorded and excluded invalid entries.
- Open-source analysis tools (e.g., NL4DV) were used to evaluate how well the sample dataset supports existing NLIs.
Research Outcomes
-
Specific Results:
- Collected and cleaned a standardized dataset of 893 natural language inputs, encompassing common chart types and data attributes.
- Conducted a twofold analysis of natural language inputs:
- Syntactic Structure: Natural language inputs were primarily categorized as imperative inputs (45%), query-based inputs (32%), question-based inputs (14%), and a small proportion of other descriptive forms.
- Input Information: User inputs contained various information tags, such as explicitly specified attributes, implicit or explicit aggregation methods, chart types, and detailed encodings.
-
Comparison with Existing Solutions:
- The data revealed that existing NLIs (e.g., the NL4DV tool) have limitations in parsing users' natural language inputs, such as weak support for mapping semantically related attributes.
- The dataset demonstrated potential for improving NLI development when used to train machine learning models (e.g., TF-IDF classifiers).
-
Experimental or Evaluation Results:
- Benchmark testing of the dataset with NL4DV showed that the system could produce partial or complete matches for approximately 75% of input tests.
- When the dataset was used for classifier development, simple machine learning models (e.g., logistic regression) achieved up to 88% accuracy.
-
Limitations and Future Directions:
- Input forms were limited to text and did not include other interaction modes such as speech.
- The dataset size is relatively small and needs to be expanded through broader crowdsourcing or targeted user groups (e.g., visualization novices).
- The current study only addresses natural language inputs for single charts; future research could explore multi-chart or dashboard input scenarios.
- The data cleaning and classification process currently relies on manual operations, necessitating the development of more automated cleaning tools and rule libraries in the future.
Conclusion
- This study represents the first systematic investigation into natural language descriptions for data visualization, collecting and analyzing users' natural language input data and providing a publicly available dataset for future research.
- The dataset and research findings offer valuable references for improving current NLI designs and open new directions for applying natural language processing technologies in data visualization.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What natural language expression forms do users commonly use when describing data visualizations?Category: Natural Language-Driven Data VisualizationSimilar questionsarrow_forward
- What information (e.g., chart type, data aggregation method) is typically included in these natural language inputs?Category: Natural Language-Driven Data VisualizationSimilar questionsarrow_forward
- What limitations do existing natural language interfaces (NLIs) have in supporting users' descriptions of data visualizations?Category: Natural Language-Driven Data VisualizationSimilar questionsarrow_forward
Practical Problems
1- Non-expert users lack technical terminology and struggle to describe data visualization needs in natural language.Category: Natural Language-Driven Data VisualizationSimilar questionsarrow_forward
- 80%
InChorus: Designing Consistent Multimodal Interactions for Data Visualization on Tablet Devices
CHI '20· Voice User Interface (VUI) Design +2
- 75%
Interactive Document Clustering Revisited: A Visual Analytics Approach
IUI '18· Interactive Data Visualization
- 75%
Data-centric disambiguation for data transformation with programming-by-example
IUI '21· Interactive Data Visualization
- 60%
What's the Difference?: Evaluating Variations of Multi-Series Bar Charts for Visual Comparison Tasks
CHI '18· Interactive Data Visualization +1
- 60%
The Impact of Word, Multiple Word, and Sentence Input on Virtual Keyboard Decoding Performance
CHI '18· Voice User Interface (VUI) Design +1
- 60%
The Effects of Adding Search Functionality to Interactive Visualizations on the Web
CHI '18· Interactive Data Visualization +1
- 60%
Interactive Repair of Tables Extracted from PDF Documents on Mobile Devices
CHI '19· Interactive Data Visualization
- 60%
EYEditor: Towards On-the-Go Heads-Up Text Editing Using Voice and Manual Input
CHI '20· Head-Up Display (HUD) & Advanced Driver Assistance Systems (ADAS) +1
- 60%
Does Interaction Improve Bayesian Reasoning with Visualization?
CHI '21· Interactive Data Visualization +1
- 60%
Integrated Visualization Editing via Parameterized Declarative Templates
CHI '21· Interactive Data Visualization +1
Based on Jaccard similarity of research subtopics & professions (≥60%)