CrossData: Leveraging Text-Data Connections for Authoring Data Documents
Interactive Data VisualizationComputational Methods in HCIData Scientists & AnalystsHCI ResearchersStatisticians & Data Scientists
Document Title
CrossData: Leveraging Text-Data Connections for Authoring Data Documents
Document Information
- Domain: Human-Computer Interaction (HCI) and data-driven document authoring
- Keywords: data documents, natural language processing, interactive articles, text-data binding, data visualization, data consistency, dynamic documents, information consumption, data-rich narratives
Research Background and Problem
- Problem: Authoring data documents often involves tedious tasks, including data analysis, converting between data and text, and maintaining consistency, leading to repetitive, inefficient, and error-prone workflows.
- Significance: Data documents are critical tools for recording and disseminating data, widely used in fields such as scientific research, finance, public health, education, and journalism. As the world becomes increasingly data-driven, optimizing the data document authoring process has become an urgent issue.
- Research Motivation:
- The lack of persistent connections between text and data is a key factor contributing to inefficiency. Most existing tools require users to manually maintain consistency between data and text.
- Current research systems supporting dynamic and interactive data document authoring often require programming knowledge, posing a high barrier for non-expert users.
- Related Work: Tools for creating data-driven content, research on linking text to other visual media, and the application of natural language interfaces in data querying and visualization provide the foundation for this study.
Solution
- Method/Solution:
- Proposes a language-driven data binding mechanism to extract potential text-data connections from descriptive text.
- Develops a prototype system called CrossData, which automatically identifies and maintains text-data connections, transforming them into persistent, interactive, first-class objects.
- Leverages natural language processing (NLP) techniques to build a connection engine that automatically detects, establishes, and maintains text-data connections during the writing process.
- Provides a suite of interaction techniques, such as data retrieval, computed values, consistency updates, and automatic adjustments for dynamic tables and charts.
- Innovations:
- Significantly lowers the usage barrier for non-technical users through a language-driven data binding design without requiring coding.
- Establishes a close connection between document authoring and data exploration.
- Introduces new concepts of data-driven text and interactive data documents.
- Implementation Steps:
- Build a Connection Engine to extract text-data connections.
- Develop a tool interface that allows users to select, update, and correct connections.
- Provide automatic data synchronization to maintain consistency among text, tables, and charts.
Research Outcomes
- Specific Outcomes:
- Proposed a language-driven data binding method and a set of interaction techniques to address user pain points.
- Validated the reliability of the connection engine through technical evaluation, successfully constructing 88.8% of text-data connections.
- Expert evaluations revealed that CrossData significantly reduced the manual workload of data document authoring while providing a seamless workflow.
- Advantages over Existing Solutions:
- Enables text-data binding without programming, reducing the learning curve for non-technical users.
- Achieves real-time data updates and dynamic text interaction, avoiding consistency issues found in traditional tools.
- Improves collaborative workflows, making it easier for authors to respond to real-time feedback.
- Experimental or Evaluation Results:
- Technical evaluations confirmed the efficiency and accuracy of the connection engine in data documents across multiple domains (88.8% connection success rate).
- Expert evaluations demonstrated that various interaction techniques effectively addressed pain points in data retrieval, consistency maintenance, and iterative data document creation, receiving widespread recognition.
- Limitations and Future Directions:
- Limitations:
- Challenges remain in handling complex expressions such as contextual references and approximate values.
- Limited scalability and customizability of operational support.
- Potential performance issues when dealing with large-scale datasets.
- Future Directions:
- Support for more complex data structures (e.g., graph or tree data) and customizable operations.
- Integration of the system into programming tools (e.g., Jupyter Notebook) to combine language-driven authoring with programming.
- Exploration of dynamic presentation forms based on text-data connections, such as data animations and videos.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can language-driven data binding mechanisms extract connections between text and data?Category: Hardware Maintenance and Asynchronous Problem-Solving SupportSimilar questionsarrow_forward
- How can language-driven text-data connections lower the barrier for non-technical users creating data documents?Category: Hardware Maintenance and Asynchronous Problem-Solving SupportSimilar questionsarrow_forward
- How can automated text-data binding improve data document creation workflows and consistency maintenance?Category: Hardware Maintenance and Asynchronous Problem-Solving SupportSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Authors find maintaining consistency between text and data in data documents tedious and error-prone.Category: Hardware Maintenance and Asynchronous Problem-Solving SupportSimilar questionsarrow_forward
- 83%
HiLT: A Library for Generating Human-in-the-Loop Data Transformation GUIs
UIST '25· Interactive Data Visualization +1
- 80%
Tessera: Discretizing Data Analysis Workflows on a Task Level
CHI '21· Interactive Data Visualization +1
- 80%
ComputableViz: Mathematical Operators as a Formalism for Visualization Processing and Analysis
CHI '22· Interactive Data Visualization +1
- 80%
Chartist: Task-driven Eye Movement Control for Chart Reading
CHI '25· Interactive Data Visualization +1
- 80%
B2: Bridging Code and Interactive Visualization in Computational Notebooks
UIST '20· Interactive Data Visualization +1
- 67%
Evaluating Preference Collection Methods for Interactive Ranking Analytics
CHI '19· Interactive Data Visualization +1
- 67%
How Data Analysts Use a Visualization Grammar in Practice
CHI '23· Interactive Data Visualization +2
- 67%
Using Open Data to Automatically Generate Localized Analogies
CHI '24· Interactive Data Visualization +2
- 67%
Data Formulator 2: Iterative Creation of Data Visualizations, with AI Transforming Data Along the Way
CHI '25· AI-Assisted Decision-Making & Automation +2
- 67%
Automatic Synthesis of Visualization Design Knowledge Bases
CHI '26· Interactive Data Visualization +2
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3517485
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Interactive Data Visualization, Computational Methods in HCI
work
Professions
Data Scientists & Analysts, HCI Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers