CrossData: Leveraging Text-Data Connections for Authoring Data Documents

Interactive Data VisualizationComputational Methods in HCIData Scientists & AnalystsHCI ResearchersStatisticians & Data Scientists

Document Title

CrossData: Leveraging Text-Data Connections for Authoring Data Documents

Document Information

  • Domain: Human-Computer Interaction (HCI) and data-driven document authoring
  • Keywords: data documents, natural language processing, interactive articles, text-data binding, data visualization, data consistency, dynamic documents, information consumption, data-rich narratives

Research Background and Problem

  • Problem: Authoring data documents often involves tedious tasks, including data analysis, converting between data and text, and maintaining consistency, leading to repetitive, inefficient, and error-prone workflows.
  • Significance: Data documents are critical tools for recording and disseminating data, widely used in fields such as scientific research, finance, public health, education, and journalism. As the world becomes increasingly data-driven, optimizing the data document authoring process has become an urgent issue.
  • Research Motivation:
    • The lack of persistent connections between text and data is a key factor contributing to inefficiency. Most existing tools require users to manually maintain consistency between data and text.
    • Current research systems supporting dynamic and interactive data document authoring often require programming knowledge, posing a high barrier for non-expert users.
  • Related Work: Tools for creating data-driven content, research on linking text to other visual media, and the application of natural language interfaces in data querying and visualization provide the foundation for this study.

Solution

  • Method/Solution:
    • Proposes a language-driven data binding mechanism to extract potential text-data connections from descriptive text.
    • Develops a prototype system called CrossData, which automatically identifies and maintains text-data connections, transforming them into persistent, interactive, first-class objects.
    • Leverages natural language processing (NLP) techniques to build a connection engine that automatically detects, establishes, and maintains text-data connections during the writing process.
    • Provides a suite of interaction techniques, such as data retrieval, computed values, consistency updates, and automatic adjustments for dynamic tables and charts.
  • Innovations:
    • Significantly lowers the usage barrier for non-technical users through a language-driven data binding design without requiring coding.
    • Establishes a close connection between document authoring and data exploration.
    • Introduces new concepts of data-driven text and interactive data documents.
  • Implementation Steps:
    • Build a Connection Engine to extract text-data connections.
    • Develop a tool interface that allows users to select, update, and correct connections.
    • Provide automatic data synchronization to maintain consistency among text, tables, and charts.

Research Outcomes

  • Specific Outcomes:
    • Proposed a language-driven data binding method and a set of interaction techniques to address user pain points.
    • Validated the reliability of the connection engine through technical evaluation, successfully constructing 88.8% of text-data connections.
    • Expert evaluations revealed that CrossData significantly reduced the manual workload of data document authoring while providing a seamless workflow.
  • Advantages over Existing Solutions:
    • Enables text-data binding without programming, reducing the learning curve for non-technical users.
    • Achieves real-time data updates and dynamic text interaction, avoiding consistency issues found in traditional tools.
    • Improves collaborative workflows, making it easier for authors to respond to real-time feedback.
  • Experimental or Evaluation Results:
    • Technical evaluations confirmed the efficiency and accuracy of the connection engine in data documents across multiple domains (88.8% connection success rate).
    • Expert evaluations demonstrated that various interaction techniques effectively addressed pain points in data retrieval, consistency maintenance, and iterative data document creation, receiving widespread recognition.
  • Limitations and Future Directions:
    • Limitations:
      • Challenges remain in handling complex expressions such as contextual references and approximate values.
      • Limited scalability and customizability of operational support.
      • Potential performance issues when dealing with large-scale datasets.
    • Future Directions:
      • Support for more complex data structures (e.g., graph or tree data) and customizable operations.
      • Integration of the system into programming tools (e.g., Jupyter Notebook) to combine language-driven authoring with programming.
      • Exploration of dynamic presentation forms based on text-data connections, such as data animations and videos.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/71927/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3517485
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Interactive Data Visualization, Computational Methods in HCI
work
Professions
Data Scientists & Analysts, HCI Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers