KTabulator: Interactive Ad hoc Table Creation Using Knowledge Graphs
Authors
Document Title
KTabulator: Interactive Ad hoc Table Creation using Knowledge Graphs
Document Information
- Subject Area: Human-Computer Interaction, Knowledge Graphs, and Data Integration Tools
- Keywords: Data tables, Data cleaning, Data integration, Databases, User interface
Research Background and Issues
-
What problems or challenges did the authors identify?
- Users frequently need to organize and construct data tables in daily life, but obtaining relevant data from the web is often hindered by its dispersion across multiple tables on different web pages. This requires manual searching, copying, and pasting, which is a tedious task.
- Even when relevant tables are found, they are often incomplete. Existing systems typically cannot easily fill in the missing information, leaving users to manually merge or join other tables using spreadsheet software.
- Knowledge graphs (e.g., DBpedia and Wikipedia) have structured many entity data into RDF triples. However, most current systems only allow software to process this data without enabling users to directly leverage knowledge graphs interactively.
-
Why is this problem important?
- Tables are a common format for users to organize, explore, and analyze data. Addressing this problem can significantly improve users' efficiency in handling data, reduce manual operations, and enhance the accuracy of data integration.
- The computational structure of knowledge graphs offers potential for information extraction and table construction, but a user-friendly tool is needed to translate this potential into practical applications.
-
Research Motivation and Related Work
- The authors highlighted that some existing systems partially provide data querying and exploration functions, such as SPARQL language and natural language question-answering tools, as well as data cleaning and table transformation tools (e.g., Wrangler and OpenRefine). However, these systems do not support users in iterative interactions to expand data tables.
- The authors designed an interactive system—KTabulator—that flexibly extracts, constructs, or expands specific data tables from knowledge graphs, empowering users to directly utilize knowledge graphs in a human-in-the-loop process.
Solution
-
What methods or solutions did the authors propose?
- Developed an interactive system called KTabulator, enabling users to efficiently construct conditional data tables using knowledge graphs (e.g., DBpedia) and Wikipedia.
- Leveraged a familiar spreadsheet interface as the primary interaction tool, combined with dynamic recommendations and guidance from knowledge graph data (e.g., entities, attributes) to help users iteratively expand data tables.
-
What are the innovative aspects of this solution?
- Human-computer collaboration: Supports users in directly adding relevant entities and attributes from knowledge graphs to expand table rows and columns.
- Table expansion operations: Dynamically expands columns based on entity attribute recommendations and joins or merges rows or columns by combining multiple related tables.
- Handling data heterogeneity: The system is designed to handle heterogeneous, inconsistent, and even sparse data, providing data density and table recommendations.
- Integration of knowledge graphs and data sources: Performs semantic analysis of existing tables on Wikipedia pages, enabling table joining and matching, and integrates data from knowledge graphs.
-
What are the implementation steps and key technologies used?
- Data collection: Extracted data from DBpedia and retrieved entities and their attributes from the knowledge graph using SPARQL query language.
- Data recommendation: Recommended the most relevant entity attributes using ontology semantic information from DBpedia.
- Table joining and matching: Extracted tables from Wikipedia pages and performed semantic matching based on column type trees.
- System architecture: Included a front-end user interface and back-end service modules, integrating a SPARQL query translator, table matcher, and attribute recommender.
Research Outcomes
-
What specific results were achieved?
- KTabulator significantly improved users' efficiency in creating ad hoc tables using knowledge graphs and Wikipedia.
- Users were able to efficiently expand table rows and columns, filter data, and find related tables that could be joined with the target table.
- User studies demonstrated that the system is engaging and particularly helpful in enabling users to collect and organize information in a human-in-the-loop process.
-
What advantages does it have compared to existing solutions?
- Compared to traditional data cleaning tools (e.g., Excel), KTabulator reduces the tedious manual data input and search operations, helping users efficiently integrate data through guidance and recommendations.
- Unlike existing knowledge graph query systems, this tool supports iterative table creation through interactive expansion and dynamic recommendation features.
-
What were the experimental or evaluation results?
- User studies revealed that most participants were able to effectively complete table creation tasks using KTabulator, including both predefined goal-oriented tasks and open-ended exploration tasks.
- Users found the system easy to operate, and the guidance features effectively helped them explore relevant attributes and data within the knowledge graph.
- Users generally appreciated features such as table recommendations, entity previews, and attribute selection, though feedback also suggested areas for functional improvement.
-
Limitations and Future Directions
- Limitations: The system has slow processing times and lacks advanced table operations (e.g., pivot tables, range filtering) and direct exploration of cross-knowledge-graph entities.
- Future directions:
- Improve system performance, such as by introducing local data storage, caching mechanisms, and progressive loading.
- Expand support for other data sources and knowledge graphs, designing cross-domain data table integration functionalities.
- Enhance user interface design to make multi-hop exploration more intuitive and efficient, and optimize table operation options to improve user experience.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can knowledge graphs (e.g., DBpedia) dynamically construct conditionalized data tables?Category: Data Tool Adoption, Analysis Interfaces, and Information Organization SupportSimilar questionsarrow_forward
- Can knowledge graphs combined with interactive table creation improve data integration and cleaning efficiency?Category: Data Tool Adoption, Analysis Interfaces, and Information Organization SupportSimilar questionsarrow_forward
- How can heterogeneity and sparsity of data be addressed in table operations?Category: Data Tool Adoption, Analysis Interfaces, and Information Organization SupportSimilar questionsarrow_forward
Practical Problems
1- Users struggle to quickly create complete tables from scattered data.Category: Data Tool Adoption, Analysis Interfaces, and Information Organization SupportSimilar questionsarrow_forward
- 75%
Interactive Context-Preserving Color Highlighting for Multiclass Scatterplots
CHI '23· Interactive Data Visualization
- 60%
T-Cal: Understanding Team Conversational Data with Calendar-based Visualization
CHI '18· Interactive Data Visualization +1
- 60%
Du Bois Wrapped Bar Chart: Visualizing Categorical Data with Disproportionate Values
CHI '20· Interactive Data Visualization +1
- 60%
Data-Driven Mark Orientation for Trend Estimation in Scatterplots
CHI '21· Interactive Data Visualization +1
- 60%
reVISit: Looking Under the Hood of Interactive Visualization Studies
CHI '21· Interactive Data Visualization +1
- 60%
ConceptScope: Organizing and Visualizing Knowledge in Documents based on Domain Ontology
CHI '21· Interactive Data Visualization +1
- 60%
Annotating Line Charts for Addressing Deception
CHI '22· Interactive Data Visualization +1
- 60%
CiteRead: Integrating Localized Citations into Scientific Paper Reading
IUI '22· Interactive Data Visualization +1
- 60%
Learning from the Past - Do Historical Data Help to Improve Progress Indicators in Web Surveys?
UIST '20· Interactive Data Visualization +1
- 60%
Threddy: An Interactive System for Personalized Thread-based Exploration and Organization of Scientific Literature
UIST '22· Interactive Data Visualization +1
Based on Jaccard similarity of research subtopics & professions (≥60%)