SemTabla: A Human-in-the-Loop Framework for Semantic Enrichment and Validation of Data Tables
Honorable MentionAuthors
Paper Title
SemTabla: A Human-in-the-Loop Framework for Semantic Enrichment and Validation of Data Tables
Publication Info
- Topic area: Semantic enrichment and validation of tabular data for enhanced reasoning and usability.
- Keywords: Semantic enrichment, human-in-the-loop, data tables, Table QA, large language models, functional dependencies, primary keys, foreign keys, interactive systems.
Background and Problem
- Problem / challenge: Metadata from table schemas often fails to capture the full business semantics of tabular data, leading to reasoning errors in Table QA systems. Existing automated approaches struggle with insufficient data utilization, narrow feature coverage, and limited interpretability.
- Significance: Accurate semantic understanding of tables is critical for improving the reasoning capabilities of Table QA systems and enabling effective decision-making in domains like finance, healthcare, and e-commerce.
- Motivation and related work: Prior work on semantic table understanding has advanced techniques like column type identification, entity linking, and relation extraction but lacks support for cross-table semantic associations, adaptability to unstructured tables, and integrated end-to-end solutions. This paper addresses these gaps.
Solution
- Proposed approach: SemTabla, an interactive system employing a human-in-the-loop mechanism to extract, validate, and refine semantic information from data tables.
- Novelty:
- A hierarchical framework for extracting semantic attributes at multiple levels (field, table, and cross-table).
- A novel sampling method for identifying critical but rare row instances.
- An interactive interface for visualization, validation, and refinement of extracted semantics.
- Evaluation of the system's usability and its impact on Table QA performance.
- Procedure and key techniques:
- Semantic Enrichment: A four-step process to extract semantic features, including field-level attributes, table relationships, column dependencies, and table-level semantic labels.
- Sampling Strategy: Iterative sampling and counterexample validation to ensure comprehensive and efficient detection.
- Semantic Validation: Interactive modules for validating extracted features using positive and negative examples, supported by SQL queries.
- Interactive Interface: Modular views for uploading datasets, visualizing semantic features, and refining results.
Results
- Concrete findings:
- Execution accuracy in Table QA improved from 36.25% to 45.18% (Qwen3-Plus) and from 43.81% to 51.83% (DeepSeek-V3) with semantic enrichment.
- Semantic enrichment particularly enhanced performance on multi-table (+11.76%), time-related (+7.86%), and aggregation (+8.76%) queries.
- Detection tasks like primary key identification and functional dependency discovery maintained low latency (milliseconds to seconds) even for large datasets (up to 1.6M rows).
- Advantage over baselines:
- Outperformed baseline systems in usability, learning cost, mental effort, result accuracy, and efficiency in a user study with 18 participants.
- Provided interactive validation and evidence-based feedback, unlike baseline systems that lacked verification mechanisms.
- Experiments / evaluation:
- User study comparing SemTabla with two baseline systems on datasets from the Bird benchmark.
- Ablation study to assess the impact of semantic enrichment on Table QA performance.
- Performance evaluation of detection tasks on large-scale datasets.
- Limitations and future work:
- Scalability: Current implementation is limited to local datasets; future work could integrate distributed databases for large-scale scenarios.
- User scenarios: Less applicable to well-documented datasets or non-technical users.
- SQL analysis: Incorporating SQL scripts for deeper semantic understanding is a potential direction.
Summary
SemTabla is an interactive system designed to address the limitations of existing semantic enrichment methods for data tables. It combines automated semantic extraction with human-in-the-loop validation, leveraging a hierarchical framework, novel sampling strategies, and an intuitive interface. The system significantly improves Table QA performance by enriching prompts with semantic information, as demonstrated in experiments with large language models. User studies highlight its usability and efficiency, though future work is needed to enhance scalability and support for diverse user scenarios.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 75%
Evaluation-First Design for Data Visualization Interfaces
CHI '26· Interactive Data Visualization +3
- 75%
Evalet: Evaluating Large Language Models through Functional Fragmentation
CHI '26· Human-LLM Collaboration +3
- 63%
Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models
CHI '19· Explainable AI (XAI) +2
- 63%
Perceptual Pat: A Virtual Human Visual System for Iterative Visualization Design
CHI '23· Interactive Data Visualization +2
- 63%
TSEditor: Interactive Time Series Editing for Privacy Preservation
CHI '26· Privacy Perception & Decision-Making +2
- 63%
"I Need to Find That One Chart": How Data Workers Navigate, Summarize and Communicate Analytical Conversations
CHI '26· User Research Methods (Interviews, Surveys, Observation) +2
- 63%
From Narrative to Numbers: Evaluating Survey Questionnaires with Large Language Models
IUI '26· Human-LLM Collaboration +2
- 63%
Rationalizer: Leveraging LLM to Support User Providing the Rationales Behind the Rating of Likert Scale Questionnaires
IUI '26· Human-LLM Collaboration +2
- 63%
OntoScope: Using a Divergent-Convergent Interaction Framework to Support LLM-based Ontology Scoping
IUI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)