SemTabla: A Human-in-the-Loop Framework for Semantic Enrichment and Validation of Data Tables

Honorable Mention
Explainable AI (XAI)Interactive Data VisualizationUser Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingData Scientists & AnalystsAI/ML Researchers & EngineersHCI Researchers

Paper Title

SemTabla: A Human-in-the-Loop Framework for Semantic Enrichment and Validation of Data Tables

Publication Info

  • Topic area: Semantic enrichment and validation of tabular data for enhanced reasoning and usability.
  • Keywords: Semantic enrichment, human-in-the-loop, data tables, Table QA, large language models, functional dependencies, primary keys, foreign keys, interactive systems.

Background and Problem

  • Problem / challenge: Metadata from table schemas often fails to capture the full business semantics of tabular data, leading to reasoning errors in Table QA systems. Existing automated approaches struggle with insufficient data utilization, narrow feature coverage, and limited interpretability.
  • Significance: Accurate semantic understanding of tables is critical for improving the reasoning capabilities of Table QA systems and enabling effective decision-making in domains like finance, healthcare, and e-commerce.
  • Motivation and related work: Prior work on semantic table understanding has advanced techniques like column type identification, entity linking, and relation extraction but lacks support for cross-table semantic associations, adaptability to unstructured tables, and integrated end-to-end solutions. This paper addresses these gaps.

Solution

  • Proposed approach: SemTabla, an interactive system employing a human-in-the-loop mechanism to extract, validate, and refine semantic information from data tables.
  • Novelty:
    1. A hierarchical framework for extracting semantic attributes at multiple levels (field, table, and cross-table).
    2. A novel sampling method for identifying critical but rare row instances.
    3. An interactive interface for visualization, validation, and refinement of extracted semantics.
    4. Evaluation of the system's usability and its impact on Table QA performance.
  • Procedure and key techniques:
    • Semantic Enrichment: A four-step process to extract semantic features, including field-level attributes, table relationships, column dependencies, and table-level semantic labels.
    • Sampling Strategy: Iterative sampling and counterexample validation to ensure comprehensive and efficient detection.
    • Semantic Validation: Interactive modules for validating extracted features using positive and negative examples, supported by SQL queries.
    • Interactive Interface: Modular views for uploading datasets, visualizing semantic features, and refining results.

Results

  • Concrete findings:
    • Execution accuracy in Table QA improved from 36.25% to 45.18% (Qwen3-Plus) and from 43.81% to 51.83% (DeepSeek-V3) with semantic enrichment.
    • Semantic enrichment particularly enhanced performance on multi-table (+11.76%), time-related (+7.86%), and aggregation (+8.76%) queries.
    • Detection tasks like primary key identification and functional dependency discovery maintained low latency (milliseconds to seconds) even for large datasets (up to 1.6M rows).
  • Advantage over baselines:
    • Outperformed baseline systems in usability, learning cost, mental effort, result accuracy, and efficiency in a user study with 18 participants.
    • Provided interactive validation and evidence-based feedback, unlike baseline systems that lacked verification mechanisms.
  • Experiments / evaluation:
    • User study comparing SemTabla with two baseline systems on datasets from the Bird benchmark.
    • Ablation study to assess the impact of semantic enrichment on Table QA performance.
    • Performance evaluation of detection tasks on large-scale datasets.
  • Limitations and future work:
    • Scalability: Current implementation is limited to local datasets; future work could integrate distributed databases for large-scale scenarios.
    • User scenarios: Less applicable to well-documented datasets or non-technical users.
    • SQL analysis: Incorporating SQL scripts for deeper semantic understanding is a potential direction.

Summary

SemTabla is an interactive system designed to address the limitations of existing semantic enrichment methods for data tables. It combines automated semantic extraction with human-in-the-loop validation, leveraging a hierarchical framework, novel sampling strategies, and an intuitive interface. The system significantly improves Table QA performance by enriching prompts with semantic information, as demonstrated in experiments with large language models. User studies highlight its usability and efficiency, though future work is needed to enhance scalability and support for diverse user scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222343/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791449
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
6 authors
sell
Subtopics
Explainable AI (XAI), Interactive Data Visualization, User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers