Text-to-SQL Domain Adaptation via Human-LLM Collaborative Data Annotation
Authors
Text-to-SQL models, which parse natural language (NL) questions to executable SQL queries, are increasingly adopted in real-world applications. However, deploying such models in the real world often requires adapting them to the highly specialized database schemas used in specific applications. We observe that the performance of existing text-to-SQL models drops dramatically when applied to a new schema, primarily due to the lack of domain-specific data for fine-tuning. Furthermore, this lack of data for the new schema also hinders our ability to effectively evaluate the model's performance in the new domain. Nevertheless, it is expensive to continuously obtain text-to-SQL data for an evolving schema in most real-world applications. To bridge this gap, we propose SQLsynth, a human-in-the-loop text-to-SQL data annotation system. SQLsynth streamlines the creation of high-quality text-to-SQL datasets through collaboration between humans and a large language model in a structured workflow. A within-subject user study comparing SQLsynth to manual annotation and ChatGPT reveals that SQLsynth significantly accelerates text-to-SQL data annotation, reduces cognitive load, and produces datasets that are more accurate, natural, and diverse. Our code is available at https://github.com/adobe/nl_sql_analyzer.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can text-to-SQL model performance degradation on new domain database schemas be addressed?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- Can human-AI collaborative dynamic annotation tools improve text-to-SQL data annotation efficiency and quality?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- How can rule generation methods combined with LLMs generate diverse, high-quality SQL and natural language query data?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
Practical Problems
1- Text-to-SQL models perform poorly in new domains due to insufficient data.Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- 100%
Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences
CHI '24· Human-LLM Collaboration +1
- 100%
InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMs
CHI '25· Human-LLM Collaboration +1
- 100%
Interactive Hyperparameter Optimization with Paintable Timelines
DIS '21· Human-LLM Collaboration +1
- 80%
AI for Low-Code for AI
IUI '24· Generative AI (Text, Image, Music, Video) +2
- 80%
Genie in the Model: Automatic Generation of Human-in-the-Loop Deep Neural Networks for Mobile Applications
UbiComp '23· Human-LLM Collaboration +2
- 67%
Towards Human-Guided Machine Learning
IUI '19· Human-LLM Collaboration +2
- 67%
CoAutoML: User Interface Framework for Machine Learning Novices using LLM-based AutoML and Test-Driven Machine Teaching
IUI '26· AutoML Interfaces +2
- 67%
Never-ending Learning of User Interfaces
UIST '23· Human-LLM Collaboration +2
- 60%
Trade-offs for Substituting a Human with an Agent in a Pair Programming Context: The Good, the Bad, and the Ugly
CHI '21· Human-LLM Collaboration +1
- 60%
Visualizing Examples of Deep Neural Networks at Scale
CHI '21· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)