DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
Authors
Paper Title
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
Publication Info
- Topic area: Text labeling frameworks combining data programming, active learning, and large language models.
- Keywords: Data labeling, data programming, active learning, large language models, text classification, labeling functions, span sets, iterative refinement, usability, labeling efficiency.
Background and Problem
- Problem / challenge: Existing text labeling approaches struggle to balance label quality and cost. Data programming requires programming expertise, active learning suffers from cold-start issues, and LLM-based labeling can be inconsistent and unreliable without human oversight.
- Significance: High-quality labeled datasets are critical for deep learning in NLP tasks, and reducing the cost and effort of labeling can accelerate advancements in the field.
- Motivation and related work: Prior work has explored pairwise combinations of data programming, active learning, and LLMs but has not fully integrated all three. Challenges include the programming burden in data programming, inefficiencies in active learning, and the need for human oversight in LLM-based labeling.
Solution
- Proposed approach: DALL, a unified text labeling framework that integrates data programming, active learning, and LLMs to improve labeling efficiency and accuracy.
- Novelty:
- Introduces a structured specification for defining labeling functions via configuration rather than code.
- Combines data programming with active learning to refine noisy labels and address cold-start issues.
- Leverages LLMs to assist in label correction, span set expansion, and labeling function refinement.
- Implements an interactive labeling system with a user-friendly interface for iterative refinement.
- Procedure and key techniques:
- Users define labeling functions using a no-code structured specification.
- Active learning selects informative instances based on uncertainty, disagreement, or abstention.
- LLMs analyze selected instances to recommend labels, expand span sets, and suggest new labeling functions.
- Iterative refinement improves label quality through user interaction with the system.
Results
- Concrete findings:
- DALL achieved comparable or higher accuracy than Doccano, Snorkel, and GPT-3.5 Turbo, with significantly lower labeling time.
- Reusing labeling functions across tasks enabled rapid convergence to high accuracy.
- Ablation studies showed that combining data programming, active learning, and LLMs improved accuracy and reduced time cost.
- Advantage over baselines:
- Faster time to reach 85% accuracy compared to alternatives.
- Sustained accuracy improvements through iterative refinement using active learning and LLM assistance.
- Experiments / evaluation:
- Comparative study with 36 participants on sentiment analysis tasks.
- Ablation study with 18 participants to evaluate the contributions of individual modules.
- Usability study with 15 participants showing high satisfaction and ease of use.
- Limitations and future work:
- Difficulty handling cross-sentence reasoning or implicit expressions.
- Need for domain-specific span sets and prompt optimization.
- Potential to build a library of reusable span sets and automate prompt tuning.
Summary
DALL is a text labeling framework that integrates data programming, active learning, and LLMs to improve labeling efficiency and accuracy. It introduces a structured specification for defining labeling functions without coding and combines it with active learning to prioritize informative instances and LLM analysis to assist in label refinement. Evaluations show that DALL outperforms existing systems in both accuracy and efficiency, with high usability ratings from participants. Future work includes addressing limitations in handling complex expressions, expanding domain-specific resources, and optimizing prompts for new tasks.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
CHI '26· Human-LLM Collaboration +3
- 100%
Live in the Loop: Rapid Run-time Feedback for Prompts
CHI '26· Human-LLM Collaboration +3
- 88%
Interview-Informed Generative Agents for Product Discovery: A Validation Study
CHI '26· Human-LLM Collaboration +3
- 88%
"Shall We Dig Deeper?": Designing and Evaluating Strategies for LLM Agents to Advance Knowledge Co-Construction in Asynchronous Online Discussions
CHI '26· Human-LLM Collaboration +3
- 88%
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
CHI '26· Human-LLM Collaboration +3
- 88%
TurnStyle: A Framework for Analyzing Human Conversational Behaviors to Predict Success in LLM-Assisted Tasks
CHI '26· Human-LLM Collaboration +3
- 75%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 75%
"Here, Let Me Help": An Empirical Study of User Interventions in Human–Web Agent Collaboration
CHI '26· Human-LLM Collaboration +3
- 75%
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
CHI '26· Human-LLM Collaboration +3
- 75%
The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)