Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
Paper Title
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
Publication Info
- Topic area: Interactive workflows for refining Large Language Models (LLMs) through co-evolution of prompts and test data.
- Keywords: Large Language Models, prompt engineering, test set evolution, human-in-the-loop, policy refinement, interactive systems, content moderation, iterative workflows, AI alignment, data synthesis.
Background and Problem
- Problem / challenge: Traditional workflows treat test data and prompt instructions as separate, static artifacts, leading to inefficiencies in refining LLM behavior. Current practices rely on ad-hoc prompt tinkering or episodic testing, which fail to systematically address edge cases or nuanced application-specific policies.
- Significance: Addressing this gap is critical for creating robust, policy-aware LLM applications, especially as LLMs are increasingly deployed in diverse domains where nuanced, context-specific behavior is required.
- Motivation and related work: Prior work has explored prompt optimization, interactive machine learning, and red teaming, but these approaches often lack integration between test set growth and prompt refinement. This paper builds on these foundations to propose a unified, iterative workflow.
Solution
- Proposed approach: Data-prompt co-evolution, an iterative workflow where test data and prompt instructions evolve together to systematically refine LLM behavior.
- Novelty:
- A structured workflow that integrates test-set growth and prompt refinement as interdependent processes.
- An interactive system that supports edge case discovery, rationale elicitation, neighborhood probing, and regression testing.
- Empirical validation showing improved alignment between user intentions and model behavior compared to a baseline.
- Procedure and key techniques:
- Discovering Failures: The system generates challenging inputs to expose prompt weaknesses.
- Articulating Rationales: Users label failures and provide rationales, supported by AI-generated suggestions.
- Neighborhood Probing: Similar examples are synthesized to test the generalizability of rationales.
- Refining the Specification: Prompt instructions are revised based on rationales and test results.
- Evaluating Against the Test Set: Revised prompts are tested on an incrementally growing test set, with automated and human evaluations.
Results
- Concrete findings:
- Co-Evolution workflow led to longer, more detailed prompt instructions (avg. 111.1 words vs. 81.2 in baseline).
- Test sets created with Co-Evolution were larger (avg. 14.19 examples vs. 7.94 in baseline) and more effective at exposing model failures (51.9% failure rate vs. 26.7% in baseline).
- Alignment between model outputs and user intentions improved significantly (F1-score: 0.69 vs. 0.56 in baseline).
- Advantage over baselines:
- Higher specificity in prompt instructions, including more explicit exceptions and concrete examples.
- More consistent and meaningful changes in model behavior.
- Enhanced user satisfaction and confidence in the system’s ability to align with their policies.
- Experiments / evaluation:
- Simulation-based experiments with deterministic personas to test failure discovery and instruction refinement.
- Controlled user study (N=16) comparing Co-Evolution against a baseline prompt-editing interface in two content moderation domains.
- Metrics included prompt length, test set size, alignment (F1-score), and subjective user satisfaction.
- Limitations and future work:
- Binary correct/incorrect labeling limits handling of ambiguous cases.
- Generated examples lacked diversity, requiring improved mechanisms for semantic variety.
- Scaling challenges with large, living test sets and collaborative team workflows.
- Potential over-reliance on LLM-as-judge, necessitating safeguards against automation bias.
- Extending the workflow to open-ended generative tasks and model parameter updates.
Summary
This paper introduces the concept of data-prompt co-evolution, where test data and prompt instructions are iteratively refined together to improve LLM behavior. The proposed interactive system enables users to identify edge cases, articulate rationales, and systematically refine prompts, resulting in richer test sets and better alignment between user intentions and model outputs. Empirical results from user studies and simulations demonstrate significant advantages over traditional prompt-editing workflows. Future directions include addressing ambiguity in labeling, enhancing example diversity, and scaling the workflow to collaborative and open-ended tasks. This approach advances the design of responsible, human-centered AI systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 100%
Live in the Loop: Rapid Run-time Feedback for Prompts
CHI '26· Human-LLM Collaboration +3
- 88%
Interview-Informed Generative Agents for Product Discovery: A Validation Study
CHI '26· Human-LLM Collaboration +3
- 88%
"Shall We Dig Deeper?": Designing and Evaluating Strategies for LLM Agents to Advance Knowledge Co-Construction in Asynchronous Online Discussions
CHI '26· Human-LLM Collaboration +3
- 88%
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
CHI '26· Human-LLM Collaboration +3
- 88%
TurnStyle: A Framework for Analyzing Human Conversational Behaviors to Predict Success in LLM-Assisted Tasks
CHI '26· Human-LLM Collaboration +3
- 75%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 75%
"Here, Let Me Help": An Empirical Study of User Interventions in Human–Web Agent Collaboration
CHI '26· Human-LLM Collaboration +3
- 75%
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
CHI '26· Human-LLM Collaboration +3
- 75%
The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)