Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationUser Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Paper Title

Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior

Publication Info

  • Topic area: Interactive workflows for refining Large Language Models (LLMs) through co-evolution of prompts and test data.
  • Keywords: Large Language Models, prompt engineering, test set evolution, human-in-the-loop, policy refinement, interactive systems, content moderation, iterative workflows, AI alignment, data synthesis.

Background and Problem

  • Problem / challenge: Traditional workflows treat test data and prompt instructions as separate, static artifacts, leading to inefficiencies in refining LLM behavior. Current practices rely on ad-hoc prompt tinkering or episodic testing, which fail to systematically address edge cases or nuanced application-specific policies.
  • Significance: Addressing this gap is critical for creating robust, policy-aware LLM applications, especially as LLMs are increasingly deployed in diverse domains where nuanced, context-specific behavior is required.
  • Motivation and related work: Prior work has explored prompt optimization, interactive machine learning, and red teaming, but these approaches often lack integration between test set growth and prompt refinement. This paper builds on these foundations to propose a unified, iterative workflow.

Solution

  • Proposed approach: Data-prompt co-evolution, an iterative workflow where test data and prompt instructions evolve together to systematically refine LLM behavior.
  • Novelty:
    1. A structured workflow that integrates test-set growth and prompt refinement as interdependent processes.
    2. An interactive system that supports edge case discovery, rationale elicitation, neighborhood probing, and regression testing.
    3. Empirical validation showing improved alignment between user intentions and model behavior compared to a baseline.
  • Procedure and key techniques:
    1. Discovering Failures: The system generates challenging inputs to expose prompt weaknesses.
    2. Articulating Rationales: Users label failures and provide rationales, supported by AI-generated suggestions.
    3. Neighborhood Probing: Similar examples are synthesized to test the generalizability of rationales.
    4. Refining the Specification: Prompt instructions are revised based on rationales and test results.
    5. Evaluating Against the Test Set: Revised prompts are tested on an incrementally growing test set, with automated and human evaluations.

Results

  • Concrete findings:
    • Co-Evolution workflow led to longer, more detailed prompt instructions (avg. 111.1 words vs. 81.2 in baseline).
    • Test sets created with Co-Evolution were larger (avg. 14.19 examples vs. 7.94 in baseline) and more effective at exposing model failures (51.9% failure rate vs. 26.7% in baseline).
    • Alignment between model outputs and user intentions improved significantly (F1-score: 0.69 vs. 0.56 in baseline).
  • Advantage over baselines:
    • Higher specificity in prompt instructions, including more explicit exceptions and concrete examples.
    • More consistent and meaningful changes in model behavior.
    • Enhanced user satisfaction and confidence in the system’s ability to align with their policies.
  • Experiments / evaluation:
    • Simulation-based experiments with deterministic personas to test failure discovery and instruction refinement.
    • Controlled user study (N=16) comparing Co-Evolution against a baseline prompt-editing interface in two content moderation domains.
    • Metrics included prompt length, test set size, alignment (F1-score), and subjective user satisfaction.
  • Limitations and future work:
    • Binary correct/incorrect labeling limits handling of ambiguous cases.
    • Generated examples lacked diversity, requiring improved mechanisms for semantic variety.
    • Scaling challenges with large, living test sets and collaborative team workflows.
    • Potential over-reliance on LLM-as-judge, necessitating safeguards against automation bias.
    • Extending the workflow to open-ended generative tasks and model parameter updates.

Summary

This paper introduces the concept of data-prompt co-evolution, where test data and prompt instructions are iteratively refined together to improve LLM behavior. The proposed interactive system enables users to identify edge cases, articulate rationales, and systematically refine prompts, resulting in richer test sets and better alignment between user intentions and model outputs. Empirical results from user studies and simulations demonstrate significant advantages over traditional prompt-editing workflows. Future directions include addressing ambiguity in labeling, enhancing example diversity, and scaling the workflow to collaborative and open-ended tasks. This approach advances the design of responsible, human-centered AI systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223228/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791222
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, User Research Methods (Interviews, Surveys, Observation), Prototyping & User Testing
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers