"Un-default" Behavior Tuning: Specifying Model Behavior outside the Norm with LLM Self-Playing and Self-Improving
Authors
Specifying model behavior is challenging—especially when the desired behavior is unpopular relative to the model’s training data. Reversing the influence of massive training corpora is both time-consuming and costly, and such interventions are typically inaccessible to end users. While Large Language Models (LLMs) make it easier to write instructions using natural language, specifying unpopular behaviors remains a difficult task. We introduce \Undefault{}, a human-in-the-loop framework that combines self-play with self-refinement to better specify such behaviors. Our system enables users to identify popular (but undesired) model behaviors through self-play, then iteratively guide the model toward preferred alternatives by refining prompts in a self-improving loop. Our first evaluation involves user study conducted on a system implementation of \Undefault{} within the context of chatbot behavior. Our system self-play itself by simulating user interactions to identify patterns and create effective prompts based on the pattern. In a within-subject study (N=12), participants pinpointed more patterns through self-playing and crafted better prompts. Surprisingly, users felt more or less success level in specifying the model behavior. Follow-up crowd studies (N=60) confirmed that the chatbot adhered to instructions without sacrificing quality. Our second evaluation is a case study on a real-world implementation using a movie rating dataset with \Undefault{}, demonstrating its effectiveness and robustness in modeling a critic's preferences across the spectrum of low to highly rated movies. Together, these results suggest how AI improves the design process of interactive AI systems. Furthermore, they suggest how the benefits of these tools may be non-obvious to end-users. We reflect on these findings and suggest future directions.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
- 100%
Belief Updating and Delegation in Multi-Task Human–AI Interaction: Evidence from Controlled Simulations
CHI '26· Human-LLM Collaboration +2
- 100%
Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
IUI '26· Human-LLM Collaboration +2
- 100%
Making Absence Visible in Intelligent Summarization Interfaces
IUI '26· Human-LLM Collaboration +2
- 83%
Which Contributions Deserve Credit? Perceptions of Attribution in Human-AI Co-Creation
CHI '25· Human-LLM Collaboration +2
- 83%
Co-Disclosing the Computer: LLM-Mediated Computing through Reflective Conversation
CHI '26· Human-LLM Collaboration +2
- 83%
What can AI do for me: Evaluating Machine Learning Interpretations in Cooperative Play
IUI '19· Human-LLM Collaboration +2
- 83%
CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent Agents
IUI '26· Human-LLM Collaboration +2
- 83%
User Reliance on AI Support for Collaborative Partner Selection
IUI '26· Human-LLM Collaboration +2
- 83%
EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences
UIST '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)