Promptimizer: User-Led Prompt Optimization for Personal Content Classification
Authors
Paper Title
Promptimizer: User-Led Prompt Optimization for Personal Content Classification
Publication Info
- Topic area: Human-in-the-loop prompt optimization for personalized content classification using LLMs.
- Keywords: Prompt optimization, human-in-the-loop, LLMs, content moderation, content curation, interpretability, personalization, active learning, YouTube creators, classifier refinement.
Background and Problem
- Problem / challenge: Existing automatic prompt optimization techniques fail to incorporate user preferences during refinement and produce interpretable prompts, limiting their effectiveness for personal content classification.
- Significance: Personalized content classification is crucial for managing diverse user preferences, protecting online communities, and enabling meaningful engagement, especially for social media users and content creators.
- Motivation and related work: Prior systems rely on rule-based approaches, supervised learning, or LLM-based classifiers, but they place a heavy burden on users for prompt engineering and fail to support evolving preferences. Automatic prompt optimization methods focus on performance but lack mechanisms for user input and interpretability.
Solution
- Proposed approach: Promptimizer, a human-in-the-loop prompt optimization workflow that integrates user input during initialization and refinement stages to produce performant and interpretable prompts.
- Novelty:
- Structured prompt filters with descriptions, positive/negative rubrics, and few-shot examples for interpretability.
- User-led steering mechanisms for diverse input during initialization and refinement.
- Active learning to prioritize informative examples for labeling.
- Clustering failure patterns to generate targeted refinements.
- Procedure and key techniques:
- Initialization: Users provide descriptions/examples and label prioritized comments using active learning.
- Refinement: Users review clustered failure patterns, fix individual mistakes, or manually edit prompts.
- Algorithm: Structured prompts are optimized iteratively using constrained edits informed by user feedback.
Results
- Concrete findings:
- In a lab study (n=16), participants unanimously preferred Promptimizer over automatic optimization, citing greater interpretability and control.
- Promptimizer achieved comparable performance to a state-of-the-art APO baseline, with F1 scores improving from 0.710 to 0.756 after refinement.
- In a field study with 10 YouTube creators, participants authored 41 diverse prompt filters, flagging ~5,700 comments (10.3% of total).
- Advantage over baselines:
- Significantly higher interpretability and user satisfaction.
- Support for nuanced refinements and evolving preferences.
- Experiments / evaluation:
- Lab study: Within-subjects comparison of Promptimizer and APO baseline using political and food-related YouTube comments.
- Field study: Three-week deployment of Puffin among YouTube creators to test real-world applicability.
- Metrics: Accuracy, precision, recall, F1 score, interpretability ratings, and usability surveys.
- Limitations and future work:
- Small sample sizes and constrained computational resources.
- Limited evaluation of long-term utility and cross-platform applicability.
- Future directions include extending to open-ended LLM tasks, incorporating contextual factors, and enabling community sharing of classifiers.
Summary
Promptimizer introduces a human-in-the-loop workflow for optimizing LLM-based content classifiers, emphasizing user input and interpretability. It enables users to initialize and refine classifiers through structured prompts, active learning, and targeted refinements. Lab experiments demonstrated unanimous user preference for Promptimizer over automatic optimization, while field studies showed its utility in real-world settings for YouTube creators managing comments. Future research could expand its applicability to broader LLM tasks and collective settings, addressing scalability and contextual challenges.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)