Promptimizer: User-Led Prompt Optimization for Personal Content Classification

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationRecommender System UXContent Creators (YouTubers, Podcasters)Software Engineers & DevelopersHCI Researchers

Paper Title

Promptimizer: User-Led Prompt Optimization for Personal Content Classification

Publication Info

  • Topic area: Human-in-the-loop prompt optimization for personalized content classification using LLMs.
  • Keywords: Prompt optimization, human-in-the-loop, LLMs, content moderation, content curation, interpretability, personalization, active learning, YouTube creators, classifier refinement.

Background and Problem

  • Problem / challenge: Existing automatic prompt optimization techniques fail to incorporate user preferences during refinement and produce interpretable prompts, limiting their effectiveness for personal content classification.
  • Significance: Personalized content classification is crucial for managing diverse user preferences, protecting online communities, and enabling meaningful engagement, especially for social media users and content creators.
  • Motivation and related work: Prior systems rely on rule-based approaches, supervised learning, or LLM-based classifiers, but they place a heavy burden on users for prompt engineering and fail to support evolving preferences. Automatic prompt optimization methods focus on performance but lack mechanisms for user input and interpretability.

Solution

  • Proposed approach: Promptimizer, a human-in-the-loop prompt optimization workflow that integrates user input during initialization and refinement stages to produce performant and interpretable prompts.
  • Novelty:
    1. Structured prompt filters with descriptions, positive/negative rubrics, and few-shot examples for interpretability.
    2. User-led steering mechanisms for diverse input during initialization and refinement.
    3. Active learning to prioritize informative examples for labeling.
    4. Clustering failure patterns to generate targeted refinements.
  • Procedure and key techniques:
    • Initialization: Users provide descriptions/examples and label prioritized comments using active learning.
    • Refinement: Users review clustered failure patterns, fix individual mistakes, or manually edit prompts.
    • Algorithm: Structured prompts are optimized iteratively using constrained edits informed by user feedback.

Results

  • Concrete findings:
    • In a lab study (n=16), participants unanimously preferred Promptimizer over automatic optimization, citing greater interpretability and control.
    • Promptimizer achieved comparable performance to a state-of-the-art APO baseline, with F1 scores improving from 0.710 to 0.756 after refinement.
    • In a field study with 10 YouTube creators, participants authored 41 diverse prompt filters, flagging ~5,700 comments (10.3% of total).
  • Advantage over baselines:
    • Significantly higher interpretability and user satisfaction.
    • Support for nuanced refinements and evolving preferences.
  • Experiments / evaluation:
    • Lab study: Within-subjects comparison of Promptimizer and APO baseline using political and food-related YouTube comments.
    • Field study: Three-week deployment of Puffin among YouTube creators to test real-world applicability.
    • Metrics: Accuracy, precision, recall, F1 score, interpretability ratings, and usability surveys.
  • Limitations and future work:
    • Small sample sizes and constrained computational resources.
    • Limited evaluation of long-term utility and cross-platform applicability.
    • Future directions include extending to open-ended LLM tasks, incorporating contextual factors, and enabling community sharing of classifiers.

Summary

Promptimizer introduces a human-in-the-loop workflow for optimizing LLM-based content classifiers, emphasizing user input and interpretability. It enables users to initialize and refine classifiers through structured prompts, active learning, and targeted refinements. Lab experiments demonstrated unanimous user preference for Promptimizer over automatic optimization, while field studies showed its utility in real-world settings for YouTube creators managing comments. Future research could expand its applicability to broader LLM tasks and collective settings, addressing scalability and contextual challenges.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223337/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790923
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Recommender System UX
work
Professions
Content Creators (YouTubers, Podcasters), Software Engineers & Developers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers