End User Authoring of Personalized Content Classifiers: Comparing Example Labeling, Rule Writing, and LLM Prompting
Authors
Research Background and Issues
-
Identified Problems or Challenges:
Current tools assume users will create custom classifiers in a single, lengthy session, which does not align with the needs of typical social media users. These users often interact with social media in fragmented, short sessions, lacking sustained attention or motivation. Additionally, existing systems provide insufficient support for iterative improvements. -
Importance of the Issue:
Social media content is growing exponentially, and the generic algorithms provided by platforms fail to meet users' diverse preferences. Supporting users in customizing classification algorithms is crucial for enhancing user experience and content management efficiency, especially in sensitive scenarios involving personalized content moderation. -
Research Motivation and Related Work:
Existing studies primarily focus on highly motivated user groups (e.g., community administrators and high-impact creators), neglecting the needs of ordinary users in lightweight and iterative classifier creation processes. This study aims to fill this gap and explore three main strategies for building personalized classifiers: label annotation, rule creation, and large language model (LLM) prompt design.
Solution
-
Proposed Methods or Solutions:
Comparison of three methods for users to create personalized classifiers:- Label Annotation: Users train supervised machine learning models by marking examples.
- Rule Creation: Users create transparent rules based on keyword matching.
- LLM Prompts: Users write natural language prompts to guide large language models in classification.
-
Innovative Aspects of the Solution:
- Incorporating advanced techniques (e.g., active learning, LLM prompt optimization) into experiments to provide a fair comparison of the three strategies.
- Exploring capabilities for rapid initialization and iterative improvement, closely aligned with typical user behavior.
- Revealing differences in performance across dimensions such as context awareness, transparency, and usability through user experiments.
-
Implementation Steps and Key Technologies:
- Design experimental systems based on the three strategies, ensuring consistent feature sets.
- Recruit 37 participants with no programming experience and evaluate each system through cross-experiments.
- Test each method's support for rapid initialization (performance within 5/10/15 minutes) and ease of iterative improvement.
- Collect quantitative performance metrics (accuracy, precision, recall, F1-score) and qualitative user feedback for comparative analysis.
Research Outcomes
-
Specific Outcomes:
- The prompt strategy achieved the best overall performance (particularly excelling in recall and F1-score).
- User preferences varied by scenario:
- Label annotation was the simplest when preferences were vague and intuitive.
- Rule creation performed better when preferences were clear and closely tied to specific topics.
- Prompt design was more efficient when preferences were broad and easily describable.
- The prompt system led in efficiency during the rapid initialization phase but performed less effectively in supporting iterative improvements.
-
Advantages Compared to Existing Solutions:
- Prompts leverage LLMs to exhibit greater flexibility in expressing broad preferences.
- Rule creation significantly outperformed other methods in transparency and control over specific tasks.
- Label annotation had the lowest cognitive load for non-technical users, making it easier to adopt.
-
Experimental or Evaluation Results:
- Time-based comparison: Prompts approached peak performance within 5 minutes, while rule creation and label annotation showed more stable incremental improvements.
- User experience: The prompt system captured complex preferences but was limited in precision and interpretability.
-
Limitations and Future Directions:
- Users struggled to resolve LLM unpredictability and conceptual alignment issues during iterative prompt refinement, suggesting the need for hybrid methods, such as using label annotation to guide prompt development.
- The diversity of content management scenarios requires more flexible tools, such as high-transparency mechanisms for community managers and optimized rapid iteration workflows for individual users.
- Investigating strategies for managing non-textual content to further expand the scope of personalized content classification.
Conclusion and Insights
This study provides valuable insights into user needs for personalized content classification. By comparing the strengths and weaknesses of different strategies, it highlights the potential of hybrid approaches—such as combining the transparency of rules with the intelligence of LLMs. Additionally, the research emphasizes the importance of considering scenario diversity, user preferences, and technological cost balance when designing content creation tools. These findings offer significant guidance for developing more efficient and user-friendly tools for custom content classification.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can three strategies (label annotation, rule creation, LLM prompting) be designed and compared to support lay users in creating personalized classifiers?Category: Computing and AI Literacy Education SupportSimilar questionsarrow_forward
- How do different strategies perform in rapid initialization, transparency, and iterative improvement support?Category: Computing and AI Literacy Education SupportSimilar questionsarrow_forward
- How do lay users' content management needs influence personalized classifier tool design?Category: Computing and AI Literacy Education SupportSimilar questionsarrow_forward
Practical Problems
1- Lay users struggle to quickly and effectively create personalized social media classifiers.Category: Computing and AI Literacy Education SupportSimilar questionsarrow_forward
- 67%
Surprise Me If You Can: Serendipity in Health Information
CHI '18· Human-LLM Collaboration +2
- 67%
Towards Complete Icon Labeling in Mobile Applications
CHI '22· Human-LLM Collaboration +1
- 67%
To Search or To Gen? Design Dimensions Integrating Web Search and Generative AI in Programmers' Information-Seeking Process
DIS '25· Human-LLM Collaboration +1
- 67%
Interactive Storytelling for Movie Recommendation through Latent Semantic Analysis
IUI '18· Human-LLM Collaboration +2
- 67%
Leveraging ChatGPT for Automated Human-centered Explanations in Recommender Systems
IUI '24· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)