Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent

Explainable AI (XAI)AutoML InterfacesComputational Methods in HCISoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Research Background and Problem

  • What issues or challenges did the authors identify?
    The authors explored how users design prompts for large language models (LLMs) to iteratively optimize their performance for data annotation tasks in the absence of gold-standard data. Previous studies have shown that gold-standard data is critical for evaluating annotation quality and system performance. However, in many real-world scenarios, gold-standard data may be unavailable, expensive, or difficult to use.

  • Why is this issue important?
    As LLMs are increasingly applied to various tasks, understanding how users can effectively utilize these models without gold-standard data is crucial for ensuring their reliability and applicability. Particularly in dynamic data annotation scenarios, enhancing users' prompt design capabilities can significantly improve annotation quality.

  • Research Motivation and Related Work
    Through a review of related literature, the authors found that gold-standard data is essential for efficient annotation, but many current methods assume its availability. Existing tools primarily rely on gold-standard data to optimize prompts, while there is limited research on scenarios where such data is absent. This study aims to delve into the "prompting in the dark" scenario, evaluating users' ability to iterate on prompts without gold-standard data.


Solution

  • What methods or solutions did the authors propose?
    The authors developed a tool called "PromptingSheet," a Google Sheets-based plugin that allows users to iteratively design prompts for data annotation tasks. Key features of the tool include task background setup, rule definition, and multiple example annotations.

  • What are the innovative aspects of this solution?

    1. It provides a user-friendly interface based on spreadsheets, making it accessible to general users rather than technical experts.
    2. It emphasizes "evolving understanding during the prompt design process," allowing users to adjust annotation rules and task goals as they observe LLM-predicted annotation results.
    3. It enables users to design prompts from scratch without requiring pre-annotated data.
  • What are the implementation steps and key technologies used?

    1. Users load data into the tool and write task background information (e.g., data source, annotation objectives).
    2. Users define rules and annotation examples for various labels using the platform's "rulebook" and "example library."
    3. The tool supports LLM-based data annotation and generates annotation explanations, enabling users to iteratively refine rules or examples to improve accuracy.
    4. Users can optimize prompts through multiple iterations until satisfactory annotation results are achieved. The tool also provides a dashboard to track task progress and monitor prompt content.

Research Outcomes

  • What specific outcomes were achieved?
    Through user testing with 20 participants, the study found that the "prompting in the dark" approach was highly unreliable. Only 9 participants improved annotation accuracy after four or more iterations. Most users faced issues with unclear rules, and the effectiveness of prompt adjustments was generally limited.

  • What advantages does it have compared to existing solutions?
    The tool simplifies the complex operations involved in prompt optimization, making it suitable for general users. However, its performance in the absence of gold-standard data did not meet expectations, falling short of the reliability offered by traditional platforms that rely on gold-standard data.

  • What were the experimental or evaluation results?

    1. The average annotation accuracy increased slightly from an initial 0.542 to 0.553 after four iterations, but the improvement was not significant and showed variability.
    2. Most participants appreciated the tool's ease of use and task completion efficiency but expressed doubts about the effectiveness of prompt improvements.
    3. Automated prompt optimization tools (e.g., DSPy) also struggled to provide significant improvements, as the small number of generated gold-standard examples was insufficient to support algorithmic optimization.
  • Limitations and Future Directions
    Limitations:

    • The "prompting in the dark" scenario heavily depends on users' individual understanding, and a lack of guidance may lead to inefficient iterations.
    • Automated tools perform poorly with limited examples or complex tasks.
    • The study was conducted in a single user research session, lacking support from long-term behavioral data.

    Future Directions:

    • Introduce initial gold-standard data to provide foundational guidance while minimizing its impact on users' exploratory freedom.
    • Incorporate more automated support features, such as methods for auto-generating rules.
    • Develop more robust prompt optimization algorithms to handle scenarios with limited examples.
    • Deploy the tool in real-world user scenarios to observe long-term impacts and usage behaviors.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188592/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714319
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), AutoML Interfaces, Computational Methods in HCI
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers