Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts

Paper Information

  • Domain: Human-Computer Interaction, Artificial Intelligence, Large Language Model (LLM) Application Design
  • Keywords: Large Language Models, Prompt Engineering, Non-Expert Users, Design Tools, Human-Computer Interaction, User Research, Education and Training, End-User Programming, Interactive Machine Learning, Design Challenges

Research Background and Problem

  • Problem Identification and Challenges: Pre-trained large language models (LLMs) like GPT-3 and ChatGPT offer new opportunities for user interaction, aiming to optimize model outputs through "prompt engineering." However, for non-AI experts, designing effective prompts is often challenging, with issues such as design fragility, unpredictable outcomes, and a lack of systematic approaches.
  • Significance: As LLM technology becomes widely adopted, prompt engineering lowers the programming barrier, enabling non-expert users to engage in design activities. This creates new opportunities for user innovation but also exposes cognitive and practical barriers in the design process.
  • Research Motivation and Related Work: The new "pre-train-prompt-predict" paradigm in prompt engineering significantly reduces the technical threshold for NLP development. However, current tools primarily serve programmers and NLP experts, lacking support for ordinary users. Existing studies have highlighted issues such as users relying on trial-and-error, lacking guidance, and having unrealistic expectations of LLM capabilities, but systematic research remains limited.

Proposed Solution

  • Method or Solution:
    • To study users' intuitive design behaviors, a no-code LLM prompt design tool called BotDesigner was created, enabling users to design and iterate chatbot prompts using natural language.
    • Observed the behaviors and challenges of non-AI experts as they completed specific tasks using the tool.
  • Innovations:
    • Introduced an open design tool that allows users to iteratively generate chatbots using plain text prompts.
    • The tool supports error tagging, comparing multi-turn dialogue outputs, and iterative improvements.
    • Adopted a design probe approach for user research, deeply uncovering barriers in intuitive prompt engineering.
  • Implementation Steps and Key Techniques:
    • Tool Design: BotDesigner provides a conversation view and an error browser. Users can iteratively design chatbots by modifying prompt templates and systematically evaluate the effectiveness of different templates.
    • Research Design: Recruited 10 non-AI experts for user behavior experiments, with tasks including testing existing prompt templates and optimizing them to improve the chatbot's performance in mimicking cooking tutorials.
    • Analysis Methods: Combined think-aloud protocols with error tagging and interaction analysis to uncover potential design thinking models and barriers.

Research Findings

  • Key Findings:
    • Users predominantly adopted opportunistic design strategies when tackling prompt design tasks, focusing on localized problem fixes rather than employing systematic global testing strategies.
    • Misconceptions based on interaction or human social experience (e.g., assuming the chatbot can "understand" instructions) posed barriers to iterative design processes.
    • Users preferred expressing design intentions through direct instructions (e.g., "tell a joke") rather than example-based samples, despite the latter being proven more effective.
    • Repetitive trial-and-error, prompt "structuring," and systematic testing were rare, highlighting an urgent need for prompt education and process guidance.
  • Advantages Over Existing Solutions:
    • Provided a framework supporting systematic testing of prompt designs, emphasizing error tagging and iterative testing at both local and global levels.
    • Developed an open-ended tool for non-experts, effectively uncovering challenges faced by non-professional users in technology adoption.
  • Experimental or Evaluation Results:
    • Most participants successfully completed specific chatbot optimization tasks but struggled to achieve comprehensive model-based designs.
    • Users' inherent tendencies toward overgeneralization and reliance on interpersonal interaction logic had a profound impact on their design behaviors.
    • Features like error tagging in the tool were underutilized in practice.
  • Limitations and Future Directions:
    • The study participants were primarily highly educated professionals, limiting generalizability to the broader public. Future research should include more diverse samples.
    • The current study focused on chatbot design; future work could extend to non-chat scenarios (e.g., search tools, automatic summarization).
    • Develop educational and training modules to enhance user intuition in prompt engineering and address disparities in design experience.
    • Suggest exploring the impact of preferences for "seamless design" vs. "non-seamless design" approaches on user behavior.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96495/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581388
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers