Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
Authors
Human-LLM CollaborationAI-Assisted Decision-Making & AutomationSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers
Title of the Paper
Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
Paper Information
- Domain: Human-Computer Interaction, Artificial Intelligence, Large Language Model (LLM) Application Design
- Keywords: Large Language Models, Prompt Engineering, Non-Expert Users, Design Tools, Human-Computer Interaction, User Research, Education and Training, End-User Programming, Interactive Machine Learning, Design Challenges
Research Background and Problem
- Problem Identification and Challenges: Pre-trained large language models (LLMs) like GPT-3 and ChatGPT offer new opportunities for user interaction, aiming to optimize model outputs through "prompt engineering." However, for non-AI experts, designing effective prompts is often challenging, with issues such as design fragility, unpredictable outcomes, and a lack of systematic approaches.
- Significance: As LLM technology becomes widely adopted, prompt engineering lowers the programming barrier, enabling non-expert users to engage in design activities. This creates new opportunities for user innovation but also exposes cognitive and practical barriers in the design process.
- Research Motivation and Related Work: The new "pre-train-prompt-predict" paradigm in prompt engineering significantly reduces the technical threshold for NLP development. However, current tools primarily serve programmers and NLP experts, lacking support for ordinary users. Existing studies have highlighted issues such as users relying on trial-and-error, lacking guidance, and having unrealistic expectations of LLM capabilities, but systematic research remains limited.
Proposed Solution
- Method or Solution:
- To study users' intuitive design behaviors, a no-code LLM prompt design tool called BotDesigner was created, enabling users to design and iterate chatbot prompts using natural language.
- Observed the behaviors and challenges of non-AI experts as they completed specific tasks using the tool.
- Innovations:
- Introduced an open design tool that allows users to iteratively generate chatbots using plain text prompts.
- The tool supports error tagging, comparing multi-turn dialogue outputs, and iterative improvements.
- Adopted a design probe approach for user research, deeply uncovering barriers in intuitive prompt engineering.
- Implementation Steps and Key Techniques:
- Tool Design: BotDesigner provides a conversation view and an error browser. Users can iteratively design chatbots by modifying prompt templates and systematically evaluate the effectiveness of different templates.
- Research Design: Recruited 10 non-AI experts for user behavior experiments, with tasks including testing existing prompt templates and optimizing them to improve the chatbot's performance in mimicking cooking tutorials.
- Analysis Methods: Combined think-aloud protocols with error tagging and interaction analysis to uncover potential design thinking models and barriers.
Research Findings
- Key Findings:
- Users predominantly adopted opportunistic design strategies when tackling prompt design tasks, focusing on localized problem fixes rather than employing systematic global testing strategies.
- Misconceptions based on interaction or human social experience (e.g., assuming the chatbot can "understand" instructions) posed barriers to iterative design processes.
- Users preferred expressing design intentions through direct instructions (e.g., "tell a joke") rather than example-based samples, despite the latter being proven more effective.
- Repetitive trial-and-error, prompt "structuring," and systematic testing were rare, highlighting an urgent need for prompt education and process guidance.
- Advantages Over Existing Solutions:
- Provided a framework supporting systematic testing of prompt designs, emphasizing error tagging and iterative testing at both local and global levels.
- Developed an open-ended tool for non-experts, effectively uncovering challenges faced by non-professional users in technology adoption.
- Experimental or Evaluation Results:
- Most participants successfully completed specific chatbot optimization tasks but struggled to achieve comprehensive model-based designs.
- Users' inherent tendencies toward overgeneralization and reliance on interpersonal interaction logic had a profound impact on their design behaviors.
- Features like error tagging in the tool were underutilized in practice.
- Limitations and Future Directions:
- The study participants were primarily highly educated professionals, limiting generalizability to the broader public. Future research should include more diverse samples.
- The current study focused on chatbot design; future work could extend to non-chat scenarios (e.g., search tools, automatic summarization).
- Develop educational and training modules to enhance user intuition in prompt engineering and address disparities in design experience.
- Suggest exploring the impact of preferences for "seamless design" vs. "non-seamless design" approaches on user behavior.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- What specific challenges do non-AI experts face when designing LLM prompts?Category: LLM Prompt Engineering and Authoring ToolsSimilar questionsarrow_forward
- Without programming background, how can ordinary users iteratively optimize prompts through tools?Category: LLM Prompt Engineering and Authoring ToolsSimilar questionsarrow_forward
- What are non-expert users' thinking patterns and behavioral preferences in prompt design?Category: LLM Prompt Engineering and Authoring ToolsSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Ordinary users struggle to design effective prompts and optimize LLM output.Category: LLM Prompt Engineering and Authoring ToolsSimilar questionsarrow_forward
- 100%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 100%
Automatic Macro Mining from Interaction Traces at Scale
CHI '24· Human-LLM Collaboration +1
- 83%
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
CHI '24· Human-LLM Collaboration +2
- 83%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 83%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 83%
Vibe Coding Entanglements – Repositioning Boundaries of Intention, Authorship, and Responsibility in Programming with Generative AI
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 83%
The CoExplorer Technology Probe: A Generative AI-Powered Adaptive Interface to Support Intentionality in Planning and Running Video Meetings
DIS '24· Human-LLM Collaboration +2
- 83%
Designing with Multi-Agent Generative AI: Insights from Industry Early Adopters
DIS '25· Human-LLM Collaboration +2
- 83%
Knowledge Graph Completion-based Question Selection for Acquiring Domain Knowledge through Dialogues
IUI '21· Conversational Chatbots +2
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581388
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers