Herding AI Cats: Lessons from Designing a Chatbot by Prompting GPT-3
Authors
Prompting Large Language Models (LLMs) is an exciting new approach to designing chatbots. But can it improve LLM’s user experience (UX) reliably enough to power chatbot products? Our attempt to design a robust chatbot by prompting GPT-3/4 alone suggests: not yet. Prompts made achieving “80%” UX goals easy, but not the remaining 20%. Fixing the few remaining interaction breakdowns resembled herding cats: We could not address one UX issue or test one design solution at a time; instead, we had to handle everything everywhere all at once. Moreover, because no prompt could make GPT reliably say “I don’t know” when it should, the user-GPT conversations had no guardrails after a breakdown occurred, often leading to UX downward spirals. These risks incentivized us to design highly prescriptive prompts and scripted bots, counter to the promises of LLM-powered chatbots. This paper describes this case study, unpacks prompting’s fickleness and its impact on UX design processes, and discusses implications for LLM-based design methods and tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
DialogLab: Authoring, Simulating, and Testing Dynamic Human-AI Group Conversations
UIST '25· Conversational Chatbots +1
- 80%
Conversation Progress Guide : UI System for Enhancing Self-Efficacy in Conversational AI
CHI '25· Conversational Chatbots +2
- 80%
AmbigChat: Interactive Hierarchical Clarification for Ambiguous Open-Domain Question Answering
UIST '25· Conversational Chatbots +1
- 75%
Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models
UIST '23· Human-LLM Collaboration
- 67%
Help me and I’ll help you: Speakers’ and listeners’ collaborative effort and the division of labour in human-agent collaborative communication
CHI '26· Conversational Chatbots +2
- 67%
Is Conversational XAI All You Need? Human-AI Decision Making With a Conversational XAI Assistant
IUI '25· Conversational Chatbots +2
- 67%
ChoiceMates: Supporting Unfamiliar Online Decision-Making with Multi-Agent Conversational Interactions
IUI '26· Human-LLM Collaboration +2
- 60%
Iris: A Conversational Agent for Complex Tasks
CHI '18· Conversational Chatbots +1
- 60%
Mapping Machine Learning Advances from HCI Research to Reveal Starting Places for Design Innovation
CHI '18· Human-LLM Collaboration
- 60%
Sketching NLP: A Case Study of Exploring the Right Things To Design with Language Intelligence
CHI '19· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)