CoPrompter: User-Centric Evaluation of LM Instruction Alignment for Improved Prompt Engineering
Authors
Ensuring large language models' (LLMs) responses align with prompt instructions is crucial for application development. Based on our formative study with industry professionals, the alignment requires heavy human involvement and tedious trial-and-error especially when there are many instructions in the prompt. To address these challenges, we introduce CoPrompter, a framework that identifies misalignment based on assessing multiple LLM responses with criteria. It proposes a method to generate evaluation criteria questions derived directly from prompt requirements and an interface to turn these questions into a user-editable checklist. Our user study with industry prompt engineers shows that CoPrompter improves the ability to identify and refine instruction alignment with prompt requirements over traditional methods, helps them understand where and how frequently models fail to follow user's prompt requirements, and helps in clarifying their own requirements, giving them greater control over the response evaluation process. We also present the design lessons to underscore our system's potential to streamline the prompt engineering process.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can LLMs' instruction alignment performance be evaluated when executing complex multi-instruction prompts?Category: LLM Prompt Engineering and Authoring ToolsSimilar questionsarrow_forward
- Can user-involved dynamic evaluation frameworks improve prompt engineering efficiency and result quality?Category: LLM Prompt Engineering and Authoring ToolsSimilar questionsarrow_forward
- Can decomposing and automatically generating evaluation criteria more precisely identify LLM alignment issues?Category: LLM Prompt Engineering and Authoring ToolsSimilar questionsarrow_forward
Practical Problems
1- Users cannot efficiently evaluate and optimize LLM instruction effectiveness in complex tasks.Category: LLM Prompt Engineering and Authoring ToolsSimilar questionsarrow_forward
- 100%
Interactive Debugging and Steering of Multi-Agent AI Systems
CHI '25· Human-LLM Collaboration +2
- 100%
DIY: Helping People Assess the Correctness of Natural Language to SQL Systems
IUI '21· Human-LLM Collaboration +2
- 83%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 83%
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
CHI '26· Human-LLM Collaboration +2
- 83%
DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows
CHI '26· Human-LLM Collaboration +2
- 83%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 83%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 83%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 83%
RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
UIST '25· Human-LLM Collaboration +2
- 80%
UMLAUT: Debugging Deep Learning Programs using Program Structure and Model Behavior
CHI '21· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)