Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking
Authors
Large Language Model (LLM) assistants, such as ChatGPT, have emerged as potential alternatives to search methods for helping users navigate complex, feature-rich software. LLMs use vast training data from domain-specific texts, software manuals, and code repositories to mimic human-like interactions, offering tailored assistance, including step-by-step instructions. In this work, we investigated LLM-generated software guidance through a within-subject experiment with 16 participants and follow-up interviews. We compared a baseline LLM assistant with an LLM optimized for particular software contexts, SoftAIBot, which also offered guidelines for constructing appropriate prompts. We assessed task completion, perceived accuracy, relevance, and trust. Surprisingly, although SoftAIBot outperformed the baseline LLM, our results revealed no significant difference in LLM usage and user perceptions with or without prompt guidelines and the integration of domain context. Most users struggled to understand how the prompt's text related to the LLM's responses and often followed the LLM's suggestions verbatim, even if they were incorrect. This resulted in difficulties when using the LLM's advice for software tasks, leading to low task completion rates. Our detailed analysis also revealed that users remained unaware of inaccuracies in the LLM's responses, indicating a gap between their lack of software expertise and their ability to evaluate the LLM's assistance. With the growing push for designing domain-specific LLM assistants, we emphasize the importance of incorporating explainable, context-aware cues into LLMs to help users understand prompt-based interactions, identify biases, and maximize the utility of LLM assistants.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
Vibrational Artificial Subtle Expressions: Conveying System’s Confidence Level to Users by Means of Smartphone Vibration
CHI '18· Vibrotactile Feedback & Skin Stimulation +1
- 67%
A Human-Computer Collaborative Editing Tool for Conceptual Diagrams
CHI '23· Human-LLM Collaboration +1
- 67%
From Text to Self: Users’ Perception of AIMC Tools on Interpersonal Communication and Self
CHI '24· Multilingual & Cross-Cultural Voice Interaction +2
- 67%
Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation
CHI '24· Human-LLM Collaboration +1
- 67%
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
CHI '25· Human-LLM Collaboration +1
- 67%
Trends, Challenges and Processes in Conversational Agent Design: Exploring Practitioners' Views through Semi-Structured Interviews
CUI '23· Conversational Chatbots +1
- 67%
Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents
CUI '25· Social & Collaborative VR +1
- 67%
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs
IUI '24· Human-LLM Collaboration +1
- 67%
FigurA11y: AI Assistance for Writing Scientific Alt Text
IUI '24· Explainable AI (XAI) +1
- 67%
An Exploratory Study on How AI Awareness Impacts Human-AI Design Collaboration
IUI '25· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)