“As an AI language model, I cannot”: Investigating LLM Denials of User Requests
Honorable MentionAuthors
Title of the Paper
“As an AI language model, I cannot”: Investigating LLM Denials of User Requests
Paper Information
- Domain: Human-Computer Interaction (HCI), Research on the Design and User Experience of AI Language Model Denials of User Requests
- Keywords: Human-Computer Interaction, Error Handling, LLM Denials, User Experience, GPT-4
Research Background and Problem
- Problem or Challenge:
- Users frequently encounter denials or errors when interacting with large language models (LLMs), reflecting a gap between user requests and system capabilities. These denial scenarios may arise from technical limitations (e.g., model capability constraints) or social reasons (e.g., policy restrictions).
- Previous research has focused on error handling and interaction repair strategies but lacks studies on the impact of LLM denial styles on user experience.
- Significance:
- The increasing use of LLMs in modern interactive technologies makes optimizing user experience critical. Understanding how users perceive denial responses can provide a foundation for better interaction design.
- Motivation and Related Work:
- Prior studies have shown that users’ perceptions of error responses and denial interactions significantly influence their trust in the system, satisfaction, and mental models.
- In the field of explainable AI, high-quality explanations help users understand system limitations and reduce frustration. This study aims to explore specific denial strategies in the context of LLMs.
Solution
- Method or Solution:
- The study compares four denial styles: baseline denial (simple denial without explanation), factual denial (providing reasons), diverting denial (redirecting to a related topic), and opinionated denial (correcting user behavior).
- Two experiments were designed to study user experiences with denials due to technical and social reasons.
- Innovations:
- This is the first systematic study on the impact of different denial styles on user perceptions (frustration, effectiveness, appropriateness, and relevance) and proposes specific recommendations for optimizing LLM denial design.
- Implementation Steps and Key Techniques:
- Develop an interaction model based on GPT-4 where all user requests are denied.
- Create scenarios and tasks to simulate real user needs, covering topics such as health, politics, and humor.
- Analyze differences in user perceptions of the four denial styles across scenarios with different denial reasons (technical and social).
Research Findings
- Specific Findings:
- The diverting denial style resulted in the lowest frustration and the highest scores for effectiveness and appropriateness.
- The baseline denial style received the lowest scores, with users perceiving it as “completely unhelpful.”
- For social reason denials, the opinionated denial style performed relatively well, though not as effectively as the diverting denial style.
- For technical reason denials, factual denial did not significantly outperform baseline denial.
- Advantages Compared to Existing Solutions:
- Optimizing denial style design can improve user experience, reduce frustration, and enhance the usability and satisfaction of LLMs.
- Experimental or Evaluation Results:
- Users expressed general dissatisfaction with baseline denials, while diverting denials received significantly higher evaluations.
- Quantitative data indicated that denial style design has a significant impact across all measured dimensions (frustration, effectiveness, appropriateness, relevance).
- Limitations and Future Directions:
- The study is limited to single-interaction scenarios and does not address denial strategies in complex dialogues.
- The sample population is based on a U.S. social context; future research could explore cross-cultural perspectives.
- Other potential denial styles (e.g., humorous denial) were not discussed.
This study provides an academic foundation for designing more humanized and interactive LLM denial styles, while also opening new directions for future research on optimizing AI user experience. The study suggests offering richer and more valuable denial responses within the constraints of technical limitations and legal-ethical frameworks to maintain positive interaction experiences.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What are the UX impacts of different LLM refusal styles?Category: Dialogue System Management and Error HandlingSimilar questionsarrow_forward
- How do different refusal styles perform in refusal scenarios caused by technical limitations versus social reasons?Category: Dialogue System Management and Error HandlingSimilar questionsarrow_forward
- Which refusal style can best minimize user frustration and improve interaction appropriateness and effectiveness?Category: Dialogue System Management and Error HandlingSimilar questionsarrow_forward
Practical Problems
1- Users feel frustrated when AI refuses requests and struggle to understand the reasons for refusal.Category: Dialogue System Management and Error HandlingSimilar questionsarrow_forward
- 80%
Black LLMirror: User (Self) Perceptions in Black American English Interactions with LLMs
CHI '26· Human-LLM Collaboration +2
- 75%
Continual Human-in-the-Loop Optimization
CHI '25· Human-LLM Collaboration
- 67%
Are Two Heads Better Than One in AI-Assisted Decision Making? Comparing the Behavior and Performance of Groups and Individuals in Human-AI Collaborative Recidivism Risk Assessment
CHI '23· Human-LLM Collaboration +2
- 67%
What is Human-Centered about Human-Centered AI? A Map of the Research Landscape
CHI '23· Human-LLM Collaboration +2
- 67%
Simulacrum of stories: Examining Large Language Models as Qualitative Research Participants
CHI '25· Human-LLM Collaboration +2
- 67%
Understanding Compliance and Conversion Dynamics in Multi-Agent Collectives
CHI '26· Human-LLM Collaboration +2
- 67%
"Are we writing an advice column for Spock here?" Understanding Stereotypes in AI Advice for Autistic Users
CHI '26· Human-LLM Collaboration +2
- 60%
Fair Machine Guidance to Enhance Fair Decision Making in Biased People
CHI '24· AI-Assisted Decision-Making & Automation +1
- 60%
EvAlignUX: Advancing UX Evaluation through LLM-Supported Metrics Exploration
CHI '25· Human-LLM Collaboration +1
- 60%
Responsibility Attribution in Human Interactions with Everyday AI Systems
CHI '25· AI Ethics, Fairness & Accountability +1
Based on Jaccard similarity of research subtopics & professions (≥60%)