Characterizing User-Reported Risks across LLM Chatbots
Authors
Paper Title
Characterizing User-Reported Risks across LLM Chatbots
Publication Info
- Topic area: User-reported risks in large language model (LLM) chatbots.
- Keywords: LLM chatbots, user-reported risks, NIST AI RMF, ChatGPT, Gemini, Claude, human-centered AI, safety, privacy, fairness, reliability.
Background and Problem
- Problem / challenge: Existing research on LLM risks is limited to single models or specific risk types, often conducted in controlled environments that fail to capture real-world user experiences across multiple chatbots.
- Significance: Understanding user-reported risks is critical for designing safer, more reliable, and user-aligned LLM chatbots, especially as these tools become integral to daily life.
- Motivation and related work: Prior work has documented risks like toxic content, hallucinations, and biases but lacks a systematic, multi-risk, cross-chatbot analysis grounded in user experiences. This paper addresses this gap by analyzing user-reported risks across seven major LLM chatbots using the NIST AI Risk Management Framework.
Solution
- Proposed approach: A hybrid methodology combining the NIST AI Risk Management Framework (top-down structure) with bottom-up user-driven topic clustering to analyze Reddit discussions about seven LLM chatbots.
- Novelty:
- Empirical characterization of user-reported risks across seven LLM chatbots.
- Identification of chatbot-specific risk patterns and uneven risk distributions.
- Development of a hybrid methodology integrating structured frameworks with user-driven clustering.
- Exploration of user trade-offs between utility and risks in daily chatbot use.
- Procedure and key techniques:
- Collect Reddit data mentioning ChatGPT, Gemini, Claude, DeepSeek, Llama, Mistral, and Qwen (4,438 posts and 48,797 comments).
- Use the NIST AI RMF to classify risks into seven categories: Valid & Reliable, Safe, Fair, Secure & Resilient, Accountable & Transparent, Explainable & Interpretable, Privacy-Enhanced.
- Apply topic modeling (BERTopic) to cluster user-reported risks and identify emergent themes.
- Construct an interactive knowledge graph to visualize relationships among chatbots, risk categories, and user experiences.
Results
- Concrete findings:
- "Valid and Reliable" risks dominate across all chatbots (58.39% of reports), followed by "Accountable and Transparent" (16.35%) and "Secure and Resilient" (9.27%).
- ChatGPT is disproportionately associated with safety and fairness concerns, Gemini with privacy issues, and Claude with security and operational resilience challenges.
- Less frequent risks like "Explainability" (1.23%) and "Privacy" (3.85%) often manifest as user trade-offs, while common risks like "Fairness" (4.52%) and "Safety" (6.39%) are experienced as direct harms.
- Advantage over baselines:
- Provides a multi-risk, cross-chatbot analysis grounded in real-world user experiences, unlike prior studies focusing on single models or risks.
- Combines structured frameworks with user-driven clustering for nuanced insights.
- Experiments / evaluation:
- Data collected from 51 subreddits spanning chatbot-specific, AI-related, and general-interest communities.
- Validation of risk extraction pipeline using human annotations (Krippendorff’s α ≥ 0.87).
- Statistical analysis (Chi-Square tests) confirms significant differences in risk distributions across chatbots.
- Limitations and future work:
- Reddit’s user base may not represent the global population of LLM users, and linguistic biases may exclude non-English-speaking users.
- Data collection via Reddit API may introduce recency or popularity bias.
- Future work should incorporate diverse platforms, multilingual sources, and longitudinal analyses.
Summary
This study provides a large-scale, empirical characterization of user-reported risks across seven LLM chatbots, revealing chatbot-specific risk patterns and the dominance of "Valid and Reliable" concerns. Using a hybrid methodology, it integrates the NIST AI RMF with user-driven topic modeling to analyze Reddit discussions. Key findings highlight uneven risk distributions, user trade-offs, and the need for human-centered risk mitigation strategies. The results emphasize the importance of aligning chatbot design and governance with user priorities, particularly reliability, safety, and fairness, while addressing less visible risks like privacy and explainability.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
AI and My Values: User Perceptions of LLMs’ Ability to Extract, Embody, and Explain Human Values from Casual Conversations
CHI '26· Human-LLM Collaboration +3
- 86%
Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges
CHI '23· Human-LLM Collaboration +2
- 86%
Be Friendly, Not Friends: How LLM Sycophancy Shapes User Trust
CHI '26· Human-LLM Collaboration +2
- 86%
Sensemaking in Multi-Agent LLM Interfaces: How Users Interpret Transparency and Trustworthiness Cues
CHI '26· Human-LLM Collaboration +2
- 75%
Who Controls the Conversation? User Perspectives On Generative AI (LLM) System Prompts
CHI '26· Human-LLM Collaboration +3
- 71%
Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information Seeking
CHI '24· Human-LLM Collaboration +2
- 71%
Plurals: A System for Guiding LLMs via Simulated Social Ensembles
CHI '25· Human-LLM Collaboration +2
- 71%
Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
CHI '26· Human-LLM Collaboration +2
- 71%
Characterizing Unintended Consequences of GUI Agents For Web Browsing
CHI '26· Human-LLM Collaboration +2
- 63%
Planning for Natural Language Failures with the AI Playbook
CHI '21· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)