Characterizing User-Reported Risks across LLM Chatbots

Human-LLM CollaborationExplainable AI (XAI)AI Ethics, Fairness & AccountabilityPrivacy by Design & User ControlAI/ML Researchers & EngineersUI/UX DesignersHCI Researchers

Paper Title

Characterizing User-Reported Risks across LLM Chatbots

Publication Info

  • Topic area: User-reported risks in large language model (LLM) chatbots.
  • Keywords: LLM chatbots, user-reported risks, NIST AI RMF, ChatGPT, Gemini, Claude, human-centered AI, safety, privacy, fairness, reliability.

Background and Problem

  • Problem / challenge: Existing research on LLM risks is limited to single models or specific risk types, often conducted in controlled environments that fail to capture real-world user experiences across multiple chatbots.
  • Significance: Understanding user-reported risks is critical for designing safer, more reliable, and user-aligned LLM chatbots, especially as these tools become integral to daily life.
  • Motivation and related work: Prior work has documented risks like toxic content, hallucinations, and biases but lacks a systematic, multi-risk, cross-chatbot analysis grounded in user experiences. This paper addresses this gap by analyzing user-reported risks across seven major LLM chatbots using the NIST AI Risk Management Framework.

Solution

  • Proposed approach: A hybrid methodology combining the NIST AI Risk Management Framework (top-down structure) with bottom-up user-driven topic clustering to analyze Reddit discussions about seven LLM chatbots.
  • Novelty:
    1. Empirical characterization of user-reported risks across seven LLM chatbots.
    2. Identification of chatbot-specific risk patterns and uneven risk distributions.
    3. Development of a hybrid methodology integrating structured frameworks with user-driven clustering.
    4. Exploration of user trade-offs between utility and risks in daily chatbot use.
  • Procedure and key techniques:
    1. Collect Reddit data mentioning ChatGPT, Gemini, Claude, DeepSeek, Llama, Mistral, and Qwen (4,438 posts and 48,797 comments).
    2. Use the NIST AI RMF to classify risks into seven categories: Valid & Reliable, Safe, Fair, Secure & Resilient, Accountable & Transparent, Explainable & Interpretable, Privacy-Enhanced.
    3. Apply topic modeling (BERTopic) to cluster user-reported risks and identify emergent themes.
    4. Construct an interactive knowledge graph to visualize relationships among chatbots, risk categories, and user experiences.

Results

  • Concrete findings:
    • "Valid and Reliable" risks dominate across all chatbots (58.39% of reports), followed by "Accountable and Transparent" (16.35%) and "Secure and Resilient" (9.27%).
    • ChatGPT is disproportionately associated with safety and fairness concerns, Gemini with privacy issues, and Claude with security and operational resilience challenges.
    • Less frequent risks like "Explainability" (1.23%) and "Privacy" (3.85%) often manifest as user trade-offs, while common risks like "Fairness" (4.52%) and "Safety" (6.39%) are experienced as direct harms.
  • Advantage over baselines:
    • Provides a multi-risk, cross-chatbot analysis grounded in real-world user experiences, unlike prior studies focusing on single models or risks.
    • Combines structured frameworks with user-driven clustering for nuanced insights.
  • Experiments / evaluation:
    • Data collected from 51 subreddits spanning chatbot-specific, AI-related, and general-interest communities.
    • Validation of risk extraction pipeline using human annotations (Krippendorff’s α ≥ 0.87).
    • Statistical analysis (Chi-Square tests) confirms significant differences in risk distributions across chatbots.
  • Limitations and future work:
    • Reddit’s user base may not represent the global population of LLM users, and linguistic biases may exclude non-English-speaking users.
    • Data collection via Reddit API may introduce recency or popularity bias.
    • Future work should incorporate diverse platforms, multilingual sources, and longitudinal analyses.

Summary

This study provides a large-scale, empirical characterization of user-reported risks across seven LLM chatbots, revealing chatbot-specific risk patterns and the dominance of "Valid and Reliable" concerns. Using a hybrid methodology, it integrates the NIST AI RMF with user-driven topic modeling to analyze Reddit discussions. Key findings highlight uneven risk distributions, user trade-offs, and the need for human-centered risk mitigation strategies. The results emphasize the importance of aligning chatbot design and governance with user priorities, particularly reliability, safety, and fairness, while addressing less visible risks like privacy and explainability.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222783/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791505
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), AI Ethics, Fairness & Accountability, Privacy by Design & User Control
work
Professions
AI/ML Researchers & Engineers, UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers