The Privacy Paradox of LLMs: User Perceptions and the Reality of PII Leakage
Authors
Paper Title
The Privacy Paradox of LLMs: User Perceptions and the Reality of PII Leakage
Publication Info
- Topic area: Privacy risks and user perceptions in large language models (LLMs).
- Keywords: LLMs, privacy leakage, PII, user perceptions, privacy paradox, training data, extraction attacks, privacy literacy, user behavior, safeguards.
Background and Problem
- Problem / challenge: LLMs are vulnerable to leaking personally identifiable information (PII) from training data, but the extent of this leakage and user perceptions of the risks are not well understood. Existing defenses are insufficient to prevent such leakage.
- Significance: PII leakage poses risks of identity theft, fraud, and other privacy violations, undermining trust in LLMs and raising ethical concerns about their deployment.
- Motivation and related work: Prior studies have demonstrated PII extraction feasibility and proposed defensive measures, but none have systematically assessed real-world leakage across mainstream LLMs or explored user perceptions and behaviors in response to these risks.
Solution
- Proposed approach: A mixed-methods study combining empirical evaluation of PII leakage in mainstream LLMs with qualitative and quantitative user studies to assess privacy perceptions, literacy, and behavioral responses.
- Novelty:
- Comprehensive evaluation of PII leakage across targeted and non-targeted attacks on ten mainstream LLMs.
- Mixed-methods user study (20 interviews, 204 survey participants) revealing gaps in privacy literacy and contradictions in user attitudes and behaviors.
- Design implications for balancing privacy and utility in future LLMs.
- Procedure and key techniques:
- Empirical evaluation using targeted and non-targeted PII extraction attacks on ten LLMs, measuring attack success rates (ASR) for emails, phone numbers, and professional PII.
- Semi-structured interviews to explore user experiences, perceptions, and suggestions regarding LLM privacy risks.
- Online survey to validate interview findings and assess user literacy, privacy concerns, and behavioral intentions.
Results
- Concrete findings:
- Targeted extraction achieved an average ASR of 78.3% for emails and 20.3% for phone numbers.
- Non-targeted extraction yielded 6,919 valid PII instances across professions, with an average ASR of 34.6%.
- Significant cross-model variability in leakage susceptibility, with some models exceeding 50% ASR.
- Users demonstrated limited understanding of LLM training data and underestimated PII extraction risks.
- Despite privacy concerns, users continued to use LLMs due to perceived utility, often exhibiting privacy cynicism.
- Advantage over baselines:
- First systematic evaluation of PII leakage across multiple LLMs using both targeted and non-targeted attacks.
- Integration of technical findings with user studies to bridge the gap between empirical risks and user perceptions.
- Experiments / evaluation:
- Empirical evaluation on ten LLMs using public datasets (e.g., Enron, OpenWebText2, CC-News) for ground truth.
- Mixed-methods user study involving 20 interviews and 204 survey participants, analyzing privacy literacy, concerns, and behavioral responses.
- Limitations and future work:
- Limited generalizability of user study findings due to participant demographics (primarily Chinese users).
- Potential underestimation of ASR for some LLMs due to insufficient parameter tuning.
- Future work should explore cross-cultural differences in privacy perceptions and develop more robust privacy-preserving mechanisms.
Summary
This study highlights significant privacy risks in mainstream LLMs, with empirical evaluations showing high success rates for PII extraction and user studies revealing gaps in privacy literacy and contradictory behaviors. Despite concerns about PII leakage, users continue to adopt LLMs due to their utility, reflecting a privacy paradox. The findings underscore the need for greater transparency, user control, and technical safeguards to balance privacy and usability in LLMs. These insights provide actionable guidance for designing trustworthy and privacy-preserving AI systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Beyond PII: How Users Attempt to Estimate and Mitigate Implicit LLM Inference
CHI '26· Explainable AI (XAI) +2
- 75%
PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice Questions
CHI '26· Explainable AI (XAI) +3
- 71%
Understanding Challenges for Developers to Create Accurate Privacy Nutrition Labels
CHI '22· Privacy by Design & User Control +1
- 71%
A Scoping Review and Guidelines on Privacy Policy's Visualization from an HCI Perspective
CHI '26· Privacy Perception & Decision-Making +2
- 71%
Understanding User Needs Underlying the Expected Roles of LLM-Based Chatbots in Privacy Decision-Making
CHI '26· Explainable AI (XAI) +2
- 71%
Uncovering Relationships Between Android Developers, User Privacy, and Developer Willingness to Reduce Fingerprinting Risks
CHI '26· Privacy by Design & User Control +2
- 71%
Helping Johnny Make Sense of Privacy Policies with LLMs
CHI '26· Privacy by Design & User Control +2
- 71%
Privy: Envisioning and Mitigating Privacy Risks for Consumer-facing AI Product Concepts
CHI '26· Explainable AI (XAI) +2
- 71%
PrivWeb: Unobtrusive and Content-aware Privacy Protection For Web Agents
CHI '26· Privacy by Design & User Control +2
- 67%
Contextualizing Privacy Decisions for Better Prediction (and Protection)
CHI '18· Privacy by Design & User Control +1
Based on Jaccard similarity of research subtopics & professions (≥60%)