The Privacy Paradox of LLMs: User Perceptions and the Reality of PII Leakage

Explainable AI (XAI)Privacy by Design & User ControlPrivacy Perception & Decision-MakingAI/ML Researchers & EngineersPrivacy Policy MakersSoftware Engineers & Developers

Paper Title

The Privacy Paradox of LLMs: User Perceptions and the Reality of PII Leakage

Publication Info

  • Topic area: Privacy risks and user perceptions in large language models (LLMs).
  • Keywords: LLMs, privacy leakage, PII, user perceptions, privacy paradox, training data, extraction attacks, privacy literacy, user behavior, safeguards.

Background and Problem

  • Problem / challenge: LLMs are vulnerable to leaking personally identifiable information (PII) from training data, but the extent of this leakage and user perceptions of the risks are not well understood. Existing defenses are insufficient to prevent such leakage.
  • Significance: PII leakage poses risks of identity theft, fraud, and other privacy violations, undermining trust in LLMs and raising ethical concerns about their deployment.
  • Motivation and related work: Prior studies have demonstrated PII extraction feasibility and proposed defensive measures, but none have systematically assessed real-world leakage across mainstream LLMs or explored user perceptions and behaviors in response to these risks.

Solution

  • Proposed approach: A mixed-methods study combining empirical evaluation of PII leakage in mainstream LLMs with qualitative and quantitative user studies to assess privacy perceptions, literacy, and behavioral responses.
  • Novelty:
    1. Comprehensive evaluation of PII leakage across targeted and non-targeted attacks on ten mainstream LLMs.
    2. Mixed-methods user study (20 interviews, 204 survey participants) revealing gaps in privacy literacy and contradictions in user attitudes and behaviors.
    3. Design implications for balancing privacy and utility in future LLMs.
  • Procedure and key techniques:
    • Empirical evaluation using targeted and non-targeted PII extraction attacks on ten LLMs, measuring attack success rates (ASR) for emails, phone numbers, and professional PII.
    • Semi-structured interviews to explore user experiences, perceptions, and suggestions regarding LLM privacy risks.
    • Online survey to validate interview findings and assess user literacy, privacy concerns, and behavioral intentions.

Results

  • Concrete findings:
    • Targeted extraction achieved an average ASR of 78.3% for emails and 20.3% for phone numbers.
    • Non-targeted extraction yielded 6,919 valid PII instances across professions, with an average ASR of 34.6%.
    • Significant cross-model variability in leakage susceptibility, with some models exceeding 50% ASR.
    • Users demonstrated limited understanding of LLM training data and underestimated PII extraction risks.
    • Despite privacy concerns, users continued to use LLMs due to perceived utility, often exhibiting privacy cynicism.
  • Advantage over baselines:
    • First systematic evaluation of PII leakage across multiple LLMs using both targeted and non-targeted attacks.
    • Integration of technical findings with user studies to bridge the gap between empirical risks and user perceptions.
  • Experiments / evaluation:
    • Empirical evaluation on ten LLMs using public datasets (e.g., Enron, OpenWebText2, CC-News) for ground truth.
    • Mixed-methods user study involving 20 interviews and 204 survey participants, analyzing privacy literacy, concerns, and behavioral responses.
  • Limitations and future work:
    • Limited generalizability of user study findings due to participant demographics (primarily Chinese users).
    • Potential underestimation of ASR for some LLMs due to insufficient parameter tuning.
    • Future work should explore cross-cultural differences in privacy perceptions and develop more robust privacy-preserving mechanisms.

Summary

This study highlights significant privacy risks in mainstream LLMs, with empirical evaluations showing high success rates for PII extraction and user studies revealing gaps in privacy literacy and contradictory behaviors. Despite concerns about PII leakage, users continue to adopt LLMs due to their utility, reflecting a privacy paradox. The findings underscore the need for greater transparency, user control, and technical safeguards to balance privacy and usability in LLMs. These insights provide actionable guidance for designing trustworthy and privacy-preserving AI systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222364/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791809
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Explainable AI (XAI), Privacy by Design & User Control, Privacy Perception & Decision-Making
work
Professions
AI/ML Researchers & Engineers, Privacy Policy Makers, Software Engineers & Developers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers