Beyond PII: How Users Attempt to Estimate and Mitigate Implicit LLM Inference
Authors
Sai Teja Peddinti
Google Inc.Paper Title
Beyond PII: How Users Attempt to Estimate and Mitigate Implicit LLM Inference
Publication Info
- Topic area: Privacy risks and user strategies in interactions with Large Language Models (LLMs).
- Keywords: LLM privacy, personal attribute inference, user perception, text sanitization, rewriting strategies, privacy risks, ChatGPT, Rescriber, SynthPAI, inference-aware design.
Background and Problem
- Problem / challenge: Users of LLMs face privacy risks from implicit inference of personal attributes (e.g., age, location, occupation) from seemingly innocuous text. Existing tools focus on explicit Personally Identifiable Information (PII) but fail to address inference-based risks. Users also lack the ability to anticipate or mitigate these risks effectively.
- Significance: Inference-based privacy risks can reveal sensitive information without explicit disclosure, affecting trust and safety in LLM interactions across domains such as workplace assistance, healthcare, and education.
- Motivation and related work: Prior research has shown that LLMs can infer personal attributes and that users often misunderstand LLM capabilities. However, no studies have systematically investigated user perceptions of implicit inference risks or evaluated user strategies for mitigating them. This paper addresses these gaps.
Solution
- Proposed approach: A mixed-methods study to evaluate user perceptions of inference risks, concern levels, and rewriting strategies to block inference. The study compares user performance with automated tools like ChatGPT and Rescriber.
- Novelty:
- First systematic investigation of user perceptions and rewriting strategies for implicit attribute inference.
- Empirical comparison of user-generated rewrites with outputs from ChatGPT and Rescriber.
- Analysis of user strategies for mitigating inference risks and their effectiveness.
- Insights into the design of inference-aware systems for privacy protection.
- Procedure and key techniques:
- Conducted a survey with 240 U.S. participants using the SynthPAI dataset, which includes text snippets with inferable attributes.
- Participants estimated inferable attributes, rated their concern levels, and attempted to rewrite text to block inference.
- Evaluated rewrites for effectiveness (using GPT-4) and semantic similarity (using BERTScore).
- Benchmarked user rewrites against ChatGPT and Rescriber outputs.
- Analyzed rewrite strategies through qualitative coding and statistical tests.
Results
- Concrete findings:
- Participants correctly estimated inferable attributes only 48% of the time, slightly above random guessing (40%).
- User rewrites were effective in 28% of cases, outperforming Rescriber (24%) but underperforming ChatGPT (50%).
- Paraphrasing was the most common rewrite strategy (60%) but had the lowest effectiveness (37%). Adding ambiguity (71%) and generalization (67%) were more effective.
- Concern levels did not vary significantly across attribute types, with 44% of participants expressing concern for all attributes.
- Advantage over baselines:
- User rewrites were more effective than Rescriber but less effective than ChatGPT.
- ChatGPT preserved semantic similarity better than users while achieving higher rewrite effectiveness.
- Experiments / evaluation:
- Used 32 text snippets from SynthPAI, each with one inferable attribute.
- Measured user estimation accuracy, concern levels, and rewrite effectiveness.
- Benchmarked against automated tools and analyzed rewrite strategies qualitatively.
- Limitations and future work:
- Synthetic dataset limits real-world applicability; future studies should use real-world conversations.
- Survey setting may not reflect real user behavior under higher stakes.
- Generalizability limited to U.S. adults; cross-cultural studies needed.
- Future work should explore interactive warnings, rewriting suggestions, and longitudinal studies of user behavior.
Summary
This study investigates how users perceive and mitigate inference-based privacy risks in LLM interactions. Participants struggled to estimate inferable attributes and often employed ineffective rewriting strategies, with paraphrasing being the most common but least effective. Automated tools like ChatGPT outperformed users in blocking inference while preserving semantic meaning. The findings highlight the need for inference-aware systems that proactively support users in mitigating privacy risks. Future research should focus on real-world datasets, interactive tools, and longitudinal studies to enhance privacy protection in LLM applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
A Scoping Review and Guidelines on Privacy Policy's Visualization from an HCI Perspective
CHI '26· Privacy Perception & Decision-Making +2
- 83%
The Privacy Paradox of LLMs: User Perceptions and the Reality of PII Leakage
CHI '26· Explainable AI (XAI) +2
- 83%
Understanding User Needs Underlying the Expected Roles of LLM-Based Chatbots in Privacy Decision-Making
CHI '26· Explainable AI (XAI) +2
- 83%
Privy: Envisioning and Mitigating Privacy Risks for Consumer-facing AI Product Concepts
CHI '26· Explainable AI (XAI) +2
- 71%
From Fragmentation to Integration: Exploring the Design Space of AI Agents for Human-as-the-Unit Privacy Management
CHI '26· Privacy by Design & User Control +3
- 71%
Privacy Control in Conversational LLM Platforms: A Walkthrough Study
CHI '26· Explainable AI (XAI) +3
- 67%
Mind the SIM: Awareness and Mental Models in a South Korean Case Study
CHI '26· Privacy by Design & User Control +2
- 63%
PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice Questions
CHI '26· Explainable AI (XAI) +3
- 60%
SIGCHI Social Impact Award Talk – Making Privacy and Security More Usable
CHI '18· Privacy by Design & User Control +1
- 60%
You 'Might' Be Affected: An Empirical Analysis of Readability and Usability Issues in Data Breach Notifications
CHI '19· Privacy by Design & User Control +1
Based on Jaccard similarity of research subtopics & professions (≥60%)