Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
Authors
Research Background and Issues
-
What problems or challenges did the authors identify?
- With the increasing application of chatbot systems driven by large language models (LLMs), users may excessively disclose sensitive personal information, such as addresses, emails, health data, etc., during interactions. This not only poses risks of privacy breaches but also challenges the classical principle of data minimization.
- Current privacy protection technologies face issues such as insufficient transparency and control for users over their data processing. These technologies fail to adequately address personalized privacy needs and the trade-off between privacy and utility.
- Users are often unwilling to spend time and effort manually deleting sensitive data, with some even unaware of the privacy risks involved.
-
Why is this issue important?
- Data breaches and memory vulnerabilities in large language models can lead to severe privacy and security risks, including identity theft or accidental disclosure of sensitive information. Additionally, privacy concerns limit users' trust in chatbots, hindering further technological adoption.
-
Research Motivation and Related Work
- Existing studies and methods have not sufficiently addressed the issue of users actively controlling privacy and data minimization. Post-processing or hardware isolation protections provide only limited solutions.
- Furthermore, widespread user privacy concerns are often overlooked due to inconvenience or lack of understanding, highlighting the need for interactive, user-friendly tools to support data control behaviors.
Solution
-
What methods or solutions did the authors propose?
- The authors proposed a browser extension called Rescriber, which detects and highlights potential sensitive information in user input and offers two methods for removal: replacement (substituting sensitive words with placeholders) and abstraction (replacing specific descriptions with generalized ones).
- They introduced the concept of "small LLM-driven user-side privacy control," enabling users to manage data minimization autonomously by running small models locally on their devices (e.g., Llama3-8B).
-
What are the innovative aspects of this solution?
- By combining user active participation with LLM-based detection technology, the solution achieves an effective balance between privacy protection and model response quality.
- It provides instant, flexible privacy control methods, granting users greater autonomy while reducing operational burdens through automation.
- The feasibility of using small LLMs as privacy agents was explored and validated. Local processing reduces reliance on large cloud models, alleviating privacy concerns while offering high transparency.
-
What are the implementation steps and key technologies used?
- The browser extension is designed with a frontend (user interface and interaction) and a backend (small LLM-based detection and abstraction).
- Key functionalities:
- Real-time detection of personal information in user input using multi-level classification (e.g., NAME, EMAIL, HEALTH_INFORMATION).
- Two data processing options: replacement and abstraction, along with a "write-back" feature to automatically restore placeholders to original information for maintaining response readability.
- Two core models were tested: Rescriber-GPT-4o (based on the GPT-4 cloud model) and Rescriber-Llama3-8B (small local model).
Research Outcomes
-
What specific results were achieved?
- The tool successfully balanced privacy protection and task generation effectiveness. Most participants reduced sensitive information disclosure while generating efficient and satisfactory outputs.
- The small model Rescriber-Llama3-8B performed comparably to GPT-4o in certain aspects, such as subjective privacy enhancement and user adoption intent, demonstrating the feasibility of locally running small models for privacy protection.
- The tool significantly improved users' awareness of data minimization and educated them on identifying and reducing potential privacy risks through its suggestions.
-
How does it compare to existing solutions?
- It provides immediate, user-visible privacy protection actions, allowing users to review and modify information before sending it.
- Local operation based on smaller models offers transparency and avoids reliance on cloud-based privacy processing.
- By integrating replacement and abstraction methods, it meets diverse privacy needs across different scenarios.
-
What were the experimental or evaluation results?
- There was no significant difference in subjective user satisfaction and adoption intent between the two models, although the GPT-4o model demonstrated higher recall rates in reducing sensitive information disclosure.
- The tool effectively helped users minimize unnecessary information sharing while maintaining output quality, particularly excelling in scenarios where manual minimization was complex and prone to oversight.
-
Limitations and Future Directions
- Limitations include: lower recall rates of small models, which may lead to missed sensitive information; and the predictive quality of abstraction results, which increases learning costs.
- Future research directions:
- Improve the detection accuracy of small LLMs, such as through fine-tuning or integrating results from multiple models.
- Enhance the continuity and consistency of abstraction outputs, enabling users to better understand and predict the tool's behavior.
- Further explore applicability in multilingual environments and across diverse domains.
Conclusion
The Rescriber system explores a user-driven data minimization approach and proposes a practical solution to privacy protection and trust issues in LLM-based models. The study demonstrates that small LLMs can serve as effective intermediary technologies for privacy management, while also playing a potential role in educating users about privacy awareness and improving privacy control.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can users effectively detect and reduce leakage of sensitive personal information when interacting with LLM-driven chatbots?Category: Trust, Transparency, and Response Latency DesignSimilar questionsarrow_forward
- Can small language models run locally on user devices to provide efficient privacy control without relying on cloud computing?Category: Trust, Transparency, and Response Latency DesignSimilar questionsarrow_forward
- How do replacement and abstraction de-identification methods compare in efficiency and user satisfaction across contexts?Category: Trust, Transparency, and Response Latency DesignSimilar questionsarrow_forward
Practical Problems
1- Users easily leak sensitive personal information when interacting with chatbots but lack convenient, effective privacy protection tools.Category: Trust, Transparency, and Response Latency DesignSimilar questionsarrow_forward
- 71%
Data-Centric Explanations: Explaining Training Data of Machine Learning Systems to Promote Transparency
CHI '21· Explainable AI (XAI) +2
- 71%
"Having Confidence in My Confidence Intervals": How Data Users Engage with Privacy-Protected Wikipedia Data
CHI '26· Explainable AI (XAI) +2
- 71%
TermSight: Making Service Contracts Approachable
CHI '26· Explainable AI (XAI) +2
- 71%
How Much Trust is Enough? Towards Calibrating Trust in Technology
CHI '26· Explainable AI (XAI) +2
- 71%
When the Codec Hallucinates: User Perceptions of Miscompressed Images
CHI '26· Explainable AI (XAI) +2
- 67%
Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations
CHI '23· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)