"Having Confidence in My Confidence Intervals": How Data Users Engage with Privacy-Protected Wikipedia Data
Authors
Paper Title
'Having Confidence in My Confidence Intervals': How Data Users Engage with Privacy-Protected Wikipedia Data
Publication Info
- Topic area: Data user engagement with privacy-preserving datasets.
- Keywords: Differential privacy, rounding, privacy-preserving data, Wikipedia, data utility, confidence intervals, data documentation, privacy-utility tradeoff.
Background and Problem
- Problem / challenge: Privacy-preserving techniques like differential privacy (DP) and rounding introduce noise into datasets, but little is known about how data users perceive and interact with such noise or how it impacts their analyses.
- Significance: Understanding these interactions is critical for designing effective documentation and tools that enable data users to work with privacy-protected data while maintaining trust and utility.
- Motivation and related work: Previous studies have focused on privacy guarantees for data subjects and usability for data curators but have largely ignored the experiences of data users. High-profile data releases like the 2020 US Census and Facebook’s Social Science One project have highlighted tensions between privacy protections and data utility, underscoring the need for better communication and support for data users.
Solution
- Proposed approach: A task-based contextual inquiry exploring how data users engage with two privacy-preserving Wikipedia datasets: one using rounding and another using differential privacy.
- Novelty:
- Empirically-driven documentation tailored to privacy-preserving datasets.
- Comparative analysis of user engagement with heuristic (rounding) and formal (DP) privacy methods.
- Identification of misconceptions about privacy-utility tradeoffs and challenges in computing confidence intervals.
- Design recommendations for improving documentation and tools for privacy-noised datasets.
- Procedure and key techniques:
- Developed documentation for datasets with expert feedback.
- Conducted a study with 15 data science practitioners using task-based contextual inquiry and semi-structured interviews.
- Analyzed participants’ interactions with datasets and their ability to compute confidence intervals and interpret privacy noise.
Results
- Concrete findings:
- Participants found rounding easier to understand but less useful for analysis due to its lack of statistical properties.
- DP data enabled better simulation-based approaches for uncertainty estimation but was initially harder to grasp.
- Most participants struggled to compute confidence intervals across multiple noisy data points, particularly for aggregated data.
- Some participants incorrectly associated higher utility of DP with weaker privacy protections.
- Advantage over baselines: DP was perceived as more analytically tractable and statistically rigorous than rounding, particularly for tasks requiring aggregation or confidence interval estimation.
- Experiments / evaluation:
- Datasets: Wikipedia pageview data (global, rounded, DP).
- Tasks: Confidence interval estimation and likelihood calculations for noisy data.
- Metrics: Participant accuracy, task completion strategies, and qualitative feedback.
- Limitations and future work:
- Limited generalizability due to the specific context of Wikipedia datasets.
- Focus on privacy and utility may have biased participant feedback.
- Future work should explore broader audiences, develop tools for confidence interval computation, and investigate perceptions of noisy data in other contexts.
Summary
This study investigates how data users engage with privacy-preserving Wikipedia datasets using rounding and differential privacy (DP). Participants found DP more suitable for statistical analysis due to its unbiasedness and Gaussian noise properties, though it required more effort to understand. Rounding was simpler but less useful for precise analyses. Many participants struggled with confidence interval calculations across multiple noisy data points and held misconceptions about the relationship between privacy and accuracy. The findings highlight the need for improved documentation, tools for uncertainty estimation, and better communication of privacy-utility tradeoffs. These insights can guide the design of future privacy-preserving data releases.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive ML
CHI '20· Explainable AI (XAI) +2
- 71%
Data-Centric Explanations: Explaining Training Data of Machine Learning Systems to Promote Transparency
CHI '21· Explainable AI (XAI) +2
- 71%
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
CHI '25· Explainable AI (XAI) +2
- 71%
TSEditor: Interactive Time Series Editing for Privacy Preservation
CHI '26· Privacy Perception & Decision-Making +2
- 71%
How Much Trust is Enough? Towards Calibrating Trust in Technology
CHI '26· Explainable AI (XAI) +2
- 71%
When the Codec Hallucinates: User Perceptions of Miscompressed Images
CHI '26· Explainable AI (XAI) +2
- 71%
Designing Effective Training Dataset Explanations: The Impact of Information Depth and Progressive Disclosure
IUI '26· Explainable AI (XAI) +2
- 67%
The Digital Landscape of Nudging: A Systematic Literature Review of Empirical Research on Digital Nudges
CHI '22· Privacy by Design & User Control +1
- 67%
Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations
CHI '23· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)