"Having Confidence in My Confidence Intervals": How Data Users Engage with Privacy-Protected Wikipedia Data

Explainable AI (XAI)Privacy by Design & User ControlPrivacy Perception & Decision-MakingData Scientists & AnalystsAI/ML Researchers & EngineersHCI Researchers

Paper Title

'Having Confidence in My Confidence Intervals': How Data Users Engage with Privacy-Protected Wikipedia Data

Publication Info

  • Topic area: Data user engagement with privacy-preserving datasets.
  • Keywords: Differential privacy, rounding, privacy-preserving data, Wikipedia, data utility, confidence intervals, data documentation, privacy-utility tradeoff.

Background and Problem

  • Problem / challenge: Privacy-preserving techniques like differential privacy (DP) and rounding introduce noise into datasets, but little is known about how data users perceive and interact with such noise or how it impacts their analyses.
  • Significance: Understanding these interactions is critical for designing effective documentation and tools that enable data users to work with privacy-protected data while maintaining trust and utility.
  • Motivation and related work: Previous studies have focused on privacy guarantees for data subjects and usability for data curators but have largely ignored the experiences of data users. High-profile data releases like the 2020 US Census and Facebook’s Social Science One project have highlighted tensions between privacy protections and data utility, underscoring the need for better communication and support for data users.

Solution

  • Proposed approach: A task-based contextual inquiry exploring how data users engage with two privacy-preserving Wikipedia datasets: one using rounding and another using differential privacy.
  • Novelty:
    1. Empirically-driven documentation tailored to privacy-preserving datasets.
    2. Comparative analysis of user engagement with heuristic (rounding) and formal (DP) privacy methods.
    3. Identification of misconceptions about privacy-utility tradeoffs and challenges in computing confidence intervals.
    4. Design recommendations for improving documentation and tools for privacy-noised datasets.
  • Procedure and key techniques:
    • Developed documentation for datasets with expert feedback.
    • Conducted a study with 15 data science practitioners using task-based contextual inquiry and semi-structured interviews.
    • Analyzed participants’ interactions with datasets and their ability to compute confidence intervals and interpret privacy noise.

Results

  • Concrete findings:
    • Participants found rounding easier to understand but less useful for analysis due to its lack of statistical properties.
    • DP data enabled better simulation-based approaches for uncertainty estimation but was initially harder to grasp.
    • Most participants struggled to compute confidence intervals across multiple noisy data points, particularly for aggregated data.
    • Some participants incorrectly associated higher utility of DP with weaker privacy protections.
  • Advantage over baselines: DP was perceived as more analytically tractable and statistically rigorous than rounding, particularly for tasks requiring aggregation or confidence interval estimation.
  • Experiments / evaluation:
    • Datasets: Wikipedia pageview data (global, rounded, DP).
    • Tasks: Confidence interval estimation and likelihood calculations for noisy data.
    • Metrics: Participant accuracy, task completion strategies, and qualitative feedback.
  • Limitations and future work:
    • Limited generalizability due to the specific context of Wikipedia datasets.
    • Focus on privacy and utility may have biased participant feedback.
    • Future work should explore broader audiences, develop tools for confidence interval computation, and investigate perceptions of noisy data in other contexts.

Summary

This study investigates how data users engage with privacy-preserving Wikipedia datasets using rounding and differential privacy (DP). Participants found DP more suitable for statistical analysis due to its unbiasedness and Gaussian noise properties, though it required more effort to understand. Rounding was simpler but less useful for precise analyses. Many participants struggled with confidence interval calculations across multiple noisy data points and held misconceptions about the relationship between privacy and accuracy. The findings highlight the need for improved documentation, tools for uncertainty estimation, and better communication of privacy-utility tradeoffs. These insights can guide the design of future privacy-preserving data releases.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/221917/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791573
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Explainable AI (XAI), Privacy by Design & User Control, Privacy Perception & Decision-Making
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers