What is Sensitive About (Sensitive) Data? Characterizing Sensitivity and Intimacy of Google Assistant Speech Records
Authors
Title of the Paper
What is Sensitive About (Sensitive) Data? Characterizing Sensitivity and Intimacy with Google Assistant Users
Paper Information
- Research Area: Human-Computer Interaction and Privacy Protection
- Keywords: voice assistants, personal data, sensitive data, intimate data, data privacy, data donation, Google Assistant, privacy norms, harmful data collection, unawareness
Research Background and Problem
-
What issues or challenges did the authors identify?
Voice assistants can store potentially sensitive or intimate data by collecting daily voice interactions. However, users, privacy experts, and regulatory bodies differ in their understanding and definition of data sensitivity. While users often perceive single interaction data as non-sensitive, aggregated data over time may reveal private information. Additionally, users have limited understanding of data collection and inference, leading to insufficient awareness of potential privacy risks. -
Why is this issue important?
As voice assistants and other digital devices increasingly integrate into daily life, they ubiquitously collect behavioral and interaction data. Understanding the definition of sensitive information and users' perceptions of it can help researchers, industry professionals, and policymakers better protect privacy. Striking a balance between data donation and privacy protection is particularly critical. -
Research Motivation and Related Work
Existing studies primarily focus on single interaction data from voice assistants, showing that users deem such data non-sensitive. However, these studies overlook the ability to aggregate data and infer information, failing to comprehensively reveal users' understanding of data sensitivity. This study aims to present a holistic view of voice interaction data, including single interactions, patterns, and inferred information, to deeply explore sensitivity and intimacy.
Solution
-
What methods or solutions did the authors propose?
The authors designed a series of scenario-based interview protocols to explore the sensitivity and intimacy of voice assistant recordings through user data donation and interaction.- Data Donation: Collected 8,735 voice recordings from 22 Google Assistant users.
- Data Analysis: Classified and annotated voice recordings to identify their sensitivity and inference capabilities.
- Data Visualization: Designed a personal data canvas for structured interaction and data presentation.
- Interviews: Conducted semi-structured interviews with 17 data donors to understand their perceptions of data sensitivity and intimacy.
-
What are the innovative aspects of the solution?
- Redefined sensitive data and intimate data.
- Introduced a sensitivity spectrum to measure variations in data sensitivity (low, medium, high).
- Developed a comprehensive method to examine data, assessing sensitivity based on single and aggregated interactions.
- Designed an interactive "personal data canvas" to help users explore and interpret data sensitivity.
-
What are the implementation steps and key technologies used?
- Data Collection: Obtained voice recordings via the Google Takeout platform.
- Data Classification: Categorized and annotated data based on interaction types (e.g., information requests, transactional requests) and sensitivity levels.
- Data Visualization: Created the "personal data canvas" to visualize data patterns and temporal distribution.
- Data Interviews: Used scenario-based questionnaires to evaluate users' data-sharing capabilities and sensitivity assessments.
- Qualitative Analysis: Applied Reflexive Thematic Analysis to analyze interview results.
Research Findings
-
What specific findings were achieved?
- Discovered that users often experience a "turning point" in their perception of sensitivity when exploring data deeply, leading to significant changes in their understanding of data sensitivity and intimacy.
- Identified three main characteristics of sensitive data: intrusiveness, specificity, and (un)availability.
- Defined intimate data as a subset of sensitive data, primarily related to private activities and spaces recorded.
- Found that sensitive data exhibits characteristics of contextual incongruity when shared.
- Quantitative analysis revealed sensitivity and intimacy scores, showing that unexpected interactions or inferred data are perceived as highly sensitive.
-
What advantages does it have compared to existing solutions?
- Expanded upon existing research focused on single interaction data by comprehensively exploring aggregated data, inference, and intimacy characteristics.
- Introduced tools for personal data exploration, enhancing users' awareness of data privacy.
- Emphasized the uncertainty in defining sensitive and intimate data, offering a new perspective for dynamic privacy boundary settings.
-
What are the experimental or evaluation results?
Interviews revealed that Google Assistant users have limited understanding of data sensitivity and privacy boundaries. However, intuitive data visualization and interaction significantly improved users' sensitivity awareness. -
Limitations and Future Directions
-
Limitations:
- The sample primarily represents Western cultures, lacking diverse global perspectives on privacy.
- Data focuses on voice assistant recordings, excluding other types of personal data.
- Some participants had prior knowledge of voice data sensitivity, potentially reducing the study's generalizability.
-
Future Directions:
- Explore sensitivity and intimacy in other types of data (e.g., mobile location data, sensor data).
- Investigate differences in privacy boundaries and sensitivity perceptions across cultural contexts.
- Design more intelligent and intuitive data visualization tools to enhance users' privacy awareness and boundary-setting capabilities.
-
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do users understand the concepts of sensitive data and intimate data recorded by voice assistants?Category: Privacy, Informed Consent, and TransparencySimilar questionsarrow_forward
- How does data sensitivity change between single data interactions and aggregated data?Category: Privacy, Informed Consent, and TransparencySimilar questionsarrow_forward
- How do data visualization tools affect users' perceptions of data sensitivity and privacy boundaries?Category: Privacy, Informed Consent, and TransparencySimilar questionsarrow_forward
Practical Problems
1- Users cannot accurately understand the sensitivity of data collected by voice assistants and struggle to set privacy boundaries.Category: Privacy, Informed Consent, and TransparencySimilar questionsarrow_forward
- 71%
A Psychometric Scale to Measure Individuals' Value of Other People's Privacy (VOPP)
CHI '23· AI Ethics, Fairness & Accountability +2
- 63%
Data-Centric Explanations: Explaining Training Data of Machine Learning Systems to Promote Transparency
CHI '21· Explainable AI (XAI) +2
- 63%
Understanding and Mitigating Technology-Facilitated Privacy Violations in the Physical World
CHI '23· Privacy by Design & User Control +2
- 63%
A World Full of Privacy and Security (Mis)conceptions? Findings of a Representative Survey in 12 Countries
CHI '23· Privacy by Design & User Control +2
- 63%
Deepfakes, Phrenology, Surveillance, and More! A Taxonomy of AI Privacy Risks
CHI '24· AI Ethics, Fairness & Accountability +2
- 63%
Encoding Privacy: Sociotechnical Dynamics of Data Protection Compliance Work
CHI '24· AI Ethics, Fairness & Accountability +2
- 63%
Exploring What People Need to Know to be AI Literate: Tailoring for a Diversity of AI Roles and Responsibilities
CHI '25· Explainable AI (XAI) +2
- 63%
TermSight: Making Service Contracts Approachable
CHI '26· Explainable AI (XAI) +2
- 63%
When Feasibility of Fairness Audits Relies on Willingness to Share Data: Examining User Acceptance of Multi-Party Computation Protocols for Fairness Monitoring
CHI '26· AI Ethics, Fairness & Accountability +2
- 63%
Decomposing Autonomy: Explaining AI Technology Acceptance Through a Liberty-Based Framework
CHI '26· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)