What is Sensitive About (Sensitive) Data? Characterizing Sensitivity and Intimacy of Google Assistant Speech Records

Explainable AI (XAI)AI Ethics, Fairness & AccountabilityPrivacy by Design & User ControlPrivacy Perception & Decision-MakingCybersecurity EngineersPrivacy Policy MakersHCI Researchers

Title of the Paper

What is Sensitive About (Sensitive) Data? Characterizing Sensitivity and Intimacy with Google Assistant Users

Paper Information

  • Research Area: Human-Computer Interaction and Privacy Protection
  • Keywords: voice assistants, personal data, sensitive data, intimate data, data privacy, data donation, Google Assistant, privacy norms, harmful data collection, unawareness

Research Background and Problem

  • What issues or challenges did the authors identify?
    Voice assistants can store potentially sensitive or intimate data by collecting daily voice interactions. However, users, privacy experts, and regulatory bodies differ in their understanding and definition of data sensitivity. While users often perceive single interaction data as non-sensitive, aggregated data over time may reveal private information. Additionally, users have limited understanding of data collection and inference, leading to insufficient awareness of potential privacy risks.

  • Why is this issue important?
    As voice assistants and other digital devices increasingly integrate into daily life, they ubiquitously collect behavioral and interaction data. Understanding the definition of sensitive information and users' perceptions of it can help researchers, industry professionals, and policymakers better protect privacy. Striking a balance between data donation and privacy protection is particularly critical.

  • Research Motivation and Related Work
    Existing studies primarily focus on single interaction data from voice assistants, showing that users deem such data non-sensitive. However, these studies overlook the ability to aggregate data and infer information, failing to comprehensively reveal users' understanding of data sensitivity. This study aims to present a holistic view of voice interaction data, including single interactions, patterns, and inferred information, to deeply explore sensitivity and intimacy.

Solution

  • What methods or solutions did the authors propose?
    The authors designed a series of scenario-based interview protocols to explore the sensitivity and intimacy of voice assistant recordings through user data donation and interaction.

    • Data Donation: Collected 8,735 voice recordings from 22 Google Assistant users.
    • Data Analysis: Classified and annotated voice recordings to identify their sensitivity and inference capabilities.
    • Data Visualization: Designed a personal data canvas for structured interaction and data presentation.
    • Interviews: Conducted semi-structured interviews with 17 data donors to understand their perceptions of data sensitivity and intimacy.
  • What are the innovative aspects of the solution?

    • Redefined sensitive data and intimate data.
    • Introduced a sensitivity spectrum to measure variations in data sensitivity (low, medium, high).
    • Developed a comprehensive method to examine data, assessing sensitivity based on single and aggregated interactions.
    • Designed an interactive "personal data canvas" to help users explore and interpret data sensitivity.
  • What are the implementation steps and key technologies used?

    1. Data Collection: Obtained voice recordings via the Google Takeout platform.
    2. Data Classification: Categorized and annotated data based on interaction types (e.g., information requests, transactional requests) and sensitivity levels.
    3. Data Visualization: Created the "personal data canvas" to visualize data patterns and temporal distribution.
    4. Data Interviews: Used scenario-based questionnaires to evaluate users' data-sharing capabilities and sensitivity assessments.
    5. Qualitative Analysis: Applied Reflexive Thematic Analysis to analyze interview results.

Research Findings

  • What specific findings were achieved?

    • Discovered that users often experience a "turning point" in their perception of sensitivity when exploring data deeply, leading to significant changes in their understanding of data sensitivity and intimacy.
    • Identified three main characteristics of sensitive data: intrusiveness, specificity, and (un)availability.
    • Defined intimate data as a subset of sensitive data, primarily related to private activities and spaces recorded.
    • Found that sensitive data exhibits characteristics of contextual incongruity when shared.
    • Quantitative analysis revealed sensitivity and intimacy scores, showing that unexpected interactions or inferred data are perceived as highly sensitive.
  • What advantages does it have compared to existing solutions?

    • Expanded upon existing research focused on single interaction data by comprehensively exploring aggregated data, inference, and intimacy characteristics.
    • Introduced tools for personal data exploration, enhancing users' awareness of data privacy.
    • Emphasized the uncertainty in defining sensitive and intimate data, offering a new perspective for dynamic privacy boundary settings.
  • What are the experimental or evaluation results?
    Interviews revealed that Google Assistant users have limited understanding of data sensitivity and privacy boundaries. However, intuitive data visualization and interaction significantly improved users' sensitivity awareness.

  • Limitations and Future Directions

    • Limitations:

      • The sample primarily represents Western cultures, lacking diverse global perspectives on privacy.
      • Data focuses on voice assistant recordings, excluding other types of personal data.
      • Some participants had prior knowledge of voice data sensitivity, potentially reducing the study's generalizability.
    • Future Directions:

      • Explore sensitivity and intimacy in other types of data (e.g., mobile location data, sensor data).
      • Investigate differences in privacy boundaries and sensitivity perceptions across cultural contexts.
      • Design more intelligent and intuitive data visualization tools to enhance users' privacy awareness and boundary-setting capabilities.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96284/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581164
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), AI Ethics, Fairness & Accountability, Privacy by Design & User Control, Privacy Perception & Decision-Making
work
Professions
Cybersecurity Engineers, Privacy Policy Makers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers