Forgetting Practices in the Data Sciences

Honorable Mention
Explainable AI (XAI)Algorithmic Fairness & BiasSoftware Engineers & DevelopersAI/ML Researchers & EngineersStatisticians & Data Scientists

Title of the Paper

Forgetting Practices in the Data Sciences

Paper Information

  • Subject Area: Forgetting Practices in Data Science and Their Impact on Human-Computer Interaction (HCI)
  • Keywords: datasets, neural networks, gaze detection, text annotation, forgetting practices, data silence, machine learning, data science, interaction design, social justice

Research Background and Issues

  • Identified Problems or Challenges:

    1. The processes of dataset selection, processing, cleaning, and annotation in data science involve "forgetting practices," which may lead to data bias or unfairness.
    2. While many studies have explored dataset bias, less attention has been given to the intrinsic biases in data handling and their resulting impacts.
    3. Big data technologies often assume data objectivity, overlooking the human decisions and manipulations behind them.
  • Significance:

    1. Data plays a critical role in social decision-making, algorithm development, and machine learning model training, but neglecting certain details in data handling can lead to unintended consequences.
    2. Current data science practices lack documentation and examination of forgetting behaviors, potentially hindering goals of fairness and transparency.
  • Research Motivation and Related Work:

    1. The study is motivated by a deep reflection on the issue of data forgetting, exploring how to build selective forgetting mechanisms while preserving and recording critical data.
    2. This research integrates HCI theories, analyses of data science practices, critical computing perspectives, and insights from gender studies and social justice.

Proposed Solutions

  • Proposed Solutions:

    1. Classification and Analysis: The authors propose a classification system for "data silence," categorizing forgetting practices into different levels and types (e.g., mild silence, strong silence, complex silence).
    2. Forgetting Practices Layered Model: This model describes how forgetting accumulates throughout the data science lifecycle (from data planning to model deployment).
    3. Academic Vocabulary and Framework: A new terminology system is introduced to incorporate forgetting, memory, and deletion into the analysis.
  • Innovations:

    1. The "Forgetting Practices Layered Model" reveals specific forgetting processes in data science work through a hierarchical structure.
    2. Enriches the social justice dimension of data science, particularly in discussions on ethics related to sensitive data handling, feature engineering, and model training.
    3. Provides new linguistic tools (e.g., "data silence," "communities of forgetters," "memory aids") to interpret forgetting phenomena in data science.
  • Implementation Steps and Key Techniques:

    1. Data Classification: Evaluate implicit practices of forgetting during data selection, cleaning, annotation, and model training.
    2. Model Evaluation: Experimentally verify whether forgetting behaviors affect model accuracy and fairness.
    3. Tool Design: Propose socio-technical tools (e.g., data frameworks with change history tracking functionality).

Research Outcomes

  • Specific Outcomes:

    1. Developed a classification system (types of data silence) to analyze how forgetting mechanisms in data science work influence final outcomes.
    2. Provided detailed descriptions of the locations and patterns of forgetting practices in data science workflows (Forgetting Layers).
    3. Suggested directions for tool design and improvement to address the issue of data silence in dynamic data handling.
  • Advantages:

    1. Analyzes the sources of implicit bias in data science from a finer-grained perspective.
    2. Enriches and extends the ethical dimensions of data science, guiding attention toward transparency and social impact in data processing.
    3. Offers concrete recommendations and practical models for academia and practitioners.
  • Experimental or Evaluation Results: Experiments and analyses demonstrate how various forgetting phenomena during data processing (e.g., missing value substitution, label disagreement handling) evolve into implicit data biases.

  • Limitations and Future Directions:

    1. Limitations: The study primarily focuses on human decision-making in data science practices, without delving deeply into algorithmic biases and their feedback effects.
    2. Future Directions:
      • Develop tools to support the documentation and reminders of forgetting practices.
      • Explore the impact of forgetting practices in AI models on real-world decision-making.
      • Expand the research scope to include the socio-cultural influences on collaboration within data science teams.

Summary and Contributions

This study not only aims to uncover forgetting practices in data science but also helps researchers systematically understand the causes and impacts of forgetting behaviors through classification and layered models. Furthermore, it proposes pathways to improve transparency and fairness in data science workflows.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/71950/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3517644
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Honorable Mention
group
Authors
2 authors
sell
Subtopics
Explainable AI (XAI), Algorithmic Fairness & Bias
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
4 related papers