When is ML data good?: Valuing in Public Health Datafication

Teleoperated DrivingAI-Assisted Decision-Making & AutomationAI Ethics, Fairness & AccountabilityCommunity Health WorkersAI/ML Researchers & Engineers

Title of the Paper

When is Machine Learning Data Good?: Valuing in Public Health Datafication

Paper Information

  • Field of Study: Human-Computer Interaction (HCI), Machine Learning, Public Health Data Processing
  • Keywords: Data work, India, public health, data value, machine learning data, ethics, data quality, data supply chain, data context, transparency and accountability

Research Background and Issues

  • Identified Problems or Challenges:

    • The quality of machine learning models depends on the quality of the data. However, public health data often suffers from "data noise," incompleteness, duplication, and unreliable labels at various stages of collection, processing, and application.
    • Lack of transparency and accountability in the data collection-to-application process leads to conflicts and differing perspectives among stakeholders.
    • Data quality issues are embedded in the complex social, cultural, and geographical contexts of data collectors (e.g., grassroots health workers), while developers often lack sufficient contextual information to optimize models.
  • Importance of the Research:

    • The rapid development of machine learning increasingly relies on large, high-quality datasets. Understanding how data quality evolves through different stages of work is critical for improving model quality, especially in high-risk domains like public health.
    • Value conflicts within the data supply chain may hinder the potential of data-driven technologies, particularly in the data work ecosystems of the Global South (e.g., India).
  • Motivation and Related Work:

    • Expands existing research on data quality, global health, and data work, particularly in understanding how the data supply chain evolves and how data is valued at different stages.
    • Integrates academic frameworks such as "data value," "data ethics," and "transparency and accountability" to explore how public health data is transformed for machine learning modeling.

Solution

  • Methods or Solutions:

    • The authors conducted interviews with 46 participants involved in the data supply chain of India's public health system, analyzing how data evolves from collection to machine learning model development.
    • Using a value research perspective, they introduced the concept of "value attribution" within the supply chain, analyzing how data collectors (grassroots health workers), data administrators, and machine learning developers define "good" and "useful" data.
  • Innovative Aspects:

    • Proposed a value attribution analysis tool based on the data supply chain framework to explore how data is "transformed" at each stage to suit specific purposes.
    • Identified key issues of transparency and accountability within the data supply chain and suggested organizational-level improvements for data work.
  • Implementation Steps and Key Techniques:

    • Data Collection: Semi-structured interviews with health workers, data administrators, and machine learning developers.
    • Data Analysis: Qualitative analysis methods to identify common themes across roles and develop a value tension model for the data flow process.
    • Case Study: In-depth examination of India's dynamic public health data environment, including regional languages, cultural contexts, and privacy concerns.

Research Outcomes

  • Specific Findings:

    • Through analysis of data work across different roles, the authors identified the following tensions:

      • Data Quality: Developers prioritize structured, standardized high-quality data, while collectors face social, cultural, and geographical constraints, leading to missing or biased data.
      • Data Context: Developers lack sufficient information to understand the context of data collection, while collectors are unaware of the ultimate purpose of the data.
      • Organizational Incentives: Collectors are driven more by financial incentives and performance evaluations, which diverge from developers' technical and research goals.
    • Summarized how different roles attribute value to "good" data:

      • Data Collectors: Data that fulfills administrative requirements and completes their tasks.
      • Data Administrators: Smooth data that supports the integrity and synchronization of information management systems.
      • Data Developers: Data suitable for model training and capable of achieving high predictive performance.
  • Advantages Compared to Existing Solutions:

    • Combines the data supply chain perspective with machine learning system design, making data work more visible and aligned with optimization needs in real-world dynamic environments.
    • Highlights the impact of social, cultural, and organizational structures on data quality, broadening the scope of technical solutions.
  • Experimental or Evaluation Results:

    • Identified security risks, uncertainties in community trust-building, and challenges posed by multilingual environments during data collection.
    • Data administrators and developers spent significant additional effort addressing data deficiencies during processing and model generation, underscoring the foundational importance of the data collection stage.
    • Transparency issues in the data supply chain, revealed through interviews, reflect poor communication among different roles in data work.
  • Limitations and Future Directions:

    • Limitations:

      • The complexity and regional specificity of the data work ecosystem may limit the applicability of findings to other countries or fields.
      • The sample size of developer reports is small and restricted to a few collaborating organizations.
    • Future Directions:

      • Recommend further research on transparency in the data supply chain, developing tools to ensure visibility and accountability at every stage of work.
      • Explore ways to enhance data literacy and agency among data collectors, improving the integration of AI tools with grassroots work.
      • Design more efficient and equitable organizational incentive schemes, aligning the motivations of health workers with the ultimate scientific goals.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68973/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501868
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Teleoperated Driving, AI-Assisted Decision-Making & Automation, AI Ethics, Fairness & Accountability
work
Professions
Community Health Workers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
2 related papers