When is ML data good?: Valuing in Public Health Datafication
Authors
Title of the Paper
When is Machine Learning Data Good?: Valuing in Public Health Datafication
Paper Information
- Field of Study: Human-Computer Interaction (HCI), Machine Learning, Public Health Data Processing
- Keywords: Data work, India, public health, data value, machine learning data, ethics, data quality, data supply chain, data context, transparency and accountability
Research Background and Issues
-
Identified Problems or Challenges:
- The quality of machine learning models depends on the quality of the data. However, public health data often suffers from "data noise," incompleteness, duplication, and unreliable labels at various stages of collection, processing, and application.
- Lack of transparency and accountability in the data collection-to-application process leads to conflicts and differing perspectives among stakeholders.
- Data quality issues are embedded in the complex social, cultural, and geographical contexts of data collectors (e.g., grassroots health workers), while developers often lack sufficient contextual information to optimize models.
-
Importance of the Research:
- The rapid development of machine learning increasingly relies on large, high-quality datasets. Understanding how data quality evolves through different stages of work is critical for improving model quality, especially in high-risk domains like public health.
- Value conflicts within the data supply chain may hinder the potential of data-driven technologies, particularly in the data work ecosystems of the Global South (e.g., India).
-
Motivation and Related Work:
- Expands existing research on data quality, global health, and data work, particularly in understanding how the data supply chain evolves and how data is valued at different stages.
- Integrates academic frameworks such as "data value," "data ethics," and "transparency and accountability" to explore how public health data is transformed for machine learning modeling.
Solution
-
Methods or Solutions:
- The authors conducted interviews with 46 participants involved in the data supply chain of India's public health system, analyzing how data evolves from collection to machine learning model development.
- Using a value research perspective, they introduced the concept of "value attribution" within the supply chain, analyzing how data collectors (grassroots health workers), data administrators, and machine learning developers define "good" and "useful" data.
-
Innovative Aspects:
- Proposed a value attribution analysis tool based on the data supply chain framework to explore how data is "transformed" at each stage to suit specific purposes.
- Identified key issues of transparency and accountability within the data supply chain and suggested organizational-level improvements for data work.
-
Implementation Steps and Key Techniques:
- Data Collection: Semi-structured interviews with health workers, data administrators, and machine learning developers.
- Data Analysis: Qualitative analysis methods to identify common themes across roles and develop a value tension model for the data flow process.
- Case Study: In-depth examination of India's dynamic public health data environment, including regional languages, cultural contexts, and privacy concerns.
Research Outcomes
-
Specific Findings:
-
Through analysis of data work across different roles, the authors identified the following tensions:
- Data Quality: Developers prioritize structured, standardized high-quality data, while collectors face social, cultural, and geographical constraints, leading to missing or biased data.
- Data Context: Developers lack sufficient information to understand the context of data collection, while collectors are unaware of the ultimate purpose of the data.
- Organizational Incentives: Collectors are driven more by financial incentives and performance evaluations, which diverge from developers' technical and research goals.
-
Summarized how different roles attribute value to "good" data:
- Data Collectors: Data that fulfills administrative requirements and completes their tasks.
- Data Administrators: Smooth data that supports the integrity and synchronization of information management systems.
- Data Developers: Data suitable for model training and capable of achieving high predictive performance.
-
-
Advantages Compared to Existing Solutions:
- Combines the data supply chain perspective with machine learning system design, making data work more visible and aligned with optimization needs in real-world dynamic environments.
- Highlights the impact of social, cultural, and organizational structures on data quality, broadening the scope of technical solutions.
-
Experimental or Evaluation Results:
- Identified security risks, uncertainties in community trust-building, and challenges posed by multilingual environments during data collection.
- Data administrators and developers spent significant additional effort addressing data deficiencies during processing and model generation, underscoring the foundational importance of the data collection stage.
- Transparency issues in the data supply chain, revealed through interviews, reflect poor communication among different roles in data work.
-
Limitations and Future Directions:
-
Limitations:
- The complexity and regional specificity of the data work ecosystem may limit the applicability of findings to other countries or fields.
- The sample size of developer reports is small and restricted to a few collaborating organizations.
-
Future Directions:
- Recommend further research on transparency in the data supply chain, developing tools to ensure visibility and accountability at every stage of work.
- Explore ways to enhance data literacy and agency among data collectors, improving the integration of AI tools with grassroots work.
- Design more efficient and equitable organizational incentive schemes, aligning the motivations of health workers with the ultimate scientific goals.
-
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What constitutes high-quality machine learning data, and how can its value be analyzed through the data supply chain?Category: Machine Learning Model Visual AnalyticsSimilar questionsarrow_forward
- How is public health data transformed from collection to application, and how does this affect model quality?Category: Machine Learning Model Visual AnalyticsSimilar questionsarrow_forward
- How can transparency and accountability issues in the data supply chain be addressed to improve data quality?Category: Machine Learning Model Visual AnalyticsSimilar questionsarrow_forward
Practical Problems
1- Machine learning developers lack public health data collection context, and data quality and trust issues are prominent.Category: Machine Learning Model Visual AnalyticsSimilar questionsarrow_forward
- 60%
For What It's Worth: Humans Overwrite Their Economic Self-Interest to Avoid Bargaining With AI Systems
CHI '22· AI-Assisted Decision-Making & Automation +1
- 60%
Autonomous Assessment of Demonstration Sufficiency via Bayesian Inverse Reinforcement Learning
HRI '24· Teleoperated Driving +1
Based on Jaccard similarity of research subtopics & professions (≥60%)