Explanatory Debiasing: Involving Domain Experts in the Data Generation Process to Mitigate Representation Bias in AI Systems
Honorable MentionAuthors
Research Background and Problem
-
Problems and Challenges:
The authors identify representation bias as a common issue in artificial intelligence (AI) systems. Sparse samples of subgroups in training data lead to poor predictive performance of AI models for these groups. Furthermore, current debiasing methods often rely on data augmentation techniques, which may retain or even amplify existing biases in the original data and generate unrealistic or unreasonable samples. -
Significance:
Representation bias affects the accuracy, fairness, and generalizability of AI systems, particularly in high-stakes domains such as healthcare. Biased data can result in discriminatory outcomes in decision-making, such as dermatological diagnostic models performing poorly for patients with darker skin tones. -
Research Motivation and Related Work:
While existing studies propose data augmentation to enhance the representation of minority groups, these methods often fail to adequately address the issue due to a lack of domain knowledge. By involving domain experts, their expertise can be leveraged to identify data issues and ensure that generated samples align with real-world observations. However, research on the involvement of domain experts in the debiasing process remains limited, especially in terms of user studies, which constrains exploration in this area.
Solution
-
Proposed Approach:
The authors propose a set of general design guidelines to integrate domain experts into the data augmentation and debiasing process. These guidelines are divided into three core stages:- Pre-Augmentation Stage: Use intuitive data interpretation and analysis to help domain experts identify representation bias.
- During Augmentation Stage: Enable experts to participate in sample generation and provide constraints to guide data augmentation algorithms.
- Post-Augmentation Stage: Facilitate the screening and validation of generated data to ensure its realism and improve data quality.
-
Innovations:
The authors combine targeted and interpretable user interaction methods with controlled data generation, systematically proposing specific guidelines for domain expert involvement in the debiasing process for the first time. Particularly in the healthcare domain, the approach integrates data-centric explanations and case-level "what-if exploration", emphasizing domain knowledge-driven interaction and validation. -
Implementation Steps and Techniques:
- Before debiasing, use interactive data visualization and model impact analysis to help experts understand how bias affects model performance.
- During data augmentation, allow experts to define specific constraints based on variables (e.g., joint influence ranges of variables) and generate samples for subgroups.
- After data augmentation, provide tools for sample screening and "what-if exploration", enabling domain experts to validate and optimize generated data.
- Developed a healthcare application based on CTGAN generative adversarial networks to predict Type 2 diabetes.
Research Outcomes
-
Specific Results:
- Proposed general design guidelines applicable across multiple domains, with prototype implementation and validation in a healthcare scenario.
- User study results indicate that domain expert involvement can significantly reduce representation bias without compromising model performance.
- Designed an open-source code framework and implementation, facilitating developers to apply these methods more easily.
-
Advantages Compared to Existing Solutions:
- Compared to traditional methods that solely rely on data augmentation algorithms, this approach significantly improves the realism of generated samples and the fairness of models by incorporating domain knowledge.
- Expert involvement further reveals latent associations in the original training set and generated data (e.g., real-world relationships between features).
-
Experimental and Evaluation Results:
- Model Performance Improvement: Compared to default models and automated methods without domain knowledge, expert involvement improved model accuracy by an average of 2.1%.
- Reduction in Representation Bias: 31 experts effectively improved sample representativeness and coverage rates through interactive tools.
- Increased Trust from Domain Experts: Experimental results show that experts' trust in AI systems increased by approximately 7% after understanding the bias.
-
Limitations and Future Directions:
- No limit was set on the volume of generated data, which may lead to performance issues if too many samples are generated.
- The current user interface lacks global inspection capabilities for generated data. Future designs should enhance data visualization and automated data validation features.
- This study focuses on healthcare data; future work should validate its applicability in other domains (e.g., finance, regulations) and with unstructured data (e.g., images, text).
Conclusion:
By introducing domain experts and proposing systematic design guidelines, this study takes a significant step toward mitigating representation bias in AI systems. It contributes to the development of fair and transparent AI systems and lays the foundation for future research and practice.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can domain experts be effectively involved during data augmentation and bias mitigation to reduce representational bias in AI systems?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- What guidance approaches help domain experts better identify, generate, and validate data to improve data quality and model fairness?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- In high-risk domains (e.g., healthcare), how does a domain-expert-involved bias mitigation design framework affect AI model performance and fairness?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
Practical Problems
1- AI models often perform poorly on minority groups, affecting fairness and reliability.Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- 71%
Are We Automating the Joy Out of Work? Designing AI to Augment Work, Not Meaning
CHI '26· AI-Assisted Decision-Making & Automation +2
- 67%
Towards Fairness in Practice: A Practitioner-Oriented Rubric for Evaluating Fair ML Toolkits
CHI '21· AI Ethics, Fairness & Accountability +1
- 67%
Jury Learning: Integrating Dissenting Voices into Machine Learning Models
CHI '22· AI Ethics, Fairness & Accountability +1
- 67%
Capable but Amoral? Comparing AI and Human Expert Collaboration in Ethical Decision Making
CHI '22· AI Ethics, Fairness & Accountability +1
- 67%
Out of Context: Investigating the Bias and Fairness Concerns of "Artificial Intelligence as a Service"
CHI '23· AI Ethics, Fairness & Accountability +1
- 67%
“It is currently hodgepodge”: Examining AI/ML Practitioners’ Challenges during Co-production of Responsible AI Values
CHI '23· AI Ethics, Fairness & Accountability +1
- 67%
STILE: Exploring and Debugging Social Biases in Pre-trained Text Representations
CHI '24· AI Ethics, Fairness & Accountability +1
- 67%
Uncovering Bias in Personal Informatics
UbiComp '23· AI Ethics, Fairness & Accountability +1
- 63%
From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms to Foster Dignified Human-AI Interaction
CHI '26· AI-Assisted Decision-Making & Automation +3
Based on Jaccard similarity of research subtopics & professions (≥60%)