Explanatory Debiasing: Involving Domain Experts in the Data Generation Process to Mitigate Representation Bias in AI Systems

Honorable Mention
AI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasPhysicians, Nurses & CliniciansSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Research Background and Problem

  • Problems and Challenges:
    The authors identify representation bias as a common issue in artificial intelligence (AI) systems. Sparse samples of subgroups in training data lead to poor predictive performance of AI models for these groups. Furthermore, current debiasing methods often rely on data augmentation techniques, which may retain or even amplify existing biases in the original data and generate unrealistic or unreasonable samples.

  • Significance:
    Representation bias affects the accuracy, fairness, and generalizability of AI systems, particularly in high-stakes domains such as healthcare. Biased data can result in discriminatory outcomes in decision-making, such as dermatological diagnostic models performing poorly for patients with darker skin tones.

  • Research Motivation and Related Work:
    While existing studies propose data augmentation to enhance the representation of minority groups, these methods often fail to adequately address the issue due to a lack of domain knowledge. By involving domain experts, their expertise can be leveraged to identify data issues and ensure that generated samples align with real-world observations. However, research on the involvement of domain experts in the debiasing process remains limited, especially in terms of user studies, which constrains exploration in this area.


Solution

  • Proposed Approach:
    The authors propose a set of general design guidelines to integrate domain experts into the data augmentation and debiasing process. These guidelines are divided into three core stages:

    1. Pre-Augmentation Stage: Use intuitive data interpretation and analysis to help domain experts identify representation bias.
    2. During Augmentation Stage: Enable experts to participate in sample generation and provide constraints to guide data augmentation algorithms.
    3. Post-Augmentation Stage: Facilitate the screening and validation of generated data to ensure its realism and improve data quality.
  • Innovations:
    The authors combine targeted and interpretable user interaction methods with controlled data generation, systematically proposing specific guidelines for domain expert involvement in the debiasing process for the first time. Particularly in the healthcare domain, the approach integrates data-centric explanations and case-level "what-if exploration", emphasizing domain knowledge-driven interaction and validation.

  • Implementation Steps and Techniques:

    • Before debiasing, use interactive data visualization and model impact analysis to help experts understand how bias affects model performance.
    • During data augmentation, allow experts to define specific constraints based on variables (e.g., joint influence ranges of variables) and generate samples for subgroups.
    • After data augmentation, provide tools for sample screening and "what-if exploration", enabling domain experts to validate and optimize generated data.
    • Developed a healthcare application based on CTGAN generative adversarial networks to predict Type 2 diabetes.

Research Outcomes

  • Specific Results:

    1. Proposed general design guidelines applicable across multiple domains, with prototype implementation and validation in a healthcare scenario.
    2. User study results indicate that domain expert involvement can significantly reduce representation bias without compromising model performance.
    3. Designed an open-source code framework and implementation, facilitating developers to apply these methods more easily.
  • Advantages Compared to Existing Solutions:

    • Compared to traditional methods that solely rely on data augmentation algorithms, this approach significantly improves the realism of generated samples and the fairness of models by incorporating domain knowledge.
    • Expert involvement further reveals latent associations in the original training set and generated data (e.g., real-world relationships between features).
  • Experimental and Evaluation Results:

    1. Model Performance Improvement: Compared to default models and automated methods without domain knowledge, expert involvement improved model accuracy by an average of 2.1%.
    2. Reduction in Representation Bias: 31 experts effectively improved sample representativeness and coverage rates through interactive tools.
    3. Increased Trust from Domain Experts: Experimental results show that experts' trust in AI systems increased by approximately 7% after understanding the bias.
  • Limitations and Future Directions:

    1. No limit was set on the volume of generated data, which may lead to performance issues if too many samples are generated.
    2. The current user interface lacks global inspection capabilities for generated data. Future designs should enhance data visualization and automated data validation features.
    3. This study focuses on healthcare data; future work should validate its applicability in other domains (e.g., finance, regulations) and with unstructured data (e.g., images, text).

Conclusion:

By introducing domain experts and proposing systematic design guidelines, this study takes a significant step toward mitigating representation bias in AI systems. It contributes to the development of fair and transparent AI systems and lays the foundation for future research and practice.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188455/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713497
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
4 authors
sell
Subtopics
AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias
work
Professions
Physicians, Nurses & Clinicians, Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers