Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
Honorable MentionAuthors
Research Background and Issues
- Issues and Challenges: The authors highlight the significant potential of third-party algorithm auditing in enhancing algorithm accountability. However, due to restricted data access, audits often rely on low-quality and incomplete data, which undermines the credibility of research. Auditors are unable to obtain the necessary sample sizes, key features, or model outputs, leading to substantial uncertainty in fairness assessments.
- Significance: As automated decision systems are increasingly applied in critical domains such as criminal justice, health insurance, and social welfare, their accountability and fairness are receiving growing attention. Existing technical and policy frameworks often fail to effectively ensure independent and thorough audits of these systems.
- Research Motivation: Despite regulations like the EU's Digital Services Act advocating for algorithm auditing, standardized methods and high-quality data for quantitative fairness assessments remain lacking. Privacy-preserving mechanisms such as differential privacy and synthetic data have been proposed as potential solutions, but their impact on audit outcomes has yet to be systematically evaluated.
Solution
-
Methods and Strategies: The authors investigate the impact of three data access scenarios on fairness metrics:
- Access to aggregated statistical data (e.g., confusion matrices).
- Access to individual-level data with model outputs.
- No model access but individual-level data available (via model reconstruction).
For each scenario, the authors simulate potential data quality issues (e.g., small sample sizes, missing features, imbalanced missing values, differential privacy, synthetic data) to quantify their effects on fairness assessments.
-
Innovations:
- Systematic comparison of the accuracy of fairness assessments under different data access strategies and varying levels of data quality degradation.
- Simulation experiments using real-world data (recidivism prediction in criminal justice and health insurance coverage prediction) to evaluate the trade-off between privacy protection and analytical reliability.
- Proposals for the applicability of differential privacy and data minimization principles in specific auditing tasks.
-
Implementation Steps and Techniques:
- Use two real-world datasets for algorithm simulations (NIJ recidivism prediction and ACS public health coverage).
- Build and train machine learning models for classification tasks (e.g., XGBoost).
- Design five specific experiments (sample reduction, feature omission, imbalanced missing values, differential privacy aggregation, synthetic data generation) to explore the impact of different data conditions on the reliability of fairness metrics.
- Quantify biases under various auditing conditions using fairness metrics such as statistical parity difference and average odds difference.
Research Findings
-
Specific Results:
- Aggregated Data Access (Scenario A): Differentially private confusion matrices can achieve highly accurate fairness assessments with sufficiently large sample sizes and appropriate privacy budgets. However, reliability significantly decreases when sample sizes are too small (e.g., fewer than 1,000).
- Access to Individual-Level Data (Scenarios B and C): Missing key features or incomplete data significantly reduces the accuracy of fairness metrics and may lead to misinterpretations (e.g., underestimating or misattributing bias).
- Synthetic Data: Synthetic data fails to reliably reflect group disparities present in real data, particularly obscuring injustices faced by disadvantaged groups, which results in a high proportion of Type II errors.
- Model Reconstruction (Scenario C): Auditing through model reconstruction is theoretically feasible but requires more high-quality data and computational resources than direct access to model outputs. Even then, data quality issues severely impact reliability.
- Comparison of Privacy Protection Methods: Differential privacy, when appropriately parameterized, can balance privacy protection and audit reliability, while the data minimization principle has minimal impact on fairness assessments when low-relevance features are removed.
-
Advantages Over Existing Methods: The study provides quantitative benchmarks and clarifies the strengths and weaknesses of conducting fairness audits under different data access and privacy protection mechanisms. It offers guidance for policymakers and technology developers.
-
Limitations: The research primarily focuses on binary classification tasks and binary group disparities, excluding multi-class classifiers or more complex fairness assessment scenarios. Additionally, the model reconstruction approach relies on idealized conditions, such as precise knowledge of model parameters.
-
Future Directions: Researchers are encouraged to explore data requirements for multi-class classifiers and intersectional fairness issues, develop more robust privacy-preserving mechanisms, and advance supportive frameworks for trustworthy third-party audits at the policy level.
Conclusion
This study identifies the issue of insufficient data access in current algorithm auditing practices and validates the effectiveness and limitations of different data access strategies and privacy-preserving technologies through simulation experiments. It emphasizes the importance of improving data access quality and flexibility while demonstrating that balancing privacy protection and fairness assessments is achievable.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do different data-access scenarios affect accuracy of algorithmic fairness evaluation?Category: Algorithm Contestability and Governance AuditingSimilar questionsarrow_forward
- How can privacy-preserving techniques such as differential privacy and data minimization balance privacy protection and analytic reliability in algorithm auditing?Category: Algorithm Contestability and Governance AuditingSimilar questionsarrow_forward
- What are specific impacts of low-quality and incomplete data on fairness-evaluation results?Category: Algorithm Contestability and Governance AuditingSimilar questionsarrow_forward
Practical Problems
1- Third-party algorithm auditing cannot conduct reliable fairness evaluation due to limited data access.Category: Algorithm Contestability and Governance AuditingSimilar questionsarrow_forward
- 86%
PASTA: A Scalable Framework for Multi-Policy AI Compliance Evaluation
CHI '26· Explainable AI (XAI) +3
- 86%
Why am I seeing this: Democratizing End User Auditing for Online Content Recommendations
UIST '25· Explainable AI (XAI) +2
- 83%
Expanding Explainability: Towards Social Transparency in AI systems
CHI '21· Explainable AI (XAI) +2
- 83%
HILL: A Hallucination Identifier for Large Language Models
CHI '24· Explainable AI (XAI) +2
- 71%
Exploring What People Need to Know to be AI Literate: Tailoring for a Diversity of AI Roles and Responsibilities
CHI '25· Explainable AI (XAI) +2
- 71%
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
CHI '25· Explainable AI (XAI) +2
- 71%
Decomposing Autonomy: Explaining AI Technology Acceptance Through a Liberty-Based Framework
CHI '26· Explainable AI (XAI) +2
- 71%
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
CHI '26· Explainable AI (XAI) +2
- 71%
Certified AI System = Trustworthy? Exploring Expert and Lay User Perceptions and Needs Regarding AI Certification
CHI '26· Explainable AI (XAI) +2
- 71%
When Life Gives You AI, Will You Turn It Into A Market for Lemons? Understanding How Information Asymmetries About AI System Capabilities Affect Market Outcomes and Adoption
CHI '26· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)