Compliant But Unsatisfactory: The Gap Between Auditing Standards and Practices for Probabilistic Genotyping Software

Explainable AI (XAI)Algorithmic Transparency & AuditabilityPrivacy by Design & User ControlAI/ML Researchers & EngineersLawyers & Legal ResearchersPrivacy Policy Makers

Paper Title

Compliant But Unsatisfactory: The Gap Between Auditing Standards and Practices for Probabilistic Genotyping Software

Publication Info

  • Topic area: Auditing standards for forensic probabilistic genotyping software in criminal justice.
  • Keywords: Probabilistic genotyping software, forensic science, audit standards, accountability, ASB 018, DNA analysis, algorithm auditing, criminal justice, AI governance, validation studies.

Background and Problem

  • Problem / challenge: Current auditing standards, such as ASB 018, allow audits to comply with formal requirements while failing to meet their intended goals, leading to inadequate assurance of software reliability in forensic contexts.
  • Significance: Probabilistic genotyping software (PGS) plays a critical role in criminal trials, influencing judicial decisions. Inadequate audits can undermine the reliability of DNA evidence and erode trust in forensic science.
  • Motivation and related work: While prior work has highlighted the importance of algorithm audits and standards for accountability, gaps remain in understanding how the design of audit standards can enable "audit washing"—where audits lend credibility to inadequate systems. This paper investigates these gaps through the lens of ASB 018.

Solution

  • Proposed approach: A qualitative analysis of ASB 018 and five audit reports to identify gaps between envisioned and actual practices, and how the standard’s design enables these gaps.
  • Novelty:
    1. Identifies specific gaps between ASB 018’s goals and compliant audit practices across five audit stages.
    2. Highlights design features of ASB 018 that enable these gaps, such as vague language and undefined terms.
    3. Provides recommendations for designing more effective audit standards that ensure accountability.
  • Procedure and key techniques:
    • Analyzed five publicly available PGS audit reports for compliance with ASB 018.
    • Used a framework of five audit stages (e.g., Audit Scope, Standards Identification) to assess gaps between ASB 018’s goals and practices.
    • Conducted qualitative coding of ASB 018’s requirements to identify design features enabling these gaps.

Results

  • Concrete findings:
    • All five audits analyzed were compliant with ASB 018 but failed to meet its envisioned goals.
    • Key gaps include inadequate evaluation of the broader sociotechnical system, failure to define falsifiable performance expectations, and insufficient documentation for third-party scrutiny.
    • Labs often framed software failures as acceptable or expected, avoiding the establishment of boundaries on software use.
  • Advantage over baselines: The study does not propose a new system but critically evaluates an existing standard (ASB 018) and identifies areas for improvement.
  • Experiments / evaluation:
    • Data: Five audit reports spanning different labs, software versions, and years (2016–2024).
    • Metrics: Compliance with ASB 018’s line-level requirements; alignment of practices with envisioned goals.
    • Methodology: Qualitative coding and thematic analysis.
  • Limitations and future work:
    • Reliance on audit reports as proxies for actual practices.
    • Limited representativeness due to the small number of publicly available reports.
    • Future work should examine broader stakeholder perspectives and explore quantitative methods to validate findings.

Summary

This paper investigates the effectiveness of ASB 018, a standard for auditing probabilistic genotyping software, by analyzing five audit reports. It identifies significant gaps between the standard’s goals and the practices it enables, such as inadequate evaluation of user input and insufficient documentation. These gaps arise from design flaws in ASB 018, including vague language and undefined terms. The authors recommend clearer articulation of goals, increased specificity in requirements, and co-designing standards with diverse stakeholders. These findings are critical for improving accountability in forensic science and ensuring the reliability of DNA evidence in criminal trials.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223224/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791552
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Explainable AI (XAI), Algorithmic Transparency & Auditability, Privacy by Design & User Control
work
Professions
AI/ML Researchers & Engineers, Lawyers & Legal Researchers, Privacy Policy Makers
article
Content Status
Full text indexed
hub
Related Papers
4 related papers