PASTA: A Scalable Framework for Multi-Policy AI Compliance Evaluation
Authors
Paper Title
PASTA: A Scalable Framework for Multi-Policy AI Compliance Evaluation
Publication Info
- Topic area: AI compliance evaluation across multiple policies and jurisdictions.
- Keywords: AI compliance, policy evaluation, large language models, model cards, regulatory alignment, multi-policy governance, cost efficiency, scalability, interpretability, responsible AI.
Background and Problem
- Problem / challenge: The rapid expansion of AI policies globally creates significant challenges for practitioners, particularly those without legal expertise, to ensure compliance across multiple jurisdictions. Existing tools are often limited to single-policy evaluations, require substantial resources, or are inaccessible to small teams.
- Significance: Ensuring compliance is critical for mitigating risks, promoting transparency, and building trust in AI systems. Scalable and cost-efficient tools are needed to make compliance evaluation accessible to resource-constrained practitioners.
- Motivation and related work: Current solutions, including questionnaire-based tools and code-scanning platforms, are either inflexible, expensive, or limited to implemented systems. Existing model cards and AI governance tools lack support for multi-policy compliance evaluation. This paper addresses these gaps by introducing PASTA, a scalable, automated, and interpretable compliance evaluation framework.
Solution
- Proposed approach: PASTA (Policy Aggregator & Scanner for Trustworthy AI) is a framework that leverages large language models (LLMs) to evaluate AI system documentation against multiple global policies, providing interpretable and actionable compliance insights.
- Novelty:
- A specialized model-card input schema designed for multi-policy compliance evaluation.
- A policy normalization pipeline that standardizes diverse regulations into a consistent, analyzable format.
- Cost-saving strategies, including policy chunking and irrelevancy mapping, to reduce computational overhead.
- An interactive interface delivering compliance heatmaps, summaries, and actionable recommendations.
- Procedure and key techniques:
- Users complete a lightweight model card capturing system-level information.
- Policies are normalized into a paragraph-level table format for consistent evaluation.
- An LLM-powered engine performs pairwise comparisons between model card sections and policy clauses, assigning violation and relevance scores.
- Results are aggregated into interpretable outputs, including heatmaps, summaries, and issue-fix tables.
Results
- Concrete findings:
- PASTA’s judgments closely align with expert evaluations (Spearman ρ = 0.626 for violation scores, ρ = 0.761 for relevance scores).
- Multi-policy evaluations complete in 1.5–1.8 minutes at a cost of $2.86–$3.06.
- Adding new policies incurs a marginal cost of $4.44 and 2.24 minutes.
- Users completed model card inputs in an average of 28.4 minutes (SD = 6.7).
- Advantage over baselines:
- Reduces evaluation volume by 44.9% through irrelevancy filtering.
- Supports multi-policy compliance in a single pass, unlike questionnaire-based tools requiring separate inputs for each policy.
- Provides interpretable outputs, enabling practitioners to identify and address compliance gaps efficiently.
- Experiments / evaluation:
- Expert evaluation: 324 section-policy pairs rated by three legal experts, showing strong alignment with PASTA’s outputs.
- Usability study: 12 AI practitioners found PASTA actionable and interpretable, with a System Usability Scale (SUS) score of 73.5 (SD = 15.2).
- Limitations and future work:
- Limited scale of expert validation; larger studies are needed.
- Current outputs focus on policy-specific evaluations; future work could explore cross-policy reasoning and conflict detection.
- Participants requested stronger guidance for remediation and improved visualization features.
Summary
PASTA introduces a scalable, cost-efficient framework for multi-policy AI compliance evaluation, addressing critical gaps in existing tools. By leveraging LLMs, a specialized model card format, and cost-saving strategies, PASTA enables practitioners to evaluate their systems against multiple global policies with minimal effort. Expert evaluations and user studies demonstrate its accuracy, usability, and practical value, making compliance evaluation accessible to resource-constrained teams. Future work will focus on expanding policy coverage, enhancing remediation guidance, and exploring educational and organizational applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
CHI '25· Explainable AI (XAI) +2
- 86%
Decomposing Autonomy: Explaining AI Technology Acceptance Through a Liberty-Based Framework
CHI '26· Explainable AI (XAI) +2
- 86%
Certified AI System = Trustworthy? Exploring Expert and Lay User Perceptions and Needs Regarding AI Certification
CHI '26· Explainable AI (XAI) +2
- 86%
Treading the Transparency Tightrope: A Taxonomy of Risks and Benefits of Foundation Model Data Transparency for Transparency Advocates
CHI '26· Explainable AI (XAI) +2
- 75%
Why am I seeing this: Democratizing End User Auditing for Online Content Recommendations
UIST '25· Explainable AI (XAI) +2
- 71%
Expanding Explainability: Towards Social Transparency in AI systems
CHI '21· Explainable AI (XAI) +2
- 71%
HILL: A Hallucination Identifier for Large Language Models
CHI '24· Explainable AI (XAI) +2
- 63%
Model Positionality and Computational Reflexivity: Promoting Reflexivity in Data Science
CHI '22· Explainable AI (XAI) +2
- 63%
Is this AI trained on Credible Data? The Effects of Labeling Quality and Performance Bias on User Trust
CHI '23· Explainable AI (XAI) +2
- 63%
Don't Look at the Data! How Differential Privacy Reconfigures the Practices of Data Science
CHI '23· AI Ethics, Fairness & Accountability +2
Based on Jaccard similarity of research subtopics & professions (≥60%)