PASTA: A Scalable Framework for Multi-Policy AI Compliance Evaluation

Explainable AI (XAI)AI Ethics, Fairness & AccountabilityAlgorithmic Transparency & AuditabilityPrivacy by Design & User ControlAI/ML Researchers & EngineersPrivacy Policy MakersHCI Researchers

Paper Title

PASTA: A Scalable Framework for Multi-Policy AI Compliance Evaluation

Publication Info

  • Topic area: AI compliance evaluation across multiple policies and jurisdictions.
  • Keywords: AI compliance, policy evaluation, large language models, model cards, regulatory alignment, multi-policy governance, cost efficiency, scalability, interpretability, responsible AI.

Background and Problem

  • Problem / challenge: The rapid expansion of AI policies globally creates significant challenges for practitioners, particularly those without legal expertise, to ensure compliance across multiple jurisdictions. Existing tools are often limited to single-policy evaluations, require substantial resources, or are inaccessible to small teams.
  • Significance: Ensuring compliance is critical for mitigating risks, promoting transparency, and building trust in AI systems. Scalable and cost-efficient tools are needed to make compliance evaluation accessible to resource-constrained practitioners.
  • Motivation and related work: Current solutions, including questionnaire-based tools and code-scanning platforms, are either inflexible, expensive, or limited to implemented systems. Existing model cards and AI governance tools lack support for multi-policy compliance evaluation. This paper addresses these gaps by introducing PASTA, a scalable, automated, and interpretable compliance evaluation framework.

Solution

  • Proposed approach: PASTA (Policy Aggregator & Scanner for Trustworthy AI) is a framework that leverages large language models (LLMs) to evaluate AI system documentation against multiple global policies, providing interpretable and actionable compliance insights.
  • Novelty:
    1. A specialized model-card input schema designed for multi-policy compliance evaluation.
    2. A policy normalization pipeline that standardizes diverse regulations into a consistent, analyzable format.
    3. Cost-saving strategies, including policy chunking and irrelevancy mapping, to reduce computational overhead.
    4. An interactive interface delivering compliance heatmaps, summaries, and actionable recommendations.
  • Procedure and key techniques:
    1. Users complete a lightweight model card capturing system-level information.
    2. Policies are normalized into a paragraph-level table format for consistent evaluation.
    3. An LLM-powered engine performs pairwise comparisons between model card sections and policy clauses, assigning violation and relevance scores.
    4. Results are aggregated into interpretable outputs, including heatmaps, summaries, and issue-fix tables.

Results

  • Concrete findings:
    • PASTA’s judgments closely align with expert evaluations (Spearman ρ = 0.626 for violation scores, ρ = 0.761 for relevance scores).
    • Multi-policy evaluations complete in 1.5–1.8 minutes at a cost of $2.86–$3.06.
    • Adding new policies incurs a marginal cost of $4.44 and 2.24 minutes.
    • Users completed model card inputs in an average of 28.4 minutes (SD = 6.7).
  • Advantage over baselines:
    • Reduces evaluation volume by 44.9% through irrelevancy filtering.
    • Supports multi-policy compliance in a single pass, unlike questionnaire-based tools requiring separate inputs for each policy.
    • Provides interpretable outputs, enabling practitioners to identify and address compliance gaps efficiently.
  • Experiments / evaluation:
    • Expert evaluation: 324 section-policy pairs rated by three legal experts, showing strong alignment with PASTA’s outputs.
    • Usability study: 12 AI practitioners found PASTA actionable and interpretable, with a System Usability Scale (SUS) score of 73.5 (SD = 15.2).
  • Limitations and future work:
    • Limited scale of expert validation; larger studies are needed.
    • Current outputs focus on policy-specific evaluations; future work could explore cross-policy reasoning and conflict detection.
    • Participants requested stronger guidance for remediation and improved visualization features.

Summary

PASTA introduces a scalable, cost-efficient framework for multi-policy AI compliance evaluation, addressing critical gaps in existing tools. By leveraging LLMs, a specialized model card format, and cost-saving strategies, PASTA enables practitioners to evaluate their systems against multiple global policies with minimal effort. Expert evaluations and user studies demonstrate its accuracy, usability, and practical value, making compliance evaluation accessible to resource-constrained teams. Future work will focus on expanding policy coverage, enhancing remediation guidance, and exploring educational and organizational applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222225/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790966
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), AI Ethics, Fairness & Accountability, Algorithmic Transparency & Auditability, Privacy by Design & User Control
work
Professions
AI/ML Researchers & Engineers, Privacy Policy Makers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers