Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors

Explainable AI (XAI)Algorithmic Transparency & AuditabilityAlgorithmic Fairness & BiasAI/ML Researchers & EngineersPrivacy Policy MakersHCI Researchers

AI policy sets boundaries on acceptable behavior for AI models, but this is challenging in the context of large language models (LLMs): how do you ensure coverage over a vast behavior space? We introduce policy maps, an approach to AI policy design inspired by the practice of physical mapmaking. Instead of aiming for full coverage, policy maps aid effective navigation through intentional design choices about which aspects to capture and which to abstract away. With Policy Projector, an interactive tool for designing LLM policy maps, an AI practitioner can survey the landscape of model input-output pairs, define custom regions (e.g., “violence”), and navigate these regions with if-then policy rules that can act on LLM outputs (e.g., if output contains “violence” and “graphic details,” then rewrite without “graphic details”). Policy Projector supports interactive policy authoring using LLM classification and steering and a map visualization reflecting the AI practitioner’s work. In an evaluation with 12 AI safety experts, our system helps policy designers craft policies around problematic model behaviors such as incorrect gender assumptions and handling of immediate physical safety threats.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/206991/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3746059.3747680
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Explainable AI (XAI), Algorithmic Transparency & Auditability, Algorithmic Fairness & Bias
work
Professions
AI/ML Researchers & Engineers, Privacy Policy Makers, HCI Researchers
article
Content Status
Abstract only
hub
Related Papers
10 related papers