Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
Authors
AI policy sets boundaries on acceptable behavior for AI models, but this is challenging in the context of large language models (LLMs): how do you ensure coverage over a vast behavior space? We introduce policy maps, an approach to AI policy design inspired by the practice of physical mapmaking. Instead of aiming for full coverage, policy maps aid effective navigation through intentional design choices about which aspects to capture and which to abstract away. With Policy Projector, an interactive tool for designing LLM policy maps, an AI practitioner can survey the landscape of model input-output pairs, define custom regions (e.g., “violence”), and navigate these regions with if-then policy rules that can act on LLM outputs (e.g., if output contains “violence” and “graphic details,” then rewrite without “graphic details”). Policy Projector supports interactive policy authoring using LLM classification and steering and a map visualization reflecting the AI practitioner’s work. In an evaluation with 12 AI safety experts, our system helps policy designers craft policies around problematic model behaviors such as incorrect gender assumptions and handling of immediate physical safety threats.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
Do Expressions Change Decisions? Exploring the Impact of AI's Explanation Tone on Decision-Making
CHI '25· Explainable AI (XAI) +2
- 71%
Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
CHI '25· Explainable AI (XAI) +2
- 71%
Treading the Transparency Tightrope: A Taxonomy of Risks and Benefits of Foundation Model Data Transparency for Transparency Advocates
CHI '26· Explainable AI (XAI) +2
- 67%
Paths Explored, Paths Omitted, Paths Obscured: Decision Points & Selective Reporting in End-to-End Data Analysis
CHI '20· Explainable AI (XAI) +1
- 67%
Manipulating and Measuring Model Interpretability
CHI '21· Explainable AI (XAI) +1
- 67%
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
CHI '21· Explainable AI (XAI) +1
- 67%
Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model Behavior
CHI '22· Explainable AI (XAI) +1
- 67%
Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning
CHI '22· Explainable AI (XAI) +1
- 67%
Fairness Evaluation in Text Classification: Machine Learning Practitioner Perspectives of Individual and Group Fairness
CHI '23· Explainable AI (XAI) +1
- 67%
Perceptions of the Fairness Impacts of Multiplicity in Machine Learning
CHI '25· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)