PolicyPad: Collaborative Prototyping of LLM Policies

Human-LLM CollaborationExplainable AI (XAI)AI-Assisted Decision-Making & AutomationPsychiatrists & PsychotherapistsPrivacy Policy MakersHCI Researchers

Paper Title

PolicyPad: Collaborative Prototyping of LLM Policies

Publication Info

  • Topic area: Collaborative design of policies for large language models (LLMs) in high-stakes domains.
  • Keywords: LLM policies, AI alignment, collaborative design, UX prototyping, mental health, legal domain, AI safety, policy iteration, real-time collaboration, human-centered AI.

Background and Problem

  • Problem / challenge: Existing LLM policies are often insular, lack input from domain experts, and fail to address critical safety concerns in high-stakes domains like mental health and law. Tools to support collaborative policy design are scarce.
  • Significance: LLM policies are essential for governing AI behavior transparently and responsibly, especially in regulated domains where model outputs can have significant consequences.
  • Motivation and related work: Prior work has explored AI alignment, red-teaming, and participatory evaluation but has not focused on tools for collaborative policy design. There is a gap in enabling domain experts to iteratively design and test policies that govern LLM behavior.

Solution

  • Proposed approach: PolicyPad, an interactive system for collaborative prototyping of LLM policies, inspired by UX prototyping practices like heuristic evaluation and storyboarding.
  • Novelty:
    1. Conceptualization of "LLM policy prototyping" as a collaborative, iterative process for designing LLM policies.
    2. Development of PolicyPad, a system integrating real-time collaboration, scenario-based testing, and heuristic evaluation.
    3. Empirical evaluation of PolicyPad with domain experts, demonstrating its ability to foster collaboration and produce novel policies.
    4. Identification of novel policy contributions, such as guidelines for deferring to human experts and eliciting critical user information.
  • Procedure and key techniques:
    1. Experts collaboratively draft policies in a shared editor.
    2. Scenarios representing real-world AI use are used to ground discussions and test policy-informed model behavior.
    3. Iterative feedback loops allow experts to refine policies based on model responses.
    4. Heuristics guide high-level policy design, while interactive widgets and spotlight scenarios facilitate collaboration and testing.

Results

  • Concrete findings:
    • 51.9% of policies created with PolicyPad were novel, compared to 18.2% in the baseline system.
    • Novel policies included specific guidelines for deferring to human experts, procedural rules for emergency situations, and elicitation of critical user information.
  • Advantage over baselines:
    • PolicyPad produced 4 times more novel policies than the baseline system.
    • Enhanced collaboration dynamics, richer discussions, and more structured workflows compared to the baseline.
  • Experiments / evaluation:
    • Conducted with 22 domain experts (10 in mental health, 12 in law) across 8 groups.
    • Tasks included drafting policies for conversational tone and safety guardrails.
    • Evaluation metrics included policy novelty, expert ratings, and qualitative feedback.
  • Limitations and future work:
    • Limited geographic diversity (mostly U.S.-based participants).
    • Scenarios may not fully capture the diversity of real-world use cases.
    • Future work could explore scaling up policy prototyping, resolving expert disagreements, and integrating global perspectives.

Summary

This paper introduces PolicyPad, a system for collaborative prototyping of LLM policies, which enables domain experts to draft, test, and refine policies in real time. Through workshops with 22 experts in mental health and law, PolicyPad demonstrated its ability to foster collaboration and produce novel policies addressing critical safety concerns. Key contributions include guidelines for deferring to human experts, procedural rules for emergencies, and eliciting critical user information. Future work could expand the geographic and cultural diversity of participants and explore scaling policy prototyping for broader participatory AI governance.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/221925/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791689
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
Psychiatrists & Psychotherapists, Privacy Policy Makers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers