Surfacing Governing Principles for Chatbots: A Workbench and Comparative Study

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationAI Ethics, Fairness & AccountabilityAlgorithmic Transparency & AuditabilityAgent Personality & AnthropomorphismAI/ML Researchers & EngineersUI/UX DesignersHCI Researchers

Paper Title

Surfacing Governing Principles for Chatbots: A Workbench and Comparative Study

Publication Info

  • Topic area: Governance and configuration of LLM-powered chatbots through principle-based workflows.
  • Keywords: Chatbot governance, principle authoring, LLMs, Trust Mediator, AI alignment, HCI, chatbot evaluation, principle coherence, principle specificity, principle coverage.

Background and Problem

  • Problem / challenge: Existing tools for chatbot governance often treat principles as isolated rules, lacking support for reasoning about principles as a coherent set. This limits the ability to detect redundancy, contradictions, or behavioral impacts when principles are combined.
  • Significance: Trust in LLM-powered chatbots depends on their behavior being governed by clear, actionable principles that align with organizational values and user expectations. Addressing the gap in principle-set governance can improve chatbot reliability and trustworthiness.
  • Motivation and related work: Prior work in AI governance and HCI has explored principles as high-level commitments or operational requirements, but these efforts often fail to connect principles to turn-by-turn chatbot behavior. Tools like EvalLM focus on evaluating prompts against criteria but do not support multi-principle coordination. This paper addresses the gap by treating principle sets as coherent artifacts for governance.

Solution

  • Proposed approach: Trust Mediator (TM), a workbench for authoring, structuring, and assessing governing principles for LLM-powered chatbots.
  • Novelty:
    1. Treating principle sets as coherent artifacts, enabling inspection and revision as a whole.
    2. Introducing LLM-based scaffolds for persona-based query generation, principle elicitation, conflict detection, and principle evaluation.
    3. Supporting iterative refinement of principles by comparing chatbot behavior with and without specific principles.
    4. Empirical comparison of manual and assisted principle authoring workflows.
  • Procedure and key techniques:
    • Persona-based query suggestion to simulate diverse user interactions.
    • Principle elicitation from user feedback using LLMs to generate candidate principles.
    • Automatic detection of conflicts or overlaps between principles.
    • Principle evaluation using tailored queries, scoring rubrics, and comparative chatbot response assessments.

Results

  • Concrete findings:
    • Manual authoring produced more specific principles (SI20: 9.36 vs. 7.24, p = 0.037).
    • Assisted authoring broadened coverage across cognitive, emotional, and organizational dimensions (coverage families: 2.8 vs. 1.9, δ = 0.83).
    • Coherence scores were slightly higher in the manual condition (CI20: 17.50 vs. 17.16), though not statistically significant.
    • Both conditions improved chatbot evaluation scores, with manual authoring showing a slightly larger improvement (1.13 vs. 0.64).
  • Advantage over baselines:
    • Assisted authoring expanded the range of concerns addressed, particularly emotional safeguards.
    • Manual authoring fostered operational clarity and reduced redundancy.
  • Experiments / evaluation:
    • A between-subjects study with 12 participants (6 per condition) compared manual and assisted workflows.
    • Metrics included principle specificity (SI20), coverage (families), coherence (CI20), and chatbot evaluation scores.
    • Data sources included authored principles, session logs, and subjective survey responses.
  • Limitations and future work:
    • Small sample size limits generalizability.
    • Study duration (30-minute sessions) may not reflect long-term governance workflows.
    • Participants were not service owners of the specific chatbot used in the study.
    • Future work should explore phased workflows combining manual and assisted authoring, and test in real-world service contexts.

Summary

This paper introduces Trust Mediator, a workbench for authoring and assessing governing principles for LLM-powered chatbots. The study compares manual and assisted workflows, finding that manual authoring fosters specificity and operational clarity, while assisted authoring broadens coverage and includes more emotional safeguards. Both approaches improved chatbot behavior, highlighting their complementary strengths. The findings suggest phased workflows that integrate manual and assisted methods, emphasizing principles as both operational requirements and communicative commitments. Future work should explore these workflows in larger, real-world contexts to enhance chatbot governance and trustworthiness.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222360/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790612
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, AI Ethics, Fairness & Accountability, Algorithmic Transparency & Auditability
work
Professions
AI/ML Researchers & Engineers, UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers