Surfacing Governing Principles for Chatbots: A Workbench and Comparative Study
Authors
Paper Title
Surfacing Governing Principles for Chatbots: A Workbench and Comparative Study
Publication Info
- Topic area: Governance and configuration of LLM-powered chatbots through principle-based workflows.
- Keywords: Chatbot governance, principle authoring, LLMs, Trust Mediator, AI alignment, HCI, chatbot evaluation, principle coherence, principle specificity, principle coverage.
Background and Problem
- Problem / challenge: Existing tools for chatbot governance often treat principles as isolated rules, lacking support for reasoning about principles as a coherent set. This limits the ability to detect redundancy, contradictions, or behavioral impacts when principles are combined.
- Significance: Trust in LLM-powered chatbots depends on their behavior being governed by clear, actionable principles that align with organizational values and user expectations. Addressing the gap in principle-set governance can improve chatbot reliability and trustworthiness.
- Motivation and related work: Prior work in AI governance and HCI has explored principles as high-level commitments or operational requirements, but these efforts often fail to connect principles to turn-by-turn chatbot behavior. Tools like EvalLM focus on evaluating prompts against criteria but do not support multi-principle coordination. This paper addresses the gap by treating principle sets as coherent artifacts for governance.
Solution
- Proposed approach: Trust Mediator (TM), a workbench for authoring, structuring, and assessing governing principles for LLM-powered chatbots.
- Novelty:
- Treating principle sets as coherent artifacts, enabling inspection and revision as a whole.
- Introducing LLM-based scaffolds for persona-based query generation, principle elicitation, conflict detection, and principle evaluation.
- Supporting iterative refinement of principles by comparing chatbot behavior with and without specific principles.
- Empirical comparison of manual and assisted principle authoring workflows.
- Procedure and key techniques:
- Persona-based query suggestion to simulate diverse user interactions.
- Principle elicitation from user feedback using LLMs to generate candidate principles.
- Automatic detection of conflicts or overlaps between principles.
- Principle evaluation using tailored queries, scoring rubrics, and comparative chatbot response assessments.
Results
- Concrete findings:
- Manual authoring produced more specific principles (SI20: 9.36 vs. 7.24, p = 0.037).
- Assisted authoring broadened coverage across cognitive, emotional, and organizational dimensions (coverage families: 2.8 vs. 1.9, δ = 0.83).
- Coherence scores were slightly higher in the manual condition (CI20: 17.50 vs. 17.16), though not statistically significant.
- Both conditions improved chatbot evaluation scores, with manual authoring showing a slightly larger improvement (1.13 vs. 0.64).
- Advantage over baselines:
- Assisted authoring expanded the range of concerns addressed, particularly emotional safeguards.
- Manual authoring fostered operational clarity and reduced redundancy.
- Experiments / evaluation:
- A between-subjects study with 12 participants (6 per condition) compared manual and assisted workflows.
- Metrics included principle specificity (SI20), coverage (families), coherence (CI20), and chatbot evaluation scores.
- Data sources included authored principles, session logs, and subjective survey responses.
- Limitations and future work:
- Small sample size limits generalizability.
- Study duration (30-minute sessions) may not reflect long-term governance workflows.
- Participants were not service owners of the specific chatbot used in the study.
- Future work should explore phased workflows combining manual and assisted authoring, and test in real-world service contexts.
Summary
This paper introduces Trust Mediator, a workbench for authoring and assessing governing principles for LLM-powered chatbots. The study compares manual and assisted workflows, finding that manual authoring fosters specificity and operational clarity, while assisted authoring broadens coverage and includes more emotional safeguards. Both approaches improved chatbot behavior, highlighting their complementary strengths. The findings suggest phased workflows that integrate manual and assisted methods, emphasizing principles as both operational requirements and communicative commitments. Future work should explore these workflows in larger, real-world contexts to enhance chatbot governance and trustworthiness.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 75%
Understanding Socio-technical Factors Configuring AI Non-Use in UX Work Practices
CHI '25· Human-LLM Collaboration +2
- 75%
Does My Chatbot Have an Agenda? Understanding Human and AI Agency in Human-Human-like Chatbot Interaction
CHI '26· Agent Personality & Anthropomorphism +2
- 75%
Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
CHI '26· Agent Personality & Anthropomorphism +2
- 67%
Who Controls the Conversation? User Perspectives On Generative AI (LLM) System Prompts
CHI '26· Human-LLM Collaboration +3
- 67%
Reactive Writers: How Co-Writing with AI Changes How We Engage with Ideas
CHI '26· Human-LLM Collaboration +3
- 67%
Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents
CHI '26· Human-LLM Collaboration +3
- 63%
Which Artificial Intelligences Do People Care About Most? A Conjoint Experiment on Moral Consideration
CHI '24· Agent Personality & Anthropomorphism +2
- 63%
Interaction Context Often Increases Sycophancy in LLMs
CHI '26· Human-LLM Collaboration +2
- 60%
“It Became My Buddy, But I’m Not Afraid to Disagree”: A Multi-Session Study of UX Evaluators Collaborating with Conversational AI Assistants
CHI '26· Human-LLM Collaboration +4
- 60%
ORAgen Fables: Advancing the Design and Management of Content Attribution
CHI '26· Generative AI (Text, Image, Music, Video) +4
Based on Jaccard similarity of research subtopics & professions (≥60%)