“I Don’t Think RAI Applies to My Model” – Engaging Non-champions with Sticky Stories for Responsible AI Work
Best PaperAuthors
Paper Title
“I Don’t Think RAI Applies to My Model” – Engaging Non-champions with Sticky Stories for Responsible AI Work
Publication Info
- Topic area: Responsible AI (RAI) engagement strategies for practitioners
- Keywords: Responsible AI, sticky stories, non-champions, harm identification, critical reflection, engagement, cognitive dissonance, transformative learning, LLM-generated narratives, AI ethics
Background and Problem
- Problem / challenge: Existing Responsible AI (RAI) tools and frameworks primarily engage motivated practitioners (RAI champions) but fail to meaningfully involve non-champions, who often see RAI as irrelevant or bureaucratic.
- Significance: Non-champions constitute the majority of practitioners, and their lack of engagement limits the real-world impact of RAI initiatives. Motivating this group is critical for broader adoption of ethical AI practices.
- Motivation and related work: Prior research has focused on RAI champions and tools like checklists, templates, and governance processes. However, these approaches often lead to superficial compliance rather than genuine engagement among non-champions. There is a gap in designing interventions that provoke deeper reflection and engagement for this group.
Solution
- Proposed approach: Introduction of "sticky stories"—narratives illustrating unexpected and severe harms caused by AI systems, designed to provoke critical reflection and engagement among non-champions.
- Novelty:
- Development of a scalable pipeline for generating sticky stories using large language models (LLMs) with qualities such as concreteness, severity, surprisingness, diversity, and relevance.
- Empirical evaluation showing sticky stories significantly increase engagement, harm identification, and critical reflection compared to baseline stories.
- Identification of distinct practitioner engagement trajectories and profiles (e.g., resistors, indifferents, followers, learners, champions).
- Integration of sticky stories into an interactive tool for RAI harm assessment.
- Procedure and key techniques:
- Conducted formative study to identify barriers to non-champion engagement.
- Designed an eight-step pipeline for generating sticky stories, including stakeholder identification, harm type pre-definition, and iterative refinement for concreteness and severity.
- Evaluated story qualities (e.g., severity, surprisingness) using human and LLM assessments.
- Conducted a mixed-design user study with 29 practitioners to measure engagement, harm identification, and critical reflection.
Results
- Concrete findings:
- Sticky stories were more concrete (+98.3%), severe (+30.9%), surprising (+42.9%), and diverse (+5.8%) than baseline stories but slightly less relevant (-8.7%).
- Practitioners exposed to sticky stories spent 207% more time on harm identification tasks and identified 4.5× more harm categories and 3.5× more subcategories compared to baseline stories.
- Critical reflection indicators (e.g., challenging assumptions, connecting to wider systems) were more frequent in the sticky story condition.
- Advantage over baselines:
- Sticky stories outperformed baseline stories in engaging practitioners, broadening harm identification, and fostering deeper reflection.
- Baseline stories often led to superficial engagement, while sticky stories prompted sustained cognitive processing and actionable insights.
- Experiments / evaluation:
- Offline evaluation of 240 stories (120 sticky, 120 baseline) confirmed the effectiveness of the sticky story pipeline.
- User study with 29 practitioners (non-champions) assessed engagement, harm diversity, and critical reflection using a mixed-design approach.
- Follow-up survey revealed more post-study actions in the sticky story group.
- Limitations and future work:
- Short-term engagement was measured; long-term behavior change remains untested.
- Small sample size limits generalizability across domains and organizations.
- Risks of unrealistic or overly dramatic stories undermining credibility.
- Future work should explore how individual story qualities affect different practitioner profiles and improve the realism of LLM-generated narratives.
Summary
This paper introduces "sticky stories," a narrative-based intervention designed to engage non-champions in Responsible AI (RAI) work by illustrating unexpected and severe harms caused by AI systems. Sticky stories were generated using a scalable LLM-based pipeline and evaluated for qualities like concreteness, severity, and surprisingness. Empirical results showed that sticky stories significantly increased engagement, harm identification, and critical reflection compared to baseline stories. The study identified distinct practitioner profiles and trajectories, highlighting how sticky stories can shift attitudes and behaviors. While promising, the approach faces challenges related to story realism and long-term impact, suggesting directions for future research and refinement.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
Farsight: Fostering Responsible AI Awareness During AI Application Prototyping
CHI '24· Explainable AI (XAI) +2
- 83%
Trust in AI-assisted Decision Making: Perspectives from Those Behind the System and Those for Whom the Decision is Made
CHI '24· Explainable AI (XAI) +2
- 83%
Trusting Autonomous Teammates in Human-AI Teams - A Literature Review
CHI '25· Explainable AI (XAI) +2
- 71%
Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models
CHI '19· Explainable AI (XAI) +2
- 71%
Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine Learning
CHI '20· Explainable AI (XAI) +2
- 71%
No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive ML
CHI '20· Explainable AI (XAI) +2
- 71%
Knowing About Knowing: An Illusion of Human Competence Can Hinder Appropriate Reliance on AI Systems
CHI '23· Explainable AI (XAI) +2
- 71%
ESCAPE: Countering Systematic Errors from Machine's Blind Spots via Interactive Visual Analysis
CHI '23· Explainable AI (XAI) +2
- 71%
RELIC: Investigating Large Language Model Responses using Self-Consistency
CHI '24· Explainable AI (XAI) +2
- 71%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)