Evaluating Behavior Change Interventions for Responsible Data Science
Authors
Paper Title
Evaluating Behavior Change Interventions for Responsible Data Science
Publication Info
- Topic area: Responsible data science and behavior change interventions in AI/ML workflows.
- Keywords: Responsible data science, behavior change interventions, fairness, Aequitas, motivational priming, cognitive load, COM-B framework, ethical AI, fairness metrics, data science workflows.
Background and Problem
- Problem / challenge: Despite awareness of algorithmic harms, responsible data science (RDS) practices are inadequately adopted. Existing methods focus on outcome metrics (e.g., fairness indices) but fail to address the behavioral and cognitive processes that shape ethical decision-making.
- Significance: Ensuring fairness and accountability in AI/ML systems is critical in high-stakes domains like healthcare, finance, and criminal justice, where algorithmic biases can lead to systemic discrimination.
- Motivation and related work: Prior work has developed fairness tools and metrics but often neglects the human factors influencing their adoption. The COM-B framework highlights the need to address Capability, Opportunity, and Motivation for sustained behavior change. However, empirical evaluations of behavior change interventions (BCIs) in RDS are lacking.
Solution
- Proposed approach: The study evaluates two BCIs: (i) Prime, a motivational priming intervention using fairness-related narratives, and (ii) Aequitas, a fairness auditing toolkit, to promote responsible behaviors in data science workflows.
- Novelty:
- First empirical evaluation of BCIs in RDS, comparing motivational priming and technical tooling.
- Development of a mixed-methods framework combining behavioral observation, fairness metrics, cognitive load assessment, and qualitative analysis.
- Quantification of cognitive load trade-offs in fairness tools, informing future tool design.
- Identification of the role of personal connection in enhancing ethical motivation.
- Procedure and key techniques:
- Participants (N=12, 5+ years of data science experience) completed two tasks (credit risk and income classification) under three conditions: Control (no intervention), Prime, and Aequitas.
- Responsible behaviors were measured using a checklist of practices across pre-processing, in-processing, and post-processing stages.
- Fairness metrics, model accuracy, cognitive load (NASA-TLX), and participant interviews were analyzed.
Results
- Concrete findings:
- Prime increased responsible behaviors from μ=2.8 (Control) to μ=5.7 (σ=0.75, p<0.01).
- Aequitas further increased responsible behaviors to μ=8.1 (σ=1.34, p<0.01).
- Aequitas significantly improved fairness metrics (false discovery rate ratio reduced to μ=1.18, p<0.01), while Prime showed a non-significant trend toward fairness improvement.
- Neither intervention compromised model accuracy (Prime: μ=0.60 vs. Control: μ=0.59; Aequitas: μ=0.58 vs. Control: μ=0.61).
- Aequitas increased cognitive load (e.g., mental demand: μ=1.67 to μ=3, p=0.02), while Prime did not.
- Advantage over baselines:
- Both interventions outperformed the Control condition in promoting responsible behaviors.
- Aequitas demonstrated superior fairness improvements compared to Prime.
- Experiments / evaluation:
- Mixed-methods study with N=12 participants using counterbalanced within-subjects design.
- Evaluation metrics included responsible behavior scores, fairness metrics, accuracy, and NASA-TLX cognitive load ratings.
- Limitations and future work:
- Small sample size (N=12) limits statistical power.
- Controlled experimental setting may not fully capture real-world complexities.
- Future work should explore hybrid interventions, longitudinal adoption, and organizational incentives for fairness.
Summary
This study evaluates two behavior change interventions—Prime (motivational priming) and Aequitas (fairness toolkit)—to promote responsible data science practices. Both interventions increased responsible behaviors, with Aequitas significantly improving fairness metrics but imposing higher cognitive load. Prime boosted motivation without affecting accuracy or cognitive load. The findings highlight the importance of balancing technical tools with motivational support and suggest that personal connection to fairness scenarios enhances intervention effectiveness. Future research should explore hybrid approaches, longitudinal impacts, and systemic support for embedding fairness across data science workflows.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
“I Don’t Think RAI Applies to My Model” – Engaging Non-champions with Sticky Stories for Responsible AI Work
CHI '26· AI Ethics, Fairness & Accountability +2
- 71%
Sensing What Surveys Miss: Understanding and Personalizing Proactive LLM Support by User Modeling
CHI '26· Human-LLM Collaboration +2
- 71%
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
CHI '26· Explainable AI (XAI) +2
- 71%
Vulnerability of LLM Outputs to Heuristics-Inducing Prompt Structures
IUI '26· Human-LLM Collaboration +2
- 63%
A Framework to Characterize Reporting on Generative AI Use
CHI '26· Generative AI (Text, Image, Music, Video) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)