Dark and Bright Side of Participatory Red-Teaming with Targets of Stereotyping for Eliciting Harmful Behaviors from Large Language Models
Honorable MentionAuthors
Paper Title
Dark and Bright Side of Participatory Red-Teaming with Targets of Stereotyping for Eliciting Harmful Behaviors from Large Language Models
Publication Info
- Topic area: Participatory red-teaming to identify and mitigate biases in generative AI systems.
- Keywords: Participatory red-teaming, stereotypes, large language models, bias detection, psychological impacts, empowerment, AI ethics, lived experience, stigma, AI safety.
Background and Problem
- Problem / challenge: Generative AI systems often reproduce or amplify societal stereotypes, causing representational harm and psychological impacts. Existing red-teaming approaches lack the inclusion of affected communities, particularly those targeted by stereotypes, and fail to address the psychological risks involved.
- Significance: Stereotypical biases in AI systems can perpetuate social hierarchies, harm marginalized groups, and undermine trust in AI. Including stereotype targets in red-teaming is critical for identifying subtle biases and ensuring AI safety.
- Motivation and related work: Prior studies highlight the limitations of automated and homogeneous red-teaming approaches in detecting context-sensitive biases. While participatory red-teaming with diverse perspectives has shown promise, little research has explored its psychological impacts on stereotype targets or how their lived experiences can be leveraged effectively and ethically.
Solution
- Proposed approach: Participatory red-teaming involving individuals targeted by stereotypes to elicit and evaluate harmful behaviors in large language models (LLMs).
- Novelty:
- Empirical investigation of psychological costs and benefits for stereotype targets participating in red-teaming.
- Analysis of how lived experiences are transformed into strategic expertise for bias detection.
- Examination of empowerment potential through participatory red-teaming.
- Design considerations for ethical and empowering red-teaming protocols.
- Procedure and key techniques:
- Recruitment of 20 participants stigmatized by stereotypes against non-prestigious university graduates in South Korea.
- Mixed-methods analysis combining psychological surveys, red-teaming documentation, and semi-structured interviews.
- Use of structured prompt templates and iterative adversarial strategies to elicit harmful AI outputs.
- Implementation of safety protocols, including distress monitoring, meditation breaks, and post-session debriefing.
Results
- Concrete findings:
- Psychological distress (SUDS) and negative affect increased significantly post-task (p < 0.001), while collective self-esteem decreased (p = 0.049).
- Participants generated 82 attack attempts, with a success rate of 63.4% (52 successful attacks).
- High cognitive and emotional workload was reported (e.g., NASA-TLX effort score: 8.40/10).
- Advantage over baselines:
- Participants’ lived experiences enabled the identification of subtle biases and the development of sophisticated attack strategies, outperforming generic or automated red-teaming approaches.
- Unique insights into harmful AI behaviors and nuanced evaluation criteria were derived from in-group perspectives.
- Experiments / evaluation:
- Psychological measures included K-PANAS, RSES, CSES, SCQ, SUDS, and NASA-TLX.
- Qualitative analysis of red-teaming documentation and interviews revealed strategic prompting patterns and emotional dynamics.
- Participants employed six prompting strategies, with success rates ranging from 43.8% to 100%.
- Limitations and future work:
- Limited to a single cultural context and small sample size (N = 20).
- Short-term study without longitudinal follow-up to assess lasting psychological impacts.
- Ethical tensions in exposing participants to discriminatory content.
- Future work should explore diverse stereotypes, protective interventions, and scalable safeguards for industrial applications.
Summary
This study investigates the psychological and strategic dynamics of participatory red-teaming with stereotype targets, focusing on graduates of non-prestigious universities in South Korea. Participants demonstrated unique expertise in detecting biases through lived experiences but faced significant psychological costs, including increased distress and reduced collective self-esteem. Despite these challenges, they reported empowerment, critical AI awareness, and a sense of contributing to AI safety. The findings highlight the need for ethical safeguards and inclusive design in participatory red-teaming to balance psychological risks with the potential for meaningful empowerment and bias mitigation.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)