Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
Honorable MentionAuthors
Paper Title
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
Publication Info
- Topic area: Red teaming large language models (LLMs) for safety and reliability evaluation.
- Keywords: Red teaming, adversarial datasets, large language models, AI safety, socio-technical practices, evaluation metrics, harm taxonomy, human-computer interaction, risk assessment, dataset creation.
Background and Problem
- Problem / challenge: Existing red teaming efforts for LLMs focus narrowly on technical benchmarks and attack success rates, neglecting the socio-technical practices involved in defining, creating, and evaluating adversarial datasets.
- Significance: Adversarial datasets are critical artifacts for assessing LLM vulnerabilities and potential harms, influencing model alignment and safety improvements. Understanding how these datasets are constructed and evaluated is essential for addressing broader societal impacts.
- Motivation and related work: Red teaming originated in security contexts and has become central to AI evaluation, particularly for LLMs. Prior work has highlighted the risks of biased, harmful, or misleading outputs but has not sufficiently examined the socio-technical dimensions of dataset creation and evaluation.
Solution
- Proposed approach: Empirical study of red teaming practices through semi-structured interviews with AI practitioners to understand how adversarial datasets are conceptualized, developed, and evaluated.
- Novelty:
- Empirical evidence on practitioners’ conceptualization of red teaming and dataset creation.
- Identification of gaps in context, interaction type, and user specificity in red teaming practices.
- Recommendations for expanding red teaming evaluations to reflect socio-technical dimensions.
- Procedure and key techniques:
- Conducted 22 semi-structured interviews with AI practitioners across academia, industry, and open-source communities.
- Thematic analysis of interview transcripts to identify patterns in dataset creation, risk definition, and evaluation practices.
- Highlighted three critical moments in red teaming: task definition, dataset development, and evaluation.
Results
- Concrete findings:
- Practitioners conceptualize red teaming either as exploration (searching for attack cases) or as classification (detecting predefined harmful content).
- Adversarial datasets are created using three approaches: repurposing existing datasets, creating datasets from scratch, or deriving datasets from human–LLM interactions.
- Evaluation methods rely heavily on automated classifiers and LLM judges, but human evaluation remains critical for nuanced judgments.
- Advantage over baselines:
- Identified socio-technical gaps in current red teaming practices, such as neglecting context, interaction types, and user-specific needs.
- Proposed actionable recommendations to address these gaps, emphasizing interdisciplinary collaboration and participatory approaches.
- Experiments / evaluation:
- Interviews covered diverse practitioner profiles, including PhD students, researchers, and engineers from eight countries.
- Analysis revealed challenges in defining harm, evaluating adversarial outputs, and ensuring dataset diversity.
- Limitations and future work:
- Limited representation of industry practitioners in the study sample.
- Future work should explore participatory methods to engage end-users and domain experts in red teaming evaluations.
Summary
This study investigates the socio-technical practices of red teaming LLMs by analyzing how practitioners define tasks, develop adversarial datasets, and evaluate model vulnerabilities. Findings reveal that harmfulness is constructed through cultural norms and technical decisions embedded in dataset creation and evaluation. The study highlights gaps in addressing context, interaction types, and user-specific needs, offering recommendations for HCI researchers to design more inclusive and context-aware red teaming approaches. These insights contribute to advancing AI safety evaluation by integrating socio-technical dimensions into red teaming practices.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
A Framework to Characterize Reporting on Generative AI Use
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 83%
Explanations as Mechanisms for Supporting Algorithmic Transparency
CHI '18· Explainable AI (XAI) +1
- 83%
HILL: A Hallucination Identifier for Large Language Models
CHI '24· Explainable AI (XAI) +2
- 71%
RELIC: Investigating Large Language Model Responses using Self-Consistency
CHI '24· Explainable AI (XAI) +2
- 71%
Do Expressions Change Decisions? Exploring the Impact of AI's Explanation Tone on Decision-Making
CHI '25· Explainable AI (XAI) +2
- 71%
Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
CHI '25· Explainable AI (XAI) +2
- 71%
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
CHI '25· Explainable AI (XAI) +2
- 71%
“I Don’t Think RAI Applies to My Model” – Engaging Non-champions with Sticky Stories for Responsible AI Work
CHI '26· AI Ethics, Fairness & Accountability +2
- 71%
Evaluating Behavior Change Interventions for Responsible Data Science
CHI '26· AI Ethics, Fairness & Accountability +2
- 71%
Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
CHI '26· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)