Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation

Honorable Mention
Explainable AI (XAI)AI Ethics, Fairness & AccountabilityAlgorithmic Transparency & AuditabilityAI/ML Researchers & EngineersHCI ResearchersData Scientists & Analysts

Paper Title

Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation

Publication Info

  • Topic area: Red teaming large language models (LLMs) for safety and reliability evaluation.
  • Keywords: Red teaming, adversarial datasets, large language models, AI safety, socio-technical practices, evaluation metrics, harm taxonomy, human-computer interaction, risk assessment, dataset creation.

Background and Problem

  • Problem / challenge: Existing red teaming efforts for LLMs focus narrowly on technical benchmarks and attack success rates, neglecting the socio-technical practices involved in defining, creating, and evaluating adversarial datasets.
  • Significance: Adversarial datasets are critical artifacts for assessing LLM vulnerabilities and potential harms, influencing model alignment and safety improvements. Understanding how these datasets are constructed and evaluated is essential for addressing broader societal impacts.
  • Motivation and related work: Red teaming originated in security contexts and has become central to AI evaluation, particularly for LLMs. Prior work has highlighted the risks of biased, harmful, or misleading outputs but has not sufficiently examined the socio-technical dimensions of dataset creation and evaluation.

Solution

  • Proposed approach: Empirical study of red teaming practices through semi-structured interviews with AI practitioners to understand how adversarial datasets are conceptualized, developed, and evaluated.
  • Novelty:
    1. Empirical evidence on practitioners’ conceptualization of red teaming and dataset creation.
    2. Identification of gaps in context, interaction type, and user specificity in red teaming practices.
    3. Recommendations for expanding red teaming evaluations to reflect socio-technical dimensions.
  • Procedure and key techniques:
    • Conducted 22 semi-structured interviews with AI practitioners across academia, industry, and open-source communities.
    • Thematic analysis of interview transcripts to identify patterns in dataset creation, risk definition, and evaluation practices.
    • Highlighted three critical moments in red teaming: task definition, dataset development, and evaluation.

Results

  • Concrete findings:
    • Practitioners conceptualize red teaming either as exploration (searching for attack cases) or as classification (detecting predefined harmful content).
    • Adversarial datasets are created using three approaches: repurposing existing datasets, creating datasets from scratch, or deriving datasets from human–LLM interactions.
    • Evaluation methods rely heavily on automated classifiers and LLM judges, but human evaluation remains critical for nuanced judgments.
  • Advantage over baselines:
    • Identified socio-technical gaps in current red teaming practices, such as neglecting context, interaction types, and user-specific needs.
    • Proposed actionable recommendations to address these gaps, emphasizing interdisciplinary collaboration and participatory approaches.
  • Experiments / evaluation:
    • Interviews covered diverse practitioner profiles, including PhD students, researchers, and engineers from eight countries.
    • Analysis revealed challenges in defining harm, evaluating adversarial outputs, and ensuring dataset diversity.
  • Limitations and future work:
    • Limited representation of industry practitioners in the study sample.
    • Future work should explore participatory methods to engage end-users and domain experts in red teaming evaluations.

Summary

This study investigates the socio-technical practices of red teaming LLMs by analyzing how practitioners define tasks, develop adversarial datasets, and evaluate model vulnerabilities. Findings reveal that harmfulness is constructed through cultural norms and technical decisions embedded in dataset creation and evaluation. The study highlights gaps in addressing context, interaction types, and user-specific needs, offering recommendations for HCI researchers to design more inclusive and context-aware red teaming approaches. These insights contribute to advancing AI safety evaluation by integrating socio-technical dimensions into red teaming practices.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222348/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790792
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Honorable Mention
group
Authors
4 authors
sell
Subtopics
Explainable AI (XAI), AI Ethics, Fairness & Accountability, Algorithmic Transparency & Auditability
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Data Scientists & Analysts
article
Content Status
Full text indexed
hub
Related Papers
10 related papers