A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism

Human-LLM CollaborationAI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasAI/ML Researchers & EngineersLawyers & Legal ResearchersHCI Researchers

Research Background and Issues

  • Issues and Challenges: The authors focus on the performance of large language models (LLMs) and humans in handling subtle gender discrimination scenarios in subjective decision-making, particularly the differences in the distribution and diversity of perspectives. Unlike tasks based on objective data, subjective decision-making relies on contextual interpretation and value judgments.
  • Significance: Subtle gender discrimination is a covert form of bias, with interpretations varying based on personal background, posing cultural and social challenges for its identification and handling. AI-assisted subjective decision-making offers a potential tool for accommodating diverse perspectives, but it may also come with biases and limited viewpoints.
  • Research Motivation and Related Work: Previous studies have focused on the capabilities of LLMs in objective tasks, but their performance in subjective decision-making has not been fully assessed. Subjective tasks require attention to complex issues such as "diversity of perspectives" and "multi-perspective understanding."

Solution

  • Methodology: The authors conduct a comparative analysis of the distribution of perspectives and stances between humans and various LLMs (e.g., GPT-3, GPT-3.5, GPT-4, and Llama 3.1) when processing 101 subtle gender discrimination scenarios. Perspectives are categorized into three types: victim, perpetrator, and decision-maker; stances are divided into four types: sexist, not sexist, depends, and no stance.
  • Innovations:
    1. Framework for Perspective and Stance Classification: Systematically defines "perspectives" and "stances," capturing differences between LLM and human decision-making in complex social interactions.
    2. Comparative Method: Combines quantitative analysis (e.g., ANOVA tests) with qualitative evaluation to systematically identify performance differences between models and humans in subjective decision-making.
    3. No Benchmark Training: Human decision data is not treated as the "gold standard" but is instead compared equally with model-generated data.
  • Implementation Steps:
    1. Data Collection: Constructs a dataset of gender discrimination scenarios; collects responses from humans and LLMs through surveys and model-generated simulations.
    2. Coding and Analysis: Iteratively develops classification standards for "stances" and "perspectives," combining frequency statistics and consistency analysis across data sources.
    3. Model Comparison: Conducts a horizontal comparison between humans and LLMs (including different generational versions of GPT and the open-source Llama model).

Research Findings

  • Specific Results:
    1. Stance Distribution: Humans and GPT-3 tend to provide simpler and more subjective responses, such as denying discrimination in scenarios by reasoning "this is a fact." In contrast, GPT-3.5, GPT-4, and Llama 3.1 are more inclined to adopt a "depends" stance, demonstrating greater analytical caution and contextual sensitivity.
    2. Perspective Distribution: Llama 3.1 significantly favors the "victim" perspective (95% of responses), while GPT-3.5 and GPT-4 more frequently combine "victim" and "perpetrator" perspectives to offer multi-perspective explanations. In comparison, human participants provide fewer multi-perspective explanations.
    3. Consistency: GPT-3 exhibits the lowest stance consistency, whereas GPT-3.5 and Llama 3.1 demonstrate higher consistency, aligning with the low response consistency observed in humans.
    4. Behavioral Differences: GPT-3 mimics human experiences in some responses (even self-identifying as female), potentially leading to "anthropomorphism" and the "uncanny valley effect." In contrast, newer versions of GPT and Llama 3.1 more explicitly acknowledge the limitations of AI models.
  • Advantages:
    1. Multi-Perspective Opinions: Newer LLMs (especially GPT-3.5, GPT-4, and Llama 3.1) provide more comprehensive and multidimensional perspectives, offering insights that humans may not have considered, showcasing collaborative potential.
    2. Model Improvements: Earlier versions (e.g., GPT-3) tend to provide simpler and more direct responses, leaving room for improvement in addressing socially complex issues.
  • Limitations and Future Directions:
    1. Single Task Domain: This study focuses on subtle gender discrimination as a case application for subjective decision-making, but its applicability to other subjective domains (e.g., racism, ageism) remains to be validated.
    2. Risk of Data Leakage: Some scenarios may appear in the training data of LLMs, potentially leading to biased results.
    3. Prompt Limitations: The impact of different prompting strategies requires further exploration in the future.
    4. Human-AI Complementarity: Designing AI as a "supplementary opinion provider" (rather than merely mimicking humans) requires more sophisticated calibration techniques and application strategies.

Conclusion

This study compares the performance of humans and different generations of LLMs in subjective decision-making, introducing "stances" and "perspectives" as key parameters for evaluating and designing subjective AI-assisted systems. Through a systematic analysis of the differences in how LLMs and humans handle complex social issues, this paper not only provides important design insights for the application of LLMs but also lays the groundwork for future evaluations of AI ethics and safety.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189194/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713248
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias
work
Professions
AI/ML Researchers & Engineers, Lawyers & Legal Researchers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers