TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories

Human-LLM CollaborationAI Ethics, Fairness & AccountabilityLow-Resource Languages & Digital InclusionAI/ML Researchers & EngineersHCI ResearchersSociologists & Anthropologists

Paper Title

TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories

Publication Info

  • Topic area: Cultural representation in AI-generated narratives
  • Keywords: Cultural misrepresentation, large language models, storytelling, taxonomy, Indian culture, multilingual evaluation, generative AI, cultural competence, Indic languages, AI evaluation

Background and Problem

  • Problem / challenge: Large language models (LLMs) often misrepresent diverse cultural identities, particularly in open-ended generative tasks like storytelling. Existing evaluations lack frameworks to assess nuanced cultural appropriateness and misrepresentation.
  • Significance: Misrepresentation in AI-generated narratives can perpetuate stereotypes, inaccuracies, and harm marginalized communities, especially in non-Western contexts. Addressing this issue is crucial for equitable AI development.
  • Motivation and related work: Prior research has focused on cultural competence in AI through datasets and quantitative metrics but has largely ignored open-ended generative tasks. Evaluations have been limited to Western contexts or narrow proxies of culture. This paper builds on these gaps by focusing on Indian cultural diversity and developing a taxonomy of misrepresentations.

Solution

  • Proposed approach: TALES (Taxonomy and Analysis of culture representation in LLM-generated Stories), a community-centered framework to evaluate cultural misrepresentations in LLM-generated stories.
  • Novelty:
    1. Development of TALES-Tax, a taxonomy of seven categories of cultural misrepresentation.
    2. Large-scale human evaluation of cultural misrepresentations across six LLMs in English and 13 Indic languages.
    3. Creation of TALES-QA, a question bank to assess cultural knowledge in models.
  • Procedure and key techniques:
    • Conducted focus groups (N=9) and surveys (N=15) to identify misrepresentation categories.
    • Developed TALES-Tax through reflexive thematic analysis of participant feedback.
    • Evaluated six LLMs using 108 annotators from 71 regions in India, collecting 2,925 annotations across 540 stories.
    • Converted misrepresentation annotations into 1,683 standalone questions in TALES-QA, validated by human experts.

Results

  • Concrete findings:
    • 88% of LLM-generated stories contained cultural misrepresentations.
    • Average misrepresentation per story: 5.42 (approximately one misrepresentation every five sentences).
    • Models answered cultural knowledge questions with 76.9% accuracy in English and 59.8% in Indic languages.
  • Advantage over baselines:
    • TALES-Tax provides a structured framework for evaluating cultural misrepresentations, addressing gaps in prior work.
    • TALES-QA highlights discrepancies between models’ cultural knowledge and their generative capabilities.
  • Experiments / evaluation:
    • Evaluated stories generated by six LLMs, including GPT 4.1, Gemini 2.5 Pro, and four open-source models.
    • Used Hofstede’s cultural onion model to design prompts targeting symbols, heroes, rituals, and values.
    • Analyzed misrepresentation frequency across languages, regions, and culturally specific items (CSIs).
  • Limitations and future work:
    • Focus groups and surveys were conducted only in English, potentially missing nuances in other languages.
    • Evaluation centered on Indian cultural contexts; broader applicability to global cultures remains unexplored.
    • Prompt design may influence observed misrepresentation rates; future work could explore richer prompts and external knowledge integration.

Summary

TALES introduces a taxonomy (TALES-Tax) and evaluation framework for analyzing cultural misrepresentations in LLM-generated stories, focusing on Indian cultural diversity. Through large-scale human evaluation, the study reveals widespread misrepresentation, particularly in Indic languages and lesser-known regions. The TALES-QA question bank demonstrates that models possess cultural knowledge but fail to apply it effectively in generative tasks. This work underscores the need for improving LLMs’ cultural competence and provides tools for future research in multilingual and culturally sensitive AI systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223387/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790519
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, AI Ethics, Fairness & Accountability, Low-Resource Languages & Digital Inclusion
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
4 related papers