TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
Authors
Paper Title
TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
Publication Info
- Topic area: Cultural representation in AI-generated narratives
- Keywords: Cultural misrepresentation, large language models, storytelling, taxonomy, Indian culture, multilingual evaluation, generative AI, cultural competence, Indic languages, AI evaluation
Background and Problem
- Problem / challenge: Large language models (LLMs) often misrepresent diverse cultural identities, particularly in open-ended generative tasks like storytelling. Existing evaluations lack frameworks to assess nuanced cultural appropriateness and misrepresentation.
- Significance: Misrepresentation in AI-generated narratives can perpetuate stereotypes, inaccuracies, and harm marginalized communities, especially in non-Western contexts. Addressing this issue is crucial for equitable AI development.
- Motivation and related work: Prior research has focused on cultural competence in AI through datasets and quantitative metrics but has largely ignored open-ended generative tasks. Evaluations have been limited to Western contexts or narrow proxies of culture. This paper builds on these gaps by focusing on Indian cultural diversity and developing a taxonomy of misrepresentations.
Solution
- Proposed approach: TALES (Taxonomy and Analysis of culture representation in LLM-generated Stories), a community-centered framework to evaluate cultural misrepresentations in LLM-generated stories.
- Novelty:
- Development of TALES-Tax, a taxonomy of seven categories of cultural misrepresentation.
- Large-scale human evaluation of cultural misrepresentations across six LLMs in English and 13 Indic languages.
- Creation of TALES-QA, a question bank to assess cultural knowledge in models.
- Procedure and key techniques:
- Conducted focus groups (N=9) and surveys (N=15) to identify misrepresentation categories.
- Developed TALES-Tax through reflexive thematic analysis of participant feedback.
- Evaluated six LLMs using 108 annotators from 71 regions in India, collecting 2,925 annotations across 540 stories.
- Converted misrepresentation annotations into 1,683 standalone questions in TALES-QA, validated by human experts.
Results
- Concrete findings:
- 88% of LLM-generated stories contained cultural misrepresentations.
- Average misrepresentation per story: 5.42 (approximately one misrepresentation every five sentences).
- Models answered cultural knowledge questions with 76.9% accuracy in English and 59.8% in Indic languages.
- Advantage over baselines:
- TALES-Tax provides a structured framework for evaluating cultural misrepresentations, addressing gaps in prior work.
- TALES-QA highlights discrepancies between models’ cultural knowledge and their generative capabilities.
- Experiments / evaluation:
- Evaluated stories generated by six LLMs, including GPT 4.1, Gemini 2.5 Pro, and four open-source models.
- Used Hofstede’s cultural onion model to design prompts targeting symbols, heroes, rituals, and values.
- Analyzed misrepresentation frequency across languages, regions, and culturally specific items (CSIs).
- Limitations and future work:
- Focus groups and surveys were conducted only in English, potentially missing nuances in other languages.
- Evaluation centered on Indian cultural contexts; broader applicability to global cultures remains unexplored.
- Prompt design may influence observed misrepresentation rates; future work could explore richer prompts and external knowledge integration.
Summary
TALES introduces a taxonomy (TALES-Tax) and evaluation framework for analyzing cultural misrepresentations in LLM-generated stories, focusing on Indian cultural diversity. Through large-scale human evaluation, the study reveals widespread misrepresentation, particularly in Indic languages and lesser-known regions. The TALES-QA question bank demonstrates that models possess cultural knowledge but fail to apply it effectively in generative tasks. This work underscores the need for improving LLMs’ cultural competence and provides tools for future research in multilingual and culturally sensitive AI systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing
CHI '26· Human-LLM Collaboration +2
- 63%
Beyond Claiming Sovereign AI: Motivations, Challenges, and Contradictions in Developing and Deploying Local Foundation Models in South Korea
CHI '26· Human-LLM Collaboration +3
- 63%
LLMs Homogenize Values in Constructive Arguments on Value-Laden Topics
CHI '26· Human-LLM Collaboration +3
- 63%
Exposing the Ideology of Large Language Models with Creative Practices
DIS '25· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)