Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions
Authors
Title of the Paper
Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions
Paper Information
- Subject Area: Human-Computer Interaction (HCI) and the application of Large Language Models (LLMs) in persona creation
- Keywords: Artificial Intelligence, Large Language Models, Human-Computer Interaction, Personas, Evaluation
Research Background and Issues
-
Identified Problems or Challenges:
- Although large language models (e.g., GPT-4) are capable of generating virtual personas, questions remain about whether these personas exhibit diversity, authenticity, and low bias.
- Do AI-generated personas overemphasize certain characteristics (e.g., age or profession), leading to bias?
- Are the generated personas of sufficient quality to meet real-world design needs?
-
Significance of the Issues:
- Personas are fictional representations used to convey user needs, behaviors, and preferences, serving as critical tools in the design process.
- If personas are biased or lack diversity, designers may overlook marginalized user groups, reducing algorithmic fairness and technological applicability.
-
Research Motivation and Related Work:
- Previous studies have demonstrated the potential of LLMs in persona generation but have not specifically examined issues of diversity and bias.
- The authors argue that LLM-generated personas require more systematic evaluation, including assessments by user experience (UX) researchers and domain experts.
Proposed Solution
-
Proposed Solution:
- The authors propose a framework for generating personas using LLMs (e.g., GPT-4) and analyze the characteristics, quality, and potential biases of the generated personas through multi-level evaluations (including internal assessments and expert reviews).
-
Innovative Aspects of the Solution:
- Using different prompts to control persona characteristics (e.g., gender, profession, and age).
- Creating a large-scale dataset (450 personas) to evaluate specific dimensions (e.g., diversity and consistency).
- Combining internal evaluations and expert feedback to examine the practical utility of personas in design contexts.
-
Implementation Steps and Key Techniques:
- Generating personas using GPT-4 and designing various prompts to explore the controllability of persona characteristics (e.g., whether gender is assigned by default).
- Conducting quantitative and qualitative analyses of persona descriptions, including demographic characteristics (age, profession, country), textual patterns (word count, diversity), biases, and attributes.
- Utilizing internal and external evaluations from UX researchers and healthcare domain experts to analyze the perceived quality of personas (e.g., credibility, consistency, informational value).
Research Findings
-
Specific Findings:
-
Diversity Analysis:
- Age range spans from 17 to 67 years, with an average of 37 years, but there is a slight bias toward younger groups.
- Gender achieves perfect balance (equal numbers of male and female personas) when unspecified.
- Professions are diverse, covering 201 different jobs, though slight gender stereotypes are observed (e.g., women are more likely to be assigned caregiving roles).
- Nationalities are predominantly concentrated in the United States (accounting for 86%), despite no explicit location specified in the prompts, indicating a "U.S.-centric" bias in the model.
-
Persona Evaluation:
- Generated personas scored highly in consistency, credibility, and design informational value, with low levels of stereotyping.
- Internal evaluators and domain experts generally found the persona descriptions credible, relatable, and positive.
- Some experts noted that certain persona descriptions (e.g., personas with addiction backgrounds) were unrealistic or inconsistent.
-
Text Analysis:
- Persona descriptions averaged 381 words, with text length slightly decreasing with age.
- Persona descriptions exhibited high diversity, with no repeated content.
-
Pain Point Analysis:
- Pain point analysis revealed that persona challenges (e.g., financial issues or interpersonal difficulties) were logically constructed and correlated with the type of addiction described.
-
-
Comparison with Existing Solutions:
- Compared to other studies with limited samples (e.g., 5–10 personas), this study’s 450-persona dataset provides more statistically significant results.
- The inclusion of expert evaluations further validates the credibility and relevance of the generated personas for real-world applications.
-
Limitations and Future Directions:
- Data generation is based on GPT-4’s general knowledge base, which may include biases from its training data, such as the overrepresentation of U.S. users.
- Lack of independent validation of persona authenticity. Future work could compare AI-generated personas with real-world data.
- Improve the robustness of prompt engineering, such as incorporating more characteristics of minority groups.
Future Research Directions:
- Explore how different prompts and data management techniques can reduce biases in persona generation.
- Compare the effectiveness of LLM-generated personas with traditionally handcrafted personas in practical design applications.
- Further investigate the application of persona generation in social domains (e.g., public health), such as more accurately describing characteristics of populations with addiction issues.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Do AI-generated virtual persona descriptions exhibit diversity, authenticity, and low bias?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- Do LLM-generated persona descriptions overemphasize certain traits (e.g., age or occupation), causing bias?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- Can generated persona descriptions meet quality requirements for actual design needs?Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
Practical Problems
1- Designers may use biased or insufficiently diverse virtual user models, neglecting marginalized group needs.Category: Bias and Fairness in Large Language Models and GenAISimilar questionsarrow_forward
- 75%
Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models
UIST '23· Human-LLM Collaboration
- 67%
Planning for Natural Language Failures with the AI Playbook
CHI '21· Human-LLM Collaboration +2
- 67%
A study of UX Practitioners Roles in Designing Real-World, Enterprise ML Systems
CHI '22· Human-LLM Collaboration +2
- 67%
Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges
CHI '23· Human-LLM Collaboration +2
- 67%
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
CHI '25· Human-LLM Collaboration +2
- 67%
Understanding Socio-technical Factors Configuring AI Non-Use in UX Work Practices
CHI '25· Human-LLM Collaboration +2
- 67%
The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language Models
CHI '26· Dark Patterns Recognition +2
- 67%
Be Friendly, Not Friends: How LLM Sycophancy Shapes User Trust
CHI '26· Human-LLM Collaboration +2
- 67%
Deception at Scale: Deceptive Designs in 1K LLM-Generated E-Commerce Components
CHI '26· AI Ethics, Fairness & Accountability +2
- 67%
Prompt Coaching for Inclusiveness: A Media Literacy Approach to Increase Users’ Awareness of Algorithmic Bias and Prompting Efficacy
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)