Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions

Human-LLM CollaborationAI Ethics, Fairness & AccountabilityUI/UX DesignersAI/ML Researchers & Engineers

Title of the Paper

Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions

Paper Information

  • Subject Area: Human-Computer Interaction (HCI) and the application of Large Language Models (LLMs) in persona creation
  • Keywords: Artificial Intelligence, Large Language Models, Human-Computer Interaction, Personas, Evaluation

Research Background and Issues

  • Identified Problems or Challenges:

    • Although large language models (e.g., GPT-4) are capable of generating virtual personas, questions remain about whether these personas exhibit diversity, authenticity, and low bias.
    • Do AI-generated personas overemphasize certain characteristics (e.g., age or profession), leading to bias?
    • Are the generated personas of sufficient quality to meet real-world design needs?
  • Significance of the Issues:

    • Personas are fictional representations used to convey user needs, behaviors, and preferences, serving as critical tools in the design process.
    • If personas are biased or lack diversity, designers may overlook marginalized user groups, reducing algorithmic fairness and technological applicability.
  • Research Motivation and Related Work:

    • Previous studies have demonstrated the potential of LLMs in persona generation but have not specifically examined issues of diversity and bias.
    • The authors argue that LLM-generated personas require more systematic evaluation, including assessments by user experience (UX) researchers and domain experts.

Proposed Solution

  • Proposed Solution:

    • The authors propose a framework for generating personas using LLMs (e.g., GPT-4) and analyze the characteristics, quality, and potential biases of the generated personas through multi-level evaluations (including internal assessments and expert reviews).
  • Innovative Aspects of the Solution:

    1. Using different prompts to control persona characteristics (e.g., gender, profession, and age).
    2. Creating a large-scale dataset (450 personas) to evaluate specific dimensions (e.g., diversity and consistency).
    3. Combining internal evaluations and expert feedback to examine the practical utility of personas in design contexts.
  • Implementation Steps and Key Techniques:

    1. Generating personas using GPT-4 and designing various prompts to explore the controllability of persona characteristics (e.g., whether gender is assigned by default).
    2. Conducting quantitative and qualitative analyses of persona descriptions, including demographic characteristics (age, profession, country), textual patterns (word count, diversity), biases, and attributes.
    3. Utilizing internal and external evaluations from UX researchers and healthcare domain experts to analyze the perceived quality of personas (e.g., credibility, consistency, informational value).

Research Findings

  • Specific Findings:

    1. Diversity Analysis:

      • Age range spans from 17 to 67 years, with an average of 37 years, but there is a slight bias toward younger groups.
      • Gender achieves perfect balance (equal numbers of male and female personas) when unspecified.
      • Professions are diverse, covering 201 different jobs, though slight gender stereotypes are observed (e.g., women are more likely to be assigned caregiving roles).
      • Nationalities are predominantly concentrated in the United States (accounting for 86%), despite no explicit location specified in the prompts, indicating a "U.S.-centric" bias in the model.
    2. Persona Evaluation:

      • Generated personas scored highly in consistency, credibility, and design informational value, with low levels of stereotyping.
      • Internal evaluators and domain experts generally found the persona descriptions credible, relatable, and positive.
      • Some experts noted that certain persona descriptions (e.g., personas with addiction backgrounds) were unrealistic or inconsistent.
    3. Text Analysis:

      • Persona descriptions averaged 381 words, with text length slightly decreasing with age.
      • Persona descriptions exhibited high diversity, with no repeated content.
    4. Pain Point Analysis:

      • Pain point analysis revealed that persona challenges (e.g., financial issues or interpersonal difficulties) were logically constructed and correlated with the type of addiction described.
  • Comparison with Existing Solutions:

    • Compared to other studies with limited samples (e.g., 5–10 personas), this study’s 450-persona dataset provides more statistically significant results.
    • The inclusion of expert evaluations further validates the credibility and relevance of the generated personas for real-world applications.
  • Limitations and Future Directions:

    1. Data generation is based on GPT-4’s general knowledge base, which may include biases from its training data, such as the overrepresentation of U.S. users.
    2. Lack of independent validation of persona authenticity. Future work could compare AI-generated personas with real-world data.
    3. Improve the robustness of prompt engineering, such as incorporating more characteristics of minority groups.

    Future Research Directions:

    • Explore how different prompts and data management techniques can reduce biases in persona generation.
    • Compare the effectiveness of LLM-generated personas with traditionally handcrafted personas in practical design applications.
    • Further investigate the application of persona generation in social domains (e.g., public health), such as more accurately describing characteristics of populations with addiction issues.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/148256/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642036
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, AI Ethics, Fairness & Accountability
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers