Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case Study

Best Paper
Human-LLM CollaborationCrowdsourcing Task Design & Quality ControlSoftware Engineers & DevelopersHCI Researchers

Document Title

Evaluating Large Language Models in Generating Synthetic HCI Research Data: A Case Study

Document Information

  • Domain: Human-Computer Interaction (HCI) research, focusing on the use of large language models (LLMs) to generate synthetic research data
  • Keywords: user experience, user modeling, language models, GPT-3, HCI simulation, synthetic data

Research Background and Issues

  • Issues and Challenges:

    • In HCI research, collecting real user data (e.g., interviews and surveys) often requires significant cost and time.
    • Most computational user modeling is limited to relatively simple behavioral metrics, making it difficult to capture aspects such as user emotions, motivations, and complex subjective experiences.
    • It remains unclear whether synthetic data can supplement real data in design and validation, and the quality of large language models in generating data for HCI has not yet been evaluated.
  • Significance:

    • If LLMs can efficiently generate research data resembling real-world data, they could significantly accelerate the iteration process of HCI research and experiments.
    • However, LLMs may also be misused by malicious users to generate falsified data, potentially undermining research reliability.
  • Research Motivation and Related Work:

    • This study builds on prior computational user modeling work, particularly focusing on the potential of GPT-3 to generate user experience-related data.
    • Unlike previous studies that emphasize the logical capabilities of language models, this study focuses on the authenticity of subjective experience data generated by LLMs.

Solution

  • Method or Solution:

    • Explore the ability of OpenAI's GPT-3 model to generate synthetic data by simulating human participants' responses to user experience questionnaires.
    • Examine whether the generated text can be distinguished as "human-created" or "computer-generated."
    • Analyze error types in synthetic data and compare the content of GPT-3-generated data with human data.
  • Innovative Aspects of the Solution:

    • The first study to validate the authenticity of synthetic data generated by LLMs in the HCI domain.
    • Introduces automated coding for thematic analysis and semantic embedding comparison to evaluate structural differences between synthetic and real data.
    • Proposes a multi-experiment validation framework, including data distinguishability tests, error analysis, and statistical similarity studies.
  • Implementation Steps and Key Techniques:

    1. Experimental Design:
      • Propose three experiments to evaluate the distinguishability, error types, and content differences of generated data.
    2. Data Generation and Processing:
      • Use GPT-3 to generate questionnaire responses based on experimental prompts in the HCI domain.
      • Compare with human data through quantitative and qualitative analyses to reveal consistencies and differences.
    3. Qualitative and Quantitative Modeling:
      • Automated semantic coding (including topic mining, embedding dimensionality reduction, and clustering analysis).
      • Use quality metrics such as Frechet distance to quantify the "human-like" nature of synthetic data.

Research Results

  • Specific Findings:

    1. Experiment 1: Distinguishability:
      • On average, participants tended to mistakenly classify GPT-3-generated text as "human-created" (accuracy ~40.45%).
    2. Experiment 2: Error Analysis:
      • Common errors included incoherent text, fabricated facts, and failure to answer questions.
    3. Experiment 3: Content Differences:
      • GPT-3-generated text showed high consistency with real data on certain topics, such as emotional expressions about games.
      • However, issues like "lack of diversity" were observed, such as repetitive bias in responses about specific games.
  • Experimental and Evaluation Results:

    • Compared to human data, the largest GPT-3 model (davinci) produced data with higher "human-like" quality.
    • Evaluation indicated that the thematic distribution and quality of generated data depended on prompt parameters and specific model versions.
    • The study revealed that the latest text-davinci-002 model generated high-quality content but lacked diversity.
  • Advantages Over Existing Solutions:

    • Demonstrated the potential of LLMs to generate subjective user experience data, which can be used for preliminary exploration or design guidance in HCI experiments.
    • Focused on experimental validation, addressing gaps in reliability assessments of generated content in language model research.
  • Limitations and Future Directions:

    • Only tested game experiences as a specific HCI context; the generalizability to other research topics remains unverified.
    • Further optimization of prompt design and generation strategies is needed to improve model applicability.
    • Calls for broader testing of LLMs' data generation capabilities across different user groups, domains, and cultural contexts.
    • Highlights the risk of misuse, such as challenges to the authenticity of online survey data, and suggests developing new verification tools.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95751/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3580688
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
Best Paper
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, Crowdsourcing Task Design & Quality Control
work
Professions
Software Engineers & Developers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers