Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case Study
Best PaperAuthors
Document Title
Evaluating Large Language Models in Generating Synthetic HCI Research Data: A Case Study
Document Information
- Domain: Human-Computer Interaction (HCI) research, focusing on the use of large language models (LLMs) to generate synthetic research data
- Keywords: user experience, user modeling, language models, GPT-3, HCI simulation, synthetic data
Research Background and Issues
-
Issues and Challenges:
- In HCI research, collecting real user data (e.g., interviews and surveys) often requires significant cost and time.
- Most computational user modeling is limited to relatively simple behavioral metrics, making it difficult to capture aspects such as user emotions, motivations, and complex subjective experiences.
- It remains unclear whether synthetic data can supplement real data in design and validation, and the quality of large language models in generating data for HCI has not yet been evaluated.
-
Significance:
- If LLMs can efficiently generate research data resembling real-world data, they could significantly accelerate the iteration process of HCI research and experiments.
- However, LLMs may also be misused by malicious users to generate falsified data, potentially undermining research reliability.
-
Research Motivation and Related Work:
- This study builds on prior computational user modeling work, particularly focusing on the potential of GPT-3 to generate user experience-related data.
- Unlike previous studies that emphasize the logical capabilities of language models, this study focuses on the authenticity of subjective experience data generated by LLMs.
Solution
-
Method or Solution:
- Explore the ability of OpenAI's GPT-3 model to generate synthetic data by simulating human participants' responses to user experience questionnaires.
- Examine whether the generated text can be distinguished as "human-created" or "computer-generated."
- Analyze error types in synthetic data and compare the content of GPT-3-generated data with human data.
-
Innovative Aspects of the Solution:
- The first study to validate the authenticity of synthetic data generated by LLMs in the HCI domain.
- Introduces automated coding for thematic analysis and semantic embedding comparison to evaluate structural differences between synthetic and real data.
- Proposes a multi-experiment validation framework, including data distinguishability tests, error analysis, and statistical similarity studies.
-
Implementation Steps and Key Techniques:
- Experimental Design:
- Propose three experiments to evaluate the distinguishability, error types, and content differences of generated data.
- Data Generation and Processing:
- Use GPT-3 to generate questionnaire responses based on experimental prompts in the HCI domain.
- Compare with human data through quantitative and qualitative analyses to reveal consistencies and differences.
- Qualitative and Quantitative Modeling:
- Automated semantic coding (including topic mining, embedding dimensionality reduction, and clustering analysis).
- Use quality metrics such as Frechet distance to quantify the "human-like" nature of synthetic data.
- Experimental Design:
Research Results
-
Specific Findings:
- Experiment 1: Distinguishability:
- On average, participants tended to mistakenly classify GPT-3-generated text as "human-created" (accuracy ~40.45%).
- Experiment 2: Error Analysis:
- Common errors included incoherent text, fabricated facts, and failure to answer questions.
- Experiment 3: Content Differences:
- GPT-3-generated text showed high consistency with real data on certain topics, such as emotional expressions about games.
- However, issues like "lack of diversity" were observed, such as repetitive bias in responses about specific games.
- Experiment 1: Distinguishability:
-
Experimental and Evaluation Results:
- Compared to human data, the largest GPT-3 model (davinci) produced data with higher "human-like" quality.
- Evaluation indicated that the thematic distribution and quality of generated data depended on prompt parameters and specific model versions.
- The study revealed that the latest text-davinci-002 model generated high-quality content but lacked diversity.
-
Advantages Over Existing Solutions:
- Demonstrated the potential of LLMs to generate subjective user experience data, which can be used for preliminary exploration or design guidance in HCI experiments.
- Focused on experimental validation, addressing gaps in reliability assessments of generated content in language model research.
-
Limitations and Future Directions:
- Only tested game experiences as a specific HCI context; the generalizability to other research topics remains unverified.
- Further optimization of prompt design and generation strategies is needed to improve model applicability.
- Calls for broader testing of LLMs' data generation capabilities across different user groups, domains, and cultural contexts.
- Highlights the risk of misuse, such as challenges to the authenticity of online survey data, and suggests developing new verification tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How authentic is synthetic data generated by LLMs in UX research?Category: Information Organization, Document Analysis, and Qualitative ResearchSimilar questionsarrow_forward
- Can GPT-3-generated text be distinguished from human-created text?Category: Information Organization, Document Analysis, and Qualitative ResearchSimilar questionsarrow_forward
- How do generated UX data compare with real data in thematic distribution and diversity?Category: Information Organization, Document Analysis, and Qualitative ResearchSimilar questionsarrow_forward
Practical Problems
1- Collecting real data in HCI research is costly and time-consuming.Category: Information Organization, Document Analysis, and Qualitative ResearchSimilar questionsarrow_forward
- 100%
Navigating the Unknown: A Chat-Based Collaborative Interface for Personalized Exploratory Tasks
IUI '25· Human-LLM Collaboration +1
- 75%
MAPLE: Mobile App Prediction Leveraging Large Language Model Embeddings
UbiComp '24· Human-LLM Collaboration
- 75%
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
UIST '24· Human-LLM Collaboration
- 67%
AI as an Agent and Collaborative Space: Exploring the role of Generative AI in Small Group Synchronous and Asynchronous Collaborative Dynamics
CHI '26· Human-LLM Collaboration +2
- 60%
Cicero: Multi-Turn, Contextual Argumentation for Accurate Crowdsourcing
CHI '19· Human-LLM Collaboration +1
- 60%
AffectiveSpotlight: Facilitating the Communication of Affective Responses from Audience Members during Online Presentations
CHI '21· Social & Collaborative VR +1
- 60%
OneLabeler: A Flexible System for Building Data Labeling Tools
CHI '22· Crowdsourcing Task Design & Quality Control
- 60%
DynEx: Dynamic Code Synthesis with Structured Design Exploration for Accelerated Exploratory Programming
CHI '25· Human-LLM Collaboration +1
- 60%
SummAct: Uncovering User Intentions Through Interactive Behaviour Summarisation
CHI '25· Human-LLM Collaboration +1
- 60%
Divergence or Convergence? A Deep Insight into the Crowd Collaboration and its Productivity in Open Source Software based on Entropy
CHI '26· Crowdsourcing Task Design & Quality Control +1
Based on Jaccard similarity of research subtopics & professions (≥60%)