Interview-Informed Generative Agents for Product Discovery: A Validation Study
Honorable MentionAuthors
Paper Title
Interview-Informed Generative Agents for Product Discovery: A Validation Study
Publication Info
- Topic area: Application of generative AI agents in product discovery and user simulation.
- Keywords: Large language models, generative agents, product discovery, user simulation, AI concept testing, Technology Acceptance Model, Net Promoter Score, qualitative feedback, document workflows, HCI.
Background and Problem
- Problem / challenge: While large language models (LLMs) have demonstrated capabilities in social science simulations, their applicability in product discovery contexts remains unclear, particularly for simulating user responses to novel concepts.
- Significance: Understanding user responses to early-stage product concepts is critical for design research and product development, but traditional methods are resource-intensive.
- Motivation and related work: Previous studies, such as Park et al. (2023), showed high accuracy for LLMs in replicating survey responses, but their utility in product discovery—where responses are constructed for hypothetical scenarios—has not been validated. This study explores whether interview-informed generative agents can simulate user evaluations of AI concepts.
Solution
- Proposed approach: Interview-informed generative agents grounded in detailed user interviews simulate responses to AI document workflow concepts.
- Novelty:
- First empirical validation of generative agents for product discovery tasks.
- Characterization of simulation fidelity as "distribution-calibrated but identity-imprecise."
- Practical guidance for integrating simulations into product development workflows.
- Procedure and key techniques:
- Conduct in-depth interviews with knowledge workers to capture workflows and technology adoption patterns.
- Use interview transcripts to create personalized generative agents.
- Compare agent responses to actual participant responses using quantitative (TAM, NPS) and qualitative metrics.
- Evaluate simulation fidelity at individual and population levels.
Results
- Concrete findings:
- Interview-based agents achieve 67% of human-human agreement in individual-level accuracy but fail to replicate specific participants reliably.
- Agents approximate population-level response distributions with better alignment than scratchpad-only or no-information baselines.
- Open-ended responses from agents capture high-level themes but lack nuance, emotional variability, and experiential fidelity.
- Advantage over baselines:
- Interview-based agents outperform scratchpad-only and no-information agents in population-level alignment (Wasserstein distance).
- Agents capture realistic variability in response distributions, especially for negative feedback.
- Experiments / evaluation:
- Study involved 51 participants evaluating four AI document workflow concepts (Multidoc Q&A Assistant, Smart Highlights Assistant, Audio Assistant, Workflow Actions Assistant).
- Metrics included MAE, correlation, Gwet’s AC2, and qualitative evaluations (sentiment, explanation, topic coverage, tone).
- Cost analysis showed agent simulations are faster (4 minutes per concept) and cheaper ($1.27 per concept) than human evaluations (30 minutes, $12.50 per concept).
- Limitations and future work:
- Small sample size (51 participants) limits statistical power.
- Agents fail to capture individual-level fidelity and nuanced qualitative insights.
- Future work should explore richer modalities, improved agent architectures, and larger, more diverse populations.
Summary
This study validates the use of interview-informed generative agents for simulating user responses in product discovery contexts. While agents achieve population-level calibration and distributional accuracy, they fail to replicate individual-level responses reliably. Practical applications include low-cost concept screening and directional exploration in early-stage design, but authentic user interviews remain essential for nuanced insights. Future research should focus on improving agent architectures and expanding validation across diverse domains and modalities.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
CHI '26· Human-LLM Collaboration +3
- 88%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 88%
"Here, Let Me Help": An Empirical Study of User Interventions in Human–Web Agent Collaboration
CHI '26· Human-LLM Collaboration +3
- 88%
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
CHI '26· Human-LLM Collaboration +3
- 88%
The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows
CHI '26· Human-LLM Collaboration +3
- 88%
Live in the Loop: Rapid Run-time Feedback for Prompts
CHI '26· Human-LLM Collaboration +3
- 78%
"Shall We Dig Deeper?": Designing and Evaluating Strategies for LLM Agents to Advance Knowledge Co-Construction in Asynchronous Online Discussions
CHI '26· Human-LLM Collaboration +3
- 78%
“It Became My Buddy, But I’m Not Afraid to Disagree”: A Multi-Session Study of UX Evaluators Collaborating with Conversational AI Assistants
CHI '26· Human-LLM Collaboration +4
- 78%
TurnStyle: A Framework for Analyzing Human Conversational Behaviors to Predict Success in LLM-Assisted Tasks
CHI '26· Human-LLM Collaboration +3
- 78%
ImpReSS: Designing and Evaluating a Lightweight Implicit Recommender System in Conversational Support Agents
IUI '26· Human-LLM Collaboration +4
Based on Jaccard similarity of research subtopics & professions (≥60%)