From Narrative to Numbers: Evaluating Survey Questionnaires with Large Language Models
Authors
As the field of intelligent interfaces is evolving, there is also a growing need for feedback mechanisms that are both expressive for participants and also contain reliable and useful information for the researcher conducting the study for survey data collection. Our study involves exploring two different kinds of survey methods: a standardized slider scale for web-based surveys and a free-form text input with a Large Language Model (LLM) acting as a backbone. The experiment includes 36 participants completing a 4×4 sliding-tile game at two different levels (easy and hard) with difficulty standardized via Manhattan-distance targets. The response mode order was counterbalanced across two sequences. This task aimed to evaluate the accuracy and quality of participant responses through different survey methodologies. Our key findings are that the LLM survey results are equivalent to the ones reported through the Web-Based slider scale questionnaire method. Our contribution is an intelligent framework that allows text-based reflections within an adaptive survey interface, helping both participants to express their experiences naturally and also researchers to gain valuable information about their system.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Rationalizer: Leveraging LLM to Support User Providing the Rationales Behind the Rating of Likert Scale Questionnaires
IUI '26· Human-LLM Collaboration +2
- 86%
Evalet: Evaluating Large Language Models through Functional Fragmentation
CHI '26· Human-LLM Collaboration +3
- 86%
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
CHI '26· Human-LLM Collaboration +3
- 86%
Integrating Complementary Feature Sets for Human-AI Decision-Making
IUI '26· Human-LLM Collaboration +3
- 83%
AI of Oz: Enhancing Wizard of Oz Studies in HCI with AI Assistance for Human Moderation
CHI '26· Human-LLM Collaboration +2
- 71%
Mapping the Wizards' Path: A Systematic Review of Wizard-of-Oz in HCI
CHI '26· Participatory Design +3
- 71%
Evaluating Generative AI in the Lab: Methodological Challenges and Guidelines
IUI '26· Generative AI (Text, Image, Music, Video) +3
- 63%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 63%
SemTabla: A Human-in-the-Loop Framework for Semantic Enrichment and Validation of Data Tables
CHI '26· Explainable AI (XAI) +3
- 63%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)