LAPS: Automating Hypothesis-Driven Statistical Analysis of Public Survey Using Large Language Models
Authors
Paper Title
LAPS: Automating Hypothesis-Driven Statistical Analysis of Public Survey Using Large Language Models
Publication Info
- Topic area: Automated frameworks for hypothesis-driven statistical analysis in social science research.
- Keywords: Public survey analysis, large language models, hypothesis testing, statistical planning, human-AI collaboration, cognitive workload, transparency, reproducibility, social science, research workflows.
Background and Problem
- Problem / challenge: Analyzing public survey data is complex due to inconsistent structures, large variable sets, and varying researcher expertise. Existing tools lack structured support for hypothesis-driven workflows, leading to instability and reduced reliability in results.
- Significance: Public surveys inform critical decisions in policy and academia, but their complexity limits accessibility and reproducibility, especially for novice researchers.
- Motivation and related work: Traditional tools (e.g., SPSS, R) and recent AI-based systems provide partial support but fail to address challenges specific to public survey data, such as complex survey designs and operationalization of hypotheses. General-purpose LLMs lack alignment with domain-specific workflows, producing outputs that may drift from researchers' intent.
Solution
- Proposed approach: LAPS (LLM-assisted Automated framework for Public Survey data analysis), an end-to-end system for hypothesis-driven statistical analysis of survey data.
- Novelty:
- Introduces a structured pipeline with four modules: Operationalization, Planning, Execution, and Reporting.
- Incorporates human-in-the-loop mechanisms to balance automation with researcher agency.
- Provides self-refinement loops for iterative improvement of variable selection and analysis plans.
- Aligns system behavior with domain-specific workflows to ensure transparency and reproducibility.
- Procedure and key techniques:
- Operationalization: Parses survey codebooks, validates hypotheses, selects variables, and refines them iteratively.
- Planning: Links descriptive statistics to structured analysis plans, incorporating complex survey design elements.
- Execution: Generates and debugs hybrid Python–R code for survey-weighted analyses.
- Reporting: Produces APA-style statistical reports with contextual interpretations and audit trails.
Results
- Concrete findings:
- SUS scores for LAPS were significantly higher than baseline tools (mean difference = 42.30, p < 0.001) and general-purpose LLMs (mean difference = 12.90, p < 0.05).
- NASA-TLX scores showed reduced cognitive workload compared to baseline tools (mean difference = 1.94, p < 0.001) and general-purpose LLMs (mean difference = 0.86, p < 0.01).
- TXAI scores indicated higher trust in LAPS explanations compared to general-purpose LLMs (overall mean difference = -0.87, p < 0.05).
- Advantage over baselines:
- Reduced cognitive burden through unified workflows.
- Higher usability and trust ratings compared to traditional tools and general-purpose LLMs.
- Enhanced researcher agency and analytical stability.
- Experiments / evaluation:
- User study with 12 social science researchers analyzing European Social Survey datasets across three environments: baseline tools, general-purpose LLMs, and LAPS.
- Measures included SUS, NASA-TLX, TXAI, and semi-structured interviews.
- Limitations and future work:
- Processing time for complex survey codebooks remains high.
- Python–R hybrid architecture increases execution complexity.
- Current support excludes advanced methods like SEM or multilevel modeling. Future work will address these limitations through OCR integration, fine-tuned LLMs, and expanded analytical capabilities.
Summary
LAPS is an automated framework that supports hypothesis-driven statistical analysis of public survey data using large language models. It preserves researcher agency, reduces cognitive workload, and produces trustworthy results through a structured pipeline of operationalization, planning, execution, and reporting. Empirical evaluation with social science researchers demonstrated significant improvements in usability, trust, and analytical stability compared to traditional tools and general-purpose LLMs. Future enhancements will focus on reducing processing time, simplifying hybrid architectures, and expanding support for advanced statistical methods. LAPS offers practical guidance for integrating LLMs into empirical research workflows and advancing human-AI collaboration in social science analysis.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 75%
InterFlow: Designing Unobtrusive AI to Empower Interviewers in Semi-Structured Interviews
CHI '26· Human-LLM Collaboration +3
- 67%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 67%
Evalet: Evaluating Large Language Models through Functional Fragmentation
CHI '26· Human-LLM Collaboration +3
- 67%
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
CHI '26· Human-LLM Collaboration +3
- 67%
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
CHI '26· Human-LLM Collaboration +3
- 67%
From Toil to Thought: Designing for Strategic Exploration and Responsible AI in Systematic Literature Reviews
IUI '26· Explainable AI (XAI) +3
- 67%
Integrating Complementary Feature Sets for Human-AI Decision-Making
IUI '26· Human-LLM Collaboration +3
- 67%
Criticality: Scaffolding Decision-Making with Interactive Critical Thinking and Evidence-Based Reasoning Traces
IUI '26· Human-LLM Collaboration +3
- 63%
LLM-based In-situ Thought Exchanges for Critical Paper Reading
IUI '26· Human-LLM Collaboration +2
- 63%
DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-Exploration
UIST '24· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)