LAPS: Automating Hypothesis-Driven Statistical Analysis of Public Survey Using Large Language Models

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationExplainable AI (XAI)User Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingUniversity Professors & ResearchersHCI ResearchersData Scientists & Analysts

Paper Title

LAPS: Automating Hypothesis-Driven Statistical Analysis of Public Survey Using Large Language Models

Publication Info

  • Topic area: Automated frameworks for hypothesis-driven statistical analysis in social science research.
  • Keywords: Public survey analysis, large language models, hypothesis testing, statistical planning, human-AI collaboration, cognitive workload, transparency, reproducibility, social science, research workflows.

Background and Problem

  • Problem / challenge: Analyzing public survey data is complex due to inconsistent structures, large variable sets, and varying researcher expertise. Existing tools lack structured support for hypothesis-driven workflows, leading to instability and reduced reliability in results.
  • Significance: Public surveys inform critical decisions in policy and academia, but their complexity limits accessibility and reproducibility, especially for novice researchers.
  • Motivation and related work: Traditional tools (e.g., SPSS, R) and recent AI-based systems provide partial support but fail to address challenges specific to public survey data, such as complex survey designs and operationalization of hypotheses. General-purpose LLMs lack alignment with domain-specific workflows, producing outputs that may drift from researchers' intent.

Solution

  • Proposed approach: LAPS (LLM-assisted Automated framework for Public Survey data analysis), an end-to-end system for hypothesis-driven statistical analysis of survey data.
  • Novelty:
    1. Introduces a structured pipeline with four modules: Operationalization, Planning, Execution, and Reporting.
    2. Incorporates human-in-the-loop mechanisms to balance automation with researcher agency.
    3. Provides self-refinement loops for iterative improvement of variable selection and analysis plans.
    4. Aligns system behavior with domain-specific workflows to ensure transparency and reproducibility.
  • Procedure and key techniques:
    • Operationalization: Parses survey codebooks, validates hypotheses, selects variables, and refines them iteratively.
    • Planning: Links descriptive statistics to structured analysis plans, incorporating complex survey design elements.
    • Execution: Generates and debugs hybrid Python–R code for survey-weighted analyses.
    • Reporting: Produces APA-style statistical reports with contextual interpretations and audit trails.

Results

  • Concrete findings:
    • SUS scores for LAPS were significantly higher than baseline tools (mean difference = 42.30, p < 0.001) and general-purpose LLMs (mean difference = 12.90, p < 0.05).
    • NASA-TLX scores showed reduced cognitive workload compared to baseline tools (mean difference = 1.94, p < 0.001) and general-purpose LLMs (mean difference = 0.86, p < 0.01).
    • TXAI scores indicated higher trust in LAPS explanations compared to general-purpose LLMs (overall mean difference = -0.87, p < 0.05).
  • Advantage over baselines:
    • Reduced cognitive burden through unified workflows.
    • Higher usability and trust ratings compared to traditional tools and general-purpose LLMs.
    • Enhanced researcher agency and analytical stability.
  • Experiments / evaluation:
    • User study with 12 social science researchers analyzing European Social Survey datasets across three environments: baseline tools, general-purpose LLMs, and LAPS.
    • Measures included SUS, NASA-TLX, TXAI, and semi-structured interviews.
  • Limitations and future work:
    • Processing time for complex survey codebooks remains high.
    • Python–R hybrid architecture increases execution complexity.
    • Current support excludes advanced methods like SEM or multilevel modeling. Future work will address these limitations through OCR integration, fine-tuned LLMs, and expanded analytical capabilities.

Summary

LAPS is an automated framework that supports hypothesis-driven statistical analysis of public survey data using large language models. It preserves researcher agency, reduces cognitive workload, and produces trustworthy results through a structured pipeline of operationalization, planning, execution, and reporting. Empirical evaluation with social science researchers demonstrated significant improvements in usability, trust, and analytical stability compared to traditional tools and general-purpose LLMs. Future enhancements will focus on reducing processing time, simplifying hybrid architectures, and expanding support for advanced statistical methods. LAPS offers practical guidance for integrating LLMs into empirical research workflows and advancing human-AI collaboration in social science analysis.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222863/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791665
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Explainable AI (XAI), User Research Methods (Interviews, Surveys, Observation)
work
Professions
University Professors & Researchers, HCI Researchers, Data Scientists & Analysts
article
Content Status
Full text indexed
hub
Related Papers
10 related papers