How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz Study
Authors
Document Title
How Do Data Analysts Respond to AI Assistants? — A Guided Experimental Study
Document Information
- Subject Area: Human-Computer Interaction (HCI), Data Science, Artificial Intelligence Assistants
- Keywords: Data analysis, statistical analysis, AI assistants, code assistants, human-AI interaction, data science tools, planning assistance, guided experiments, LLM-assisted computation
Research Background and Problem
-
Problem or Challenge:
- During the data analysis process, analysts need to make various complex decisions, such as data cleaning and statistical modeling. These decisions are highly flexible and may lead to conflicting or divergent conclusions among analysts.
- Current tools based on large language models (LLMs, such as GitHub Copilot) primarily focus on code execution assistance but lack support for analysis "planning."
- The nonlinear and iterative nature of data analysis makes it difficult for existing AI assistants to accurately capture the analyst's current context and needs, such as proposing potential alternative decisions during the planning stage.
- Execution preferences in data analysis may lead analysts to overlook the diversity and depth required for high-quality analysis.
-
Importance:
- Data analysis serves as the foundation for scientific research and decision-making. However, deficiencies in analysis planning may result in unreliable outcomes and exacerbate the reproducibility crisis in scientific research.
- Providing high-quality data analysis suggestions can not only enhance the robustness of analysis results but also help analysts identify potential biases.
-
Research Motivation and Related Work:
- Although the advancements of LLMs have been applied in code generation and auto-completion functionalities, their role in supporting analysts in planning contexts remains underexplored.
- Research in areas such as "explainable AI" and "human-AI collaboration" reveals new opportunities for designing analysis assistants, but specific design directions in the context of data analysis are still lacking.
Solution
-
Method or Solution:
- Designed an AI data analysis assistant framework that integrates "planning assistance + execution assistance."
- Reviewed literature and analyzed a crowdsourced analysis study to categorize various types of actionable suggestions for the analysis process (e.g., domain background information, operational constructs, model selection, summarized in Table 1).
- Conducted a guided experiment (Wizard-of-Oz Study) in which 13 data analysts interacted with an AI assistant simulated by human experimenters to complete specified tasks, observing their reactions and preferences regarding planning and execution assistance.
-
Innovations:
- Proposed a categorization framework for planning suggestion types spanning the entire data analysis lifecycle.
- Simulated an LLM-driven "suggestion generation–timely delivery" intelligent interaction process in the experiment, focusing on user acceptance behavior during the planning stage.
- Explored the limitations of current LLM assistants and provided design recommendations for data analysis assistants that go beyond traditional execution assistance perspectives.
-
Implementation Steps and Techniques:
- Developed a JupyterLab extension as the user interface for interacting with the assistant.
- In the guided experiment, participants received real-time suggestions generated by experiment controllers (Wizard).
- Conducted qualitative and quantitative analysis on user preferences regarding suggestion types and timing.
Research Outcomes
-
Specific Outcomes:
- Planning facilitation: The study identified the usefulness of planning-related suggestions, such as high-level planning and high-dimensional variable selection, which helped analysts consider unnoticed decisions.
- The experiment demonstrated that planning and execution assistance have respective strengths and weaknesses in terms of timing and task alignment, with planning assistance significantly improving decision quality and depth.
- User acceptance of suggestions was influenced by multiple factors, including "content category," "expert background comparison," and "timing accuracy" (Figures 5 and 6 provide detailed analyses of these dynamics).
-
Comparison with Existing Solutions and Advantages:
- Compared to tools that only offer code completion (e.g., ChatGPT, Copilot), the assistant designed in this study integrates both execution and planning assistance while considering dynamic contextual adaptation.
- Emphasized the "planning advantage" in data analysis tasks rather than direct task completion, significantly enhancing the breadth and depth of analytical decision-making.
-
Experimental Results:
- On average, each analyst received 11.85 suggestions, with planning-related content accounting for 9.85 suggestions. The acceptance rate for planning suggestions that could be embedded into analysis tasks (e.g., variable operationalization, model selection) was 51.6%.
- Case studies demonstrated that better timing coordination and contextual control could greatly enhance user experience.
-
Limitations and Future Directions:
- Limitations:
- The study primarily recruited analysts with high statistical skills, which may limit the generalizability of results to beginners.
- The experiment duration was relatively short (approximately 2 hours), making it difficult to capture changes in long-term interaction habits.
- The study did not directly evaluate the specific improvement in analysis quality or accuracy due to AI suggestions.
- Future Directions:
- Develop practical systems for long-term use to study behavioral evolution.
- Assess the assistant's overall impact on analysis correctness, diversity, and interpretability.
- Expand the study sample to include analysts with broader skill backgrounds.
- Limitations:
Through the experiments and design presented in this study, the research team provides important insights for the development of future AI planning assistance tools, demonstrating that such assistants can significantly enhance the robustness and reliability of data analysis workflows.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do data analysts respond to planning assistance suggestions from AI assistants?Category: AI Understanding, Task Delegation, and Algorithm GovernanceSimilar questionsarrow_forward
- Can AI assistants improve the quality and depth of decisions during the planning phase of data analysis?Category: AI Understanding, Task Delegation, and Algorithm GovernanceSimilar questionsarrow_forward
- What factors influence data analysts' acceptance of planning suggestions from AI assistants?Category: AI Understanding, Task Delegation, and Algorithm GovernanceSimilar questionsarrow_forward
Practical Problems
1- Data analysts struggle to balance planning depth and decision diversity in analytical tasks.Category: AI Understanding, Task Delegation, and Algorithm GovernanceSimilar questionsarrow_forward
- 80%
Towards Feature Engineering with Human and AI’s Knowledge: Understanding Data Science Practitioners' Perceptions in Human&AI-Assisted Feature Engineering Design
DIS '24· Human-LLM Collaboration +2
- 80%
EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences
UIST '25· Human-LLM Collaboration +1
- 67%
Dango: A Mixed-Initiative Data Wrangling System using Large Language Model
CHI '25· Human-LLM Collaboration +2
- 67%
ChainBuddy: An AI-assisted Agent System for Generating LLM Pipelines
CHI '25· Human-LLM Collaboration +1
- 67%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
- 67%
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
CHI '26· Human-LLM Collaboration +2
- 67%
DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows
CHI '26· Human-LLM Collaboration +2
- 67%
Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging System
CHI '26· Interactive Data Visualization +2
- 67%
The Bots of Persuasion: Examining How Conversational Agents' Linguistic Expressions of Personality Affect User Perceptions and Decisions
CHI '26· Agent Personality & Anthropomorphism +2
- 67%
Belief Updating and Delegation in Multi-Task Human–AI Interaction: Evidence from Controlled Simulations
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)