How Do Analysts Understand and Verify AI-Assisted Data Analyses?
Authors
Human-LLM CollaborationExplainable AI (XAI)Interactive Data VisualizationUI/UX DesignersData Scientists & AnalystsAI/ML Researchers & Engineers
Title of the Paper
How to Understand and Validate AI-Assisted Data Analysis?
Bibliographic Information
- Domain: Human-Computer Interaction (HCI), Natural Language Interfaces, Data Science Tools and Validation
- Keywords: AI-assisted data analysis, Explainable AI, Human-Computer Interaction, Data Science Assistants, Automated Data Science, Design Probes
Research Background and Problem
- What problems or challenges did the authors identify?
Data analysis, which combines domain knowledge, statistical skills, and programming expertise, is often challenging. AI-assisted tools (e.g., ChatGPT) can help analysts by converting natural language instructions into code to facilitate data processing. However, AI-generated analyses may misinterpret user intent or produce errors, and validating the accuracy of these analyses remains difficult. - Why is this issue important?
Conclusions drawn from data analysis often influence critical decisions, such as those in scientific research, business operations, and government policies. If AI-generated analyses contain errors and users blindly rely on automated outputs, it could lead to incorrect conclusions and high-cost decision-making risks. - Research motivation and related work
From an HCI perspective, the study explores how to design tools that enable analysts to effectively understand and validate AI-generated analyses. Existing work focuses on programming support tools (e.g., Copilot) but rarely addresses the specific needs of data analysis users. This research aims to fill the gap in studying the validation process for data analysis.
Solution
- What methods or solutions did the authors propose?
The authors developed a design probe that integrates AI-generated code, code annotations, natural language explanations, data visualizations, and interactive data tables to support analysts in performing validation tasks. - What is innovative about this solution?
Compared to existing tools (e.g., Code Interpreter), the design probe introduces interactive support for intermediate data tables and combines them with visual summaries. This helps analysts trace the data processing workflow and better validate analysis results. - What are the implementation steps and key technologies used?
- Design Probe: A multi-panel interface that combines AI-generated code and natural language explanations with interactive data tables. The side panel supports data browsing, filtering, and sorting functionalities.
- Task Preparation: Selection of tasks involving real-world datasets and problems, ensuring a diversity of errors (e.g., data errors, computational errors, or misinterpretation of the analysis problem).
- User Study: Conducted with 22 participants who performed validation processes on 10 data analysis tasks, observing their behaviors and recording the tools and methods they used.
Research Findings
- What specific findings were obtained?
- Analysts alternated between "process-based behaviors" and "data-based behaviors" during the validation process. For instance, they typically began by validating the AI's process and then delved into the data to confirm findings.
- Data tables and visualizations played a crucial role in error detection, while natural language explanations and code were used to understand the data processing logic.
- What advantages does it have compared to existing solutions?
- The provided data summaries and intermediate data tables enabled participants to identify errors more effectively.
- The tight coupling of data and process improved analysts' validation efficiency.
- What were the experimental or evaluation results?
- The majority of participants used process-based behaviors to validate AI analysis steps, while about half transitioned to data-oriented behaviors to understand data structures or identify errors.
- Validation behaviors were significantly influenced by participants' programming backgrounds and data analysis experience.
- Limitations and future directions
- The study did not cover validation behaviors involving more complex data analyses (e.g., machine learning models or statistical inference).
- It did not explore how analysts use their own data and design AI interactions in real-world scenarios.
- Future research is proposed to optimize AI's explicit articulation of data assumptions to assist users in constructing validation strategies.
Impact and Discussion
- Impact on data analysts
Improving data literacy and learning common validation strategies (e.g., quick data verification) are critical skills. Understanding the strengths and limitations of AI tools can enhance analysis quality and reliability. - Recommendations for system designers
- Integrate interactive features for data and code to enhance data traceability.
- Provide multiple forms of explanation (e.g., code translation, natural language conversion, or visual representation) for data operations.
- Explicitly display AI's assumptions about the data and the analysis context.
This study enriches the understanding of AI-assisted analysis validation and provides important insights for designing smarter data validation tools.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- In AI-assisted data analysis, how can analysts understand and verify AI-generated analytical results?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- During data validation, which key features help analysts better identify and correct errors?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- What behavioral patterns do analysts exhibit when validating AI-generated analytical results?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
lightbulb
Practical Problems
1- AI-generated data analysis results may contain errors that are difficult to verify, undermining user trust.Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- 86%
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
CHI '25· Human-LLM Collaboration +2
- 83%
iScore: Visual Analytics for Interpreting How Language Models Automatically Score Summaries
IUI '24· Explainable AI (XAI) +1
- 75%
Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
IUI '26· Human-LLM Collaboration +3
- 71%
Talk to the Hand: an LLM-powered Chatbot with Visual Pointer as Proactive Companion for On-Screen Tasks
CHI '25· Voice User Interface (VUI) Design +2
- 71%
Interactive Explainable Ranking
CHI '26· Explainable AI (XAI) +2
- 71%
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
CHI '26· Human-LLM Collaboration +2
- 71%
OntoScope: Using a Divergent-Convergent Interaction Framework to Support LLM-based Ontology Scoping
IUI '26· Human-LLM Collaboration +2
- 71%
SCSimulator: An Exploratory Visual Analytics Framework for Partner Selection in Supply Chains through LLM-driven Multi-Agent Simulation
IUI '26· Human-LLM Collaboration +2
- 71%
Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
UIST '24· Human-LLM Collaboration +2
- 67%
FDHelper: Assist Unsupervised Fraud Detection Experts with Interactive Feature Selection and Evaluation
CHI '20· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642497
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), Interactive Data Visualization
work
Professions
UI/UX Designers, Data Scientists & Analysts, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers