AutoDS: Towards Human-Centered Automation of Data Science
Authors
Document Title
AutoDS: Towards Human-Centered Automation of Data Science
Document Information
- Topic Area: Automated Data Science, Human-Computer Collaboration, User Experience Design
- Keywords: Data Science, Automated Data Science, Automated Machine Learning, AutoML, AutoDS, Human-Computer Collaboration, Explainable AI, User Research, Automated Model Building, XAI
Research Background and Issues
-
Issues and Challenges:
- Tasks within the data science lifecycle, such as data exploration and model training, are labor-intensive and time-consuming, preventing data scientists from focusing on high-value knowledge discovery and decision-making activities.
- It remains unclear whether automation technologies (e.g., AutoML) can optimize these processes to support data science tasks and what their actual impact is on work practices and user behavior.
- Trust and user experience with automation tools (e.g., AutoDS) require deeper understanding.
-
Importance:
- Automation can reduce the time spent on low-level tasks, improve efficiency, and help data scientists focus on decision-making and knowledge discovery.
- Human-computer collaboration is a critical trend in future data science practices, necessitating exploration of how automation technologies can integrate into human workflows while ensuring interpretability and trustworthiness.
-
Research Motivation and Related Work:
- Previous studies indicate that 80% of data science time is spent on data preparation and model selection, tasks that can be optimized through automation.
- AutoML technologies are rapidly evolving (e.g., Google AutoML, H2O, DataRobot), yet research on their interaction with real-world workflows remains insufficient.
- HCI research has explored future directions for human collaboration with AutoDS systems, but experimental validation and user behavior data are still lacking.
Solution
-
Methods and Solutions:
- Propose a prototype automated data science system, AutoDS, capable of automatically suggesting machine learning configurations, preprocessing data, selecting algorithms, and training models.
- Provide two user interfaces: a web-based graphical interface and a programming notebook interface.
- Design experiments to compare data scientists' behaviors and outcomes when using AutoDS versus traditional Jupyter Notebook for task completion.
-
Innovations:
- Automatically generate readable Python code to help users understand and modify models.
- Provide real-time visualization of tree structures and model rankings to support users in monitoring model training and filtering results.
- Explore shifts in user work patterns and their trust and acceptance of tools assisted by automation.
-
Implementation Steps and Technologies:
- The system accepts user-uploaded datasets, automatically suggests task configurations, and generates multiple model pipelines.
- High-performance models are created using a joint optimization algorithm combining data preprocessing, feature engineering, algorithm selection, and hyperparameter tuning.
- Users can browse model details, download and edit corresponding Python notebook code, or directly deploy models as API endpoints.
Research Outcomes
-
Specific Outcomes:
- In experiments, AutoDS significantly improved productivity (average of 8 models generated per user) and model quality (ROC AUC 0.919 compared to 0.899 using traditional methods), while reducing human errors.
- Despite high model quality, user confidence in AutoDS-generated models was lower than manually crafted models (2.4 vs. 3.3 on a 5-point scale).
-
Advantages Over Existing Solutions:
- Automation supports the entire data science lifecycle, saving substantial time.
- Provides more understandable and modifiable code output, enhancing system transparency.
- Encourages users to focus on understanding models and data rather than repetitive coding tasks.
-
Experimental or Evaluation Results:
- Experiments validated the efficiency gains and quality improvements brought by AutoDS.
- Users showed acceptance of AutoDS functionalities, but confidence in the models requires further design optimization.
-
Limitations and Future Directions:
- The simplicity of the experimental task datasets may limit direct applicability to complex data science projects.
- Trust issues with AutoDS need to be addressed, such as by enhancing its interpretability and transparency.
- Future research could expand to diverse user groups (e.g., domain experts and end-users) and applications in complex, multi-stage data science projects.
- Investigate further optimization of human-computer collaboration within data science workflows.
The above is a structured summary and key point extraction of the PDF content. This article proposes an automation platform tailored for human-centered data science work and analyzes its impact on user behavior, productivity, and model quality through experimental research, while identifying areas for improvement.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do automation tools (e.g., AutoDS) affect data scientists' work behavior and efficiency?Category: Data Tool Adoption, Analysis Interfaces, and Information Organization SupportSimilar questionsarrow_forward
- How can automated data science platforms improve model quality and reduce human error?Category: Data Tool Adoption, Analysis Interfaces, and Information Organization SupportSimilar questionsarrow_forward
- How do users perceive automatically generated models, and how do trust and acceptance change?Category: Data Tool Adoption, Analysis Interfaces, and Information Organization SupportSimilar questionsarrow_forward
Practical Problems
1- Inefficient repetitive work in data science consumes time and affects high-value tasks.Category: Data Tool Adoption, Analysis Interfaces, and Information Organization SupportSimilar questionsarrow_forward
- 71%
EvalAssist: Insights on Task-Specific Evaluations and AI-Assisted Judgment Strategy Preferences
UIST '25· Human-LLM Collaboration +1
- 63%
More Isn't Always Better: Balancing Decision Accuracy and Conformity Pressures in Multi-AI Advice
CHI '26· Human-LLM Collaboration +2
- 63%
Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging System
CHI '26· Interactive Data Visualization +2
- 63%
The Bots of Persuasion: Examining How Conversational Agents' Linguistic Expressions of Personality Affect User Perceptions and Decisions
CHI '26· Agent Personality & Anthropomorphism +2
- 63%
Belief Updating and Delegation in Multi-Task Human–AI Interaction: Evidence from Controlled Simulations
CHI '26· Human-LLM Collaboration +2
- 63%
Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
IUI '26· Human-LLM Collaboration +2
- 63%
Making Absence Visible in Intelligent Summarization Interfaces
IUI '26· Human-LLM Collaboration +2
- 63%
"Un-default" Behavior Tuning: Specifying Model Behavior outside the Norm with LLM Self-Playing and Self-Improving
IUI '26· Human-LLM Collaboration +2
- 63%
Never-ending Learning of User Interfaces
UIST '23· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)