Fits and Starts: Enterprise Use of AutoML and the Role of Humans in the Loop

Honorable Mention
AutoML InterfacesInteractive Data VisualizationSoftware Engineers & DevelopersUI/UX DesignersData Scientists & AnalystsAI/ML Researchers & Engineers

Title of the Paper

Fits and Starts: Enterprise Use of AutoML and the Role of Humans in the Loop

Paper Information

  • Subject Area: Human-Computer Interaction (HCI), Automated Machine Learning (AutoML), Data Science
  • Keywords: Data Science, Automation, Machine Learning, Human-Computer Interaction, Data Visualization, Artificial Intelligence, AutoML, Data Analysis, Collaboration, Production Environment

Research Background and Issues

  • Problems and Challenges:

    1. Enterprises face numerous challenges when adopting Automated Machine Learning (AutoML) technologies, particularly in data preparation, model monitoring, deployment, and communication processes.
    2. Current AutoML systems have not achieved end-to-end automation of the data science workflow and still require significant human intervention.
    3. A key issue is how personnel with varying technical backgrounds within enterprises can effectively use AutoML technologies.
    4. The role of data visualization in effectively supporting human-machine collaboration within AutoML systems remains underexplored.
  • Significance: AutoML has the potential to accelerate the application of machine learning, lower technical barriers, and enable non-technical users to utilize complex data science tools, thereby driving data-driven decision-making in enterprises. However, improper use could lead to severe consequences. Understanding these issues has profound implications for designing effective AutoML and visualization tools.

  • Research Motivation: The authors aim to explore the real-world use of AutoML in enterprise environments, including how data visualization can integrate human-machine collaboration into the data science workflow and address current limitations.

Solutions

  • Methods and Solutions:

    1. Research Methodology: Conducted interviews with 29 participants (data scientists, business analysts, team managers) from enterprises of varying sizes to analyze the practical use and challenges of AutoML.
    2. Proposed a framework summarizing the levels of automation required by users with different technical proficiencies.
    3. Identified three primary use cases: automation of routine tasks, rapid prototyping, and democratization of data science.
  • Innovations:

    1. Introduced a framework addressing the automation needs of users with varying technical expertise.
    2. Highlighted the potential and limitations of data visualization in fostering human-machine collaboration.
    3. Identified critical design directions, such as improving tool integration and automating data preparation.
  • Implementation Steps and Techniques:

    1. Collected user feedback on AutoML usage to identify key issues.
    2. Built an analytical framework based on existing data science workflows, including data preparation, analysis, deployment, and communication.
    3. Developed a hierarchical automation model to guide the design of human-machine collaborative tools.

Research Findings

  • Specific Findings:

    1. Provided a framework outlining the ideal levels of automation for users with different technical proficiencies in the data science workflow.
    2. Identified the main use cases of AutoML:
      • A. Accelerating routine tasks
      • B. Rapid prototyping and exploration
      • C. Democratization (enabling non-technical users to participate in data work)
    3. Demonstrated that data preparation remains the primary bottleneck for enterprise adoption of AutoML.
    4. Emphasized the importance of a cautious approach to "human-machine collaboration," suggesting that human intervention should be limited in certain stages.
    5. Highlighted the shortcomings of data visualization in monitoring and communication, as well as potential areas for improvement.
  • Comparison with Existing Solutions:
    Compared to traditional AutoML systems that focus solely on model selection and hyperparameter tuning, this study shows that an integrated solution encompassing data preparation, deployment, and communication tools is essential to address current enterprise pain points.

  • Experimental or Evaluation Results:

    1. Data preparation consumes a significant amount of users' time; existing tools (e.g., Alteryx) partially address the issue but require further improvement.
    2. AutoML is being explored more extensively by large enterprises, but its adoption risks exacerbating the technical knowledge gap.
    3. Visualization tools are not well-integrated into existing enterprise data tool ecosystems.
  • Limitations and Future Directions:

    1. The study primarily involved data science technical experts; further research is needed to examine the needs of "citizen data scientists" (non-technical users).
    2. More no-code or low-code AutoML technologies should be developed, and tool ecosystems should be more tightly integrated.
    3. The sample size was limited; future studies should expand to include more industries and types of data work.
    4. Future research should explore new visualization tools to support the "human-machine collaboration" model in AutoML systems, enhancing collaboration and trust.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47318/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445775
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
Honorable Mention
group
Authors
2 authors
sell
Subtopics
AutoML Interfaces, Interactive Data Visualization
work
Professions
Software Engineers & Developers, UI/UX Designers, Data Scientists & Analysts, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
6 related papers