Tracing and Visualizing Human-ML/AI Collaborative Processes through Artifacts of Data Work
Honorable MentionTitle of the Paper
Tracing and Visualizing Human-ML/AI Collaborative Processes through Artifacts of Data Work
Paper Information
- Domain: Human-computer collaborative work, particularly the traceability and visualization of human collaboration with automated machine learning (AutoML) technologies in data science workflows.
- Keywords: AutoML, human-machine collaboration, data visualization, artifact classification, end-to-end data science.
Research Background and Issues
-
Identified Problems or Challenges:
- Current AutoML systems primarily focus on automating model selection and data preparation, but substantial manual coordination and human-machine collaboration are still required in practice.
- The "black-box" nature of automation often leads to collaboration friction within teams, making it difficult to determine "who did what and when."
- Existing AutoML tools rarely address the human element, with limited functionality to trace and audit human involvement in data science workflows.
-
Significance of the Research: Capturing and tracing the collaboration between human and machine intelligence in data work is becoming increasingly important. Enhancing transparency and interpretability can help both technical and non-technical users better understand data flows and decision-making processes.
-
Motivation and Related Work:
- Inspired by popular visual analytics methods and AutoML literature, the authors aim to create a tool to bridge the gap between automation and manual operations.
- Drawing on prior research on human-machine collaboration, traceability, and visual-assisted analysis, the authors designed a general artifact classification to describe various artifacts (human inputs or machine-generated) in data operations.
Solution
-
Methods and Solutions:
- Development of an AutoML Artifact Classification: Creation of a classification system to summarize and describe artifacts in AutoML and human collaboration processes. Artifacts include not only inputs and outputs but also metadata and documentation from automated processes.
- Development of the AutoML Trace Interactive Visualization Prototype: An interactive visualization tool that displays the artifacts generated during data workflows and their dynamic evolution in human-machine collaboration.
- Definition of Data Work Traceability: Proposed a definition encompassing three core aspects: artifact provenance, transparency, and context.
-
Innovations:
- Extended the artifact classification to a broader context of end-to-end data workflows, covering stages such as preparation, analysis, deployment, and communication.
- Provided a classification system adaptable to various systems, supporting the capture, tracing, and visualization of artifacts in automated data science processes.
-
Implementation Steps and Key Techniques:
- Artifact Classification Development:
- Through an iterative process combining literature review and theoretical construction, distilled 52 final artifacts from an initial set of 400.
- The classification includes four dimensions: provenance, transmission mode, format, and task.
- Development of AutoML Trace:
- Leveraged APIs in a real-world enterprise AutoML system to capture artifacts, record timestamps, and dependencies, enabling artifact tracing and version comparison.
- Provided three visualization views (provenance view, dependency view, and version history view) to explore artifact origins, dependencies, and evolution.
- Artifact Classification Development:
Research Outcomes
-
Specific Outcomes:
- Proposed an AutoML Artifact Classification suitable for capturing human-machine collaboration artifacts, encompassing 52 grouped artifacts across four major stages: data preparation, analysis, deployment, and communication.
- Developed the AutoML Trace Visualization Prototype, which demonstrates human-machine collaboration in data work and feedback on human interventions from AutoML systems.
- Validated the tool's design effectiveness through collaboration scenarios with enterprise teams, helping them better understand their AutoML systems' functionality and areas for improvement.
-
Comparative Advantages over Existing Solutions:
- Considered human factors and socio-technical relationships in AutoML systems, addressing the lack of human element focus in existing tools.
- The classification and tool's generalizability allow adaptation to diverse AutoML systems, supporting a broader range of tasks.
-
Experiment or Evaluation Results:
- In collaboration with enterprise teams, the tool and classification provided new perspectives, enabling them to analyze human-machine interactions and identify improvement needs in their systems. For example, they could clearly see which human interventions were accepted or ignored by the system.
-
Limitations and Future Directions:
- Limitations:
- The classification development and tool validation were based on a limited number of enterprise team contexts, and generalizability needs further verification.
- The system currently lacks optimization for multi-user collaboration and complex cooperative scenarios.
- Future Directions:
- Expand validation to multi-team, cross-system contexts to further refine and enhance the classification.
- Develop more complex interactive designs to support asynchronous collaboration analysis between humans and machines, as well as among humans.
- Extend visualization research to different types of AutoML systems and emerging ML/AI technologies.
- Limitations:
By proposing a novel classification method and implementing it as a visualization tool, the authors provide an important theoretical framework and practical tool for studying AutoML and human collaboration, with potential for further development in both academia and industry.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can human-AI collaborative data artifacts in AutoML systems be captured and described through classification?Category: Machine Learning Model Visualization, Debugging, and Explainability SupportSimilar questionsarrow_forward
- How can visualization tools track and present human-AI collaborative workflows and their evolution?Category: Machine Learning Model Visualization, Debugging, and Explainability SupportSimilar questionsarrow_forward
- What definition can comprehensively reflect traceability of data work, including artifact provenance, transparency, and contextual relations?Category: Machine Learning Model Visualization, Debugging, and Explainability SupportSimilar questionsarrow_forward
Practical Problems
1- Team members struggle to understand who did what and when in AutoML systems.Category: Machine Learning Model Visualization, Debugging, and Explainability SupportSimilar questionsarrow_forward
- 100%
Visualizing Examples of Deep Neural Networks at Scale
CHI '21· Human-LLM Collaboration +1
- 80%
Method for Exploring Generative Adversarial Networks (GANs) via Automatically Generated Image Galleries
CHI '21· Generative AI (Text, Image, Music, Video) +2
- 67%
Crystalline: Lowering the Cost for Developers to Collect and Organize Information for Decision Making
CHI '22· Human-LLM Collaboration +2
- 67%
Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs
CHI '25· Human-LLM Collaboration +2
- 67%
Dango: A Mixed-Initiative Data Wrangling System using Large Language Model
CHI '25· Human-LLM Collaboration +2
- 67%
Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 67%
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
CHI '26· Human-LLM Collaboration +2
- 67%
CodeVoyager: Integrating Interactive Visual Aids with LLMs for Code Comprehension
IUI '26· Human-LLM Collaboration +2
- 67%
SCSimulator: An Exploratory Visual Analytics Framework for Partner Selection in Supply Chains through LLM-driven Multi-Agent Simulation
IUI '26· Human-LLM Collaboration +2
- 67%
Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition
UIST '24· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)