Facilitating Knowledge Sharing from Domain Experts to Data Scientists for Building NLP Models
Authors
Document Title
Facilitating Knowledge Sharing from Domain Experts to Data Scientists for Building NLP Models
Document Information
- Subject Areas: Human-Computer Interaction, Machine Learning, Multidisciplinary Collaboration
- Keywords: Human-in-the-Loop Machine Learning, CSCW (Computer-Supported Cooperative Work), Natural Language Processing, Domain Knowledge Sharing, Multidisciplinary Collaboration, Label Inference, Knowledge Extraction, Interactive Systems, Cold Start, Scalability
- Conference: 26th International Conference on Intelligent User Interfaces (IUI '21)
- Primary Authors and Affiliations:
- Soya Park (MIT)
- April Wang (University of Michigan)
- Ban Kawas, Q. Vera Liao, David Piorkowski, Marina Danilevsky (IBM Research)
Research Background and Problem
-
Challenges and Issues:
- Data scientists need to acquire domain-specific knowledge to develop machine learning models, but this learning process is often time-consuming and inefficient.
- Extracting domain experts' knowledge relies on sporadic and complex interactions with data scientists, and there is a lack of standardized tools to facilitate collaboration.
- Current machine learning development often requires a large amount of labeled data, making it particularly challenging to develop and debug models in scenarios with scarce or non-existent labels.
-
Motivation and Importance:
- As machine learning becomes widely applied in domains such as healthcare and finance, it is crucial to address the inefficiency of domain knowledge transfer and the scarcity of labels to accelerate model development.
- More efficient methods for domain knowledge extraction can enhance model development transparency, scalability, and reduce the burden on domain experts.
-
Related Work:
- Technical and interpersonal challenges in data science workflows, such as the limitations of collaborative data science tools and the participation patterns of domain experts in model development.
- Existing domain knowledge acquisition methods focus on document annotation tools or intuitive interactive tools for data collection, but they still face issues such as difficulty in knowledge transfer and insufficient support for data scientists' needs.
Solution
-
Core Contribution:
- Proposed a domain knowledge transfer framework "Ziva" (for natural language processing tasks).
- Designed and developed a user interface to help domain experts effectively transfer knowledge through domain concept extraction and label inference methods.
-
Innovations:
- Introduced a diverse set of inference methods for domain experts to share knowledge via representative data instances (supporting five forms: keyword selection, instance simplification, instance perturbation, concept keyword annotation, and concept annotation).
- Provided asynchronous interaction and reuse capabilities for knowledge.
- Supported model development needs in cold start scenarios.
-
Implementation Steps and Techniques:
- Extracted a representative subset of text instances from unlabeled data (using tf-idf and k-means clustering).
- Designed an interactive interface for domain experts to perform the following tasks:
- Concept Extraction: Group instance cards through a drag-and-drop interface to extract domain topics and descriptions.
- Label Inference: Use five specific methods for label explanation to help data scientists understand domain instances.
- Data Usage: Data scientists utilized domain knowledge extracted by Ziva as the basis for feature engineering, rapid label generation, and model debugging.
Research Outcomes
-
Experimental Design:
- Experiment 1: Small-scale lab user study (12 participants) comparing task completion time, cognitive load, and task experience under different label inference methods for domain experts.
- Experiment 2: Extended to a large-scale crowdsourcing experiment (88 participants), reporting the usability, mental stress, and subjective work experience of each method based on field surveys.
- Experiment 3: Interview study with data scientists (7 experienced professionals) to examine the practical value of Ziva outputs in machine learning workflows.
-
Key Findings:
- Domain Knowledge Transfer:
- Ziva supported domain experts in completing knowledge extraction tasks with relatively low cognitive load, performing particularly well under the "instance simplification" and "concept annotation" task conditions.
- Utility for Data Scientists:
- Data scientists noted that the outputs could be used for label propagation, feature engineering, enhancing interpretability, and designing few-shot learning models.
- "Concept annotation" was the most favored, as its structured information could be directly applied to multi-task modeling.
- Experimental Performance:
- Test results based on domain expert annotations showed that features generated from concept annotations exhibited stronger consistency in machine learning models compared to existing "keyword (Bag of Words)" methods.
- Feedback:
- Data scientists acknowledged Ziva's contribution to seamless domain knowledge collaboration and suggested future extensions to different domains and multi-model architectures.
- Domain Knowledge Transfer:
-
Limitations and Future Directions:
- Currently validated only on sentiment analysis tasks and restaurant review data. Future work should extend to other task contexts and domains (e.g., healthcare, education, machine translation).
- Further research is needed on sparse domain knowledge, such as domain background information and data sparsity.
- Explore key design elements for tools that support efficient collaboration and knowledge reuse.
Conclusion
This paper provides a standardized tool (Ziva) for facilitating collaboration between data scientists and domain experts in building NLP models. Through design evaluations, it demonstrates how novel interaction methods and an open framework can significantly improve the efficiency of knowledge sharing while addressing common challenges in multidisciplinary collaboration.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can interactive systems be designed to efficiently support domain experts in transferring domain knowledge to data scientists?Category: Scientific Anomaly Detection and Causal Analysis SupportSimilar questionsarrow_forward
- How do five label inference methods affect task performance and cognitive burden in domain expert knowledge transfer?Category: Scientific Anomaly Detection and Causal Analysis SupportSimilar questionsarrow_forward
- Can information extracted from domain concept annotations be directly used for feature engineering and debugging machine learning models?Category: Scientific Anomaly Detection and Causal Analysis SupportSimilar questionsarrow_forward
Practical Problems
1- Data scientists struggle to rapidly acquire domain knowledge, lowering machine learning development efficiency.Category: Scientific Anomaly Detection and Causal Analysis SupportSimilar questionsarrow_forward
- 100%
Understanding and Supporting Knowledge Decomposition for Machine Teaching
DIS '20· Human-LLM Collaboration +1
- 100%
Mallard: Turn the Web into a Contextualized Prototyping Environment for Machine Learning
UIST '19· Human-LLM Collaboration +1
- 80%
Rapsai: Accelerating Machine Learning Prototyping of Multimedia Applications through Visual Programming
CHI '23· Human-LLM Collaboration +1
- 80%
Exploring Empty Spaces: Human-in-the-Loop Data Augmentation
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 80%
CoLadder: Manipulating Code Generation via Multi-Level Blocks
UIST '24· Human-LLM Collaboration +1
- 75%
Enabling Data-Driven API Design with Community Usage Data: A Need-Finding Study
CHI '20· Computational Methods in HCI
- 67%
Crystalline: Lowering the Cost for Developers to Collect and Organize Information for Decision Making
CHI '22· Human-LLM Collaboration +2
- 67%
Beyond Code Generation: LLM-supported Exploration of the Program Design Space
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 67%
Linting Style and Substance in READMEs
CHI '26· Computational Methods in HCI +2
- 67%
Cerebra: Aligning Implicit Knowledge in Interactive SQL Authoring
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)