Towards Guidelines for Designing Human-in-the-Loop Machine Training Interfaces
Title of the Paper
Towards Guidelines for Designing Human-in-the-Loop Machine Training Interfaces
Paper Information
- Domain: Human-Computer Interaction, Interactive Machine Learning, Artificial Intelligence System Design
- Keywords: Human-Computer Interaction, Machine Learning, Data Labeling, Algorithm Training, User Interface, User Experience, UX Evaluation
Research Background and Problem Statement
-
What problems or challenges did the authors identify?
- The "human-in-the-loop" process in Interactive Machine Learning (IML) poses a bottleneck to the scalability of training processes. Relying solely on human labeling tasks may fail to meet the demands of large-scale datasets and is prone to issues such as fatigue and labeling quality.
- There is a lack of design guidelines specifically tailored for creating interactive interfaces for IML systems, particularly regarding how to balance user agency, interaction burden, and model efficiency.
- Many existing interfaces fail to adequately support user agency or provide intuitive options, leading to user frustration.
-
Why is this problem important? Designing effective user interfaces for interactive machine learning can enhance system usability, reduce the burden of human labeling, and optimize the performance of machine learning systems. Systems that balance user agency and learning efficiency are critical for practical industrial applications, sustainable user experiences, and reducing training costs.
-
Research Motivation and Related Work
- Traditional machine learning systems are difficult to debug and may produce biased, irrelevant, or offensive results.
- Existing studies (e.g., Fails and Olsen [7]) have explored how interactive machine learning allows users to directly participate in training but lack systematic interaction design guidelines.
- Issues such as transparency and contextual adaptability have been identified as common pain points in several studies (e.g., Kulesza et al. [13][15], Amershi et al. [1]), encouraging the development of better design standards.
Proposed Solution
-
What methods or solutions did the authors propose? This study designed an experiment using four different interface variants for "human-in-the-loop machine learning training," allowing users to train a recommendation system through these interfaces. The interfaces were evaluated based on interaction speed, user experience, and efficiency parameters.
-
What is innovative about this solution?
- The authors proposed the IMLIQ scoring method (Interactive Machine Learning Interaction Quality), a comprehensive evaluation metric combining user experience and system learning efficiency.
- Preliminary design guidelines for interactive learning interfaces were proposed to optimize the balance between user experience and efficiency.
-
What were the implementation steps and key technologies used?
- Experimental Design:
- Participants interacted with four different recommendation system interfaces to complete a task of recommending a specific target item (a red chair).
- User interaction data (e.g., number of interactions and time to complete tasks) and subjective questionnaire evaluations (NASA-TLX for user burden) were collected.
- Interface Variants:
- Interface 1: A binary labeling task where users selected "like" or "dislike" for specific items.
- Interface 2: Similar to Interface 1 but allowed users to temporarily disable the influence of specific features (e.g., color or style).
- Interface 3: Comparative selection, where users chose the closest match to the target item from three options.
- Interface 4: After making a selection, the system asked users to specify the reason for their choice.
- Evaluation Metrics:
- Efficiency: Recorded the number of interactions and time required to complete tasks.
- User Experience: Collected subjective scores on psychological, physical, and temporal workload using NASA-TLX, along with additional assessments of interactivity and enjoyment.
- Experimental Design:
Research Findings
-
What specific results were achieved?
- Users preferred interfaces with greater flexibility and expressive capabilities (Interfaces 2 and 3 received the most votes).
- Users tended to favor interfaces that avoided "irreversible" or repetitive tasks, perceiving them as more efficient and enjoyable.
- By comparing the interaction performance of the four interfaces, preliminary design guidelines for interactive learning interfaces were derived.
-
What advantages does this solution have over existing ones? The proposed design guidelines and IMLIQ scoring model intuitively integrate user experience and system performance, offering a new perspective for the integration of human-computer interaction and machine learning interfaces.
-
What were the experimental or evaluation results?
- User Preference Distribution: Interfaces 2 and 3 received the most user votes (6 votes each); Interface 1 received 2 votes, and Interface 4 received only 1 vote.
- Performance Comparison: In the IMLIQ model scoring, Interface 3 achieved the highest score (8.2), while Interface 4 scored the lowest (-1.5).
| Interface | Avg. Interaction Count | Effort Score (1-10) | Interactivity Score (1-10) | Enjoyment Score (1-10) |
|---|---|---|---|---|
| Interface 1 | 6.0 | 2.4 | 4.5 | 4.9 |
| Interface 2 | 5.0 | 2.3 | 6.6 | 6.1 |
| Interface 3 | 5.0 | 2.6 | 6.0 | 6.6 |
| Interface 4 | 6.0 | 3.2 | 5.7 | 4.7 |
- Limitations and Future Directions:
- The current experiment had a limited sample size, with participants primarily from higher education backgrounds, potentially biasing results toward users familiar with ML systems.
- The test scenarios were relatively simple and may not fully represent real-world industrial applications.
- The proposed IMLIQ scoring method requires further large-scale experimental validation, particularly in establishing empirical weights for subjective scores.
Conclusion
The design of interactive machine learning system interfaces requires a balance between user experience and system efficiency. This study highlighted key issues in designing human-in-the-loop training systems and explored the applicability of different interfaces through experimentation. It proposed four design guidelines and the IMLIQ scoring method as a reference framework for future research and design. This work provides new insights into the intersection of HCI and machine learning, with future validation needed in more complex scenarios and with broader user groups to confirm these preliminary findings.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can interfaces in interactive machine learning be designed to balance UX and system efficiency?Category: Programming, Computing, and Physical Prototyping EducationSimilar questionsarrow_forward
- How do different types of human-in-the-loop interfaces perform in efficiency and UX during machine learning training?Category: Programming, Computing, and Physical Prototyping EducationSimilar questionsarrow_forward
- What design guidelines do interactive machine learning systems need to optimize user agency and reduce operational burden?Category: Programming, Computing, and Physical Prototyping EducationSimilar questionsarrow_forward
Practical Problems
1- Users face redundant, burdensome, and inefficient tasks when training machine learning systems.Category: Programming, Computing, and Physical Prototyping EducationSimilar questionsarrow_forward
- 88%
DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
CHI '26· Human-LLM Collaboration +3
- 75%
PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule Synthesis
CHI '23· Explainable AI (XAI) +2
- 75%
Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
CHI '23· Explainable AI (XAI) +2
- 75%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 75%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 75%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 75%
Unakite: Scaffolding Developers’ Decision-Making Using the Web
UIST '19· Explainable AI (XAI) +2
- 67%
When Help Hurts: Verification Load and Fatigue with AI Coding Assistants
CHI '26· Human-LLM Collaboration +3
- 63%
Understanding the Effect of Accuracy on Trust in Machine Learning Models
CHI '19· Explainable AI (XAI) +1
- 63%
Automation Accuracy Is Good, but High Controllability May Be Better
CHI '19· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)