Towards Guidelines for Designing Human-in-the-Loop Machine Training Interfaces

Human-LLM CollaborationExplainable AI (XAI)AI-Assisted Decision-Making & AutomationComputational Methods in HCISoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI ResearchersStatisticians & Data Scientists

Title of the Paper

Towards Guidelines for Designing Human-in-the-Loop Machine Training Interfaces

Paper Information

  • Domain: Human-Computer Interaction, Interactive Machine Learning, Artificial Intelligence System Design
  • Keywords: Human-Computer Interaction, Machine Learning, Data Labeling, Algorithm Training, User Interface, User Experience, UX Evaluation

Research Background and Problem Statement

  • What problems or challenges did the authors identify?

    1. The "human-in-the-loop" process in Interactive Machine Learning (IML) poses a bottleneck to the scalability of training processes. Relying solely on human labeling tasks may fail to meet the demands of large-scale datasets and is prone to issues such as fatigue and labeling quality.
    2. There is a lack of design guidelines specifically tailored for creating interactive interfaces for IML systems, particularly regarding how to balance user agency, interaction burden, and model efficiency.
    3. Many existing interfaces fail to adequately support user agency or provide intuitive options, leading to user frustration.
  • Why is this problem important? Designing effective user interfaces for interactive machine learning can enhance system usability, reduce the burden of human labeling, and optimize the performance of machine learning systems. Systems that balance user agency and learning efficiency are critical for practical industrial applications, sustainable user experiences, and reducing training costs.

  • Research Motivation and Related Work

    1. Traditional machine learning systems are difficult to debug and may produce biased, irrelevant, or offensive results.
    2. Existing studies (e.g., Fails and Olsen [7]) have explored how interactive machine learning allows users to directly participate in training but lack systematic interaction design guidelines.
    3. Issues such as transparency and contextual adaptability have been identified as common pain points in several studies (e.g., Kulesza et al. [13][15], Amershi et al. [1]), encouraging the development of better design standards.

Proposed Solution

  • What methods or solutions did the authors propose? This study designed an experiment using four different interface variants for "human-in-the-loop machine learning training," allowing users to train a recommendation system through these interfaces. The interfaces were evaluated based on interaction speed, user experience, and efficiency parameters.

  • What is innovative about this solution?

    1. The authors proposed the IMLIQ scoring method (Interactive Machine Learning Interaction Quality), a comprehensive evaluation metric combining user experience and system learning efficiency.
    2. Preliminary design guidelines for interactive learning interfaces were proposed to optimize the balance between user experience and efficiency.
  • What were the implementation steps and key technologies used?

    1. Experimental Design:
      • Participants interacted with four different recommendation system interfaces to complete a task of recommending a specific target item (a red chair).
      • User interaction data (e.g., number of interactions and time to complete tasks) and subjective questionnaire evaluations (NASA-TLX for user burden) were collected.
    2. Interface Variants:
      • Interface 1: A binary labeling task where users selected "like" or "dislike" for specific items.
      • Interface 2: Similar to Interface 1 but allowed users to temporarily disable the influence of specific features (e.g., color or style).
      • Interface 3: Comparative selection, where users chose the closest match to the target item from three options.
      • Interface 4: After making a selection, the system asked users to specify the reason for their choice.
    3. Evaluation Metrics:
      • Efficiency: Recorded the number of interactions and time required to complete tasks.
      • User Experience: Collected subjective scores on psychological, physical, and temporal workload using NASA-TLX, along with additional assessments of interactivity and enjoyment.

Research Findings

  • What specific results were achieved?

    1. Users preferred interfaces with greater flexibility and expressive capabilities (Interfaces 2 and 3 received the most votes).
    2. Users tended to favor interfaces that avoided "irreversible" or repetitive tasks, perceiving them as more efficient and enjoyable.
    3. By comparing the interaction performance of the four interfaces, preliminary design guidelines for interactive learning interfaces were derived.
  • What advantages does this solution have over existing ones? The proposed design guidelines and IMLIQ scoring model intuitively integrate user experience and system performance, offering a new perspective for the integration of human-computer interaction and machine learning interfaces.

  • What were the experimental or evaluation results?

    • User Preference Distribution: Interfaces 2 and 3 received the most user votes (6 votes each); Interface 1 received 2 votes, and Interface 4 received only 1 vote.
    • Performance Comparison: In the IMLIQ model scoring, Interface 3 achieved the highest score (8.2), while Interface 4 scored the lowest (-1.5).
InterfaceAvg. Interaction CountEffort Score (1-10)Interactivity Score (1-10)Enjoyment Score (1-10)
Interface 16.02.44.54.9
Interface 25.02.36.66.1
Interface 35.02.66.06.6
Interface 46.03.25.74.7
  • Limitations and Future Directions:
    1. The current experiment had a limited sample size, with participants primarily from higher education backgrounds, potentially biasing results toward users familiar with ML systems.
    2. The test scenarios were relatively simple and may not fully represent real-world industrial applications.
    3. The proposed IMLIQ scoring method requires further large-scale experimental validation, particularly in establishing empirical weights for subjective scores.

Conclusion

The design of interactive machine learning system interfaces requires a balance between user experience and system efficiency. This study highlighted key issues in designing human-in-the-loop training systems and explored the applicability of different interfaces through experimentation. It proposed four design guidelines and the IMLIQ scoring method as a reference framework for future research and design. This work provides new insights into the intersection of HCI and machine learning, with future validation needed in more complex scenarios and with broader user groups to confirm these preliminary findings.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/57994/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3397481.3450668
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), AI-Assisted Decision-Making & Automation, Computational Methods in HCI
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers