HINT: Integration Testing for AI-based features with Humans in the Loop

AI-Assisted Decision-Making & AutomationComputational Methods in HCIAI/ML Researchers & EngineersHCI Researchers

Document Title

HINT: Integration Testing for AI-based Features with Humans in the Loop

Document Information

  • Subject Areas: Artificial Intelligence, Human-Computer Interaction, Software Testing, User Experience
  • Keywords: Human-Computer Interaction, Prototype Testing, Testing Framework, Crowdsourcing, User Experience, AI Integration Testing, Evolutionary Testing, Multi-session Interaction, System Performance Evaluation

Research Background and Issues

  • Issues and Challenges

    • The dynamic nature of AI technology poses challenges for testing human-computer interaction and collaboration, especially in simulating real-world scenarios prior to deployment.
    • Current testing methods (e.g., offline performance evaluation, small-scale user studies, and A/B testing) have limitations, such as insufficient analysis of user behavior or high costs.
    • Isolated testing of AI models fails to adequately capture the dynamic changes in user-AI collaboration.
  • Importance

    • Artificial intelligence is widely applied in real-life scenarios (e.g., email management, content recommendation), with its success largely dependent on the effectiveness of human-AI collaboration.
    • Complex human-computer interaction experiences often require multiple interactions to reveal relevant issues or performance, which may be overlooked in traditional single-point testing.
    • In-depth testing prior to deployment can help reduce debugging costs and mitigate user attrition risks.
  • Research Motivation and Related Work

    • Current methods such as offline testing and traditional lab-based user studies are costly and limited in scope, making them ineffective for meeting the needs of rapid iteration.
    • This work builds upon and extends existing research on AI testing and human-AI collaboration testing, designing a scalable testing framework called HINT based on the concept of "integration testing."

Solution

  • Method and Framework

    • The proposed HINT (Human-AI Integration Testing) framework enables rapid and flexible testing of AI-driven features in dynamic human-computer interaction experiences.
    • HINT is inspired by the concept of integration testing in software development, integrating AI models with applications while involving real users (typically crowdsourced) in multi-session testing.
  • Innovations

    • Provides a crowdsourcing-based workflow that automates data collection and summarization.
    • Tests not only the offline performance of AI but also delves into user behavior changes and overall user experience after multiple interactions.
    • Quantifies user reactions to AI errors and potential long-term effects, such as changes in trust toward AI.
    • Generates detailed test reports, including interaction behavior data, subjective user feedback, and comprehensive analysis.
  • Implementation Steps and Key Techniques

    1. Test Preparation: Define task scenarios, design user tasks, and prototype AI functionalities for testing.
    2. Task Execution: Organize participants via crowdsourcing platforms to perform multiple rounds of tasks, simulating real-world usage scenarios.
    3. Data Collection: Record interaction processes through the AI user interface while collecting subjective user evaluations (e.g., trust, perceived utility).
    4. Report Generation: Compile interaction and feedback data into summary reports, enabling developers to visualize test results and identify issues.

Research Outcomes

  • Specific Outcomes

    • Experimental validation of HINT applied to two AI-driven email management functionalities: AI-based search and event detection features.
    • Tested HINT's sensitivity to various AI performance evolution patterns, including static performance, cross-session changes, and intra-session changes.
    • HINT revealed key user behavior patterns, such as when users trust AI, when they abandon reliance on AI, and the long-term impact of AI performance evolution on user perception.
  • Outcome Analysis

    • Experiments demonstrated that the HINT framework effectively captures dynamic user behaviors and correlates them with changes in AI model performance.
    • Highlighted user sensitivity to AI performance changes and corresponding behavioral adjustments (e.g., task completion time, adoption rates).
  • Experimental Results

    • HINT showcased subtle differences in user behavior under various AI evolution scenarios. For instance, users exhibited distinct trust and task adaptation patterns when using high-performance static AI versus low-performance dynamic AI.
    • The experiments involved 313 participants, covering diverse dynamic possibilities, validating HINT's effectiveness in assessing user task performance and subjective experiences.
  • Limitations and Future Directions

    • Limitations:
      • Highly dependent on the accuracy of task definitions, making it unsuitable for applications lacking clear goal-oriented tasks.
      • Current use of novice crowdsourced participants has not been validated with users possessing high domain expertise.
    • Future Work:
      • Explore extensions for personalized and open-ended AI functionalities (e.g., testing content recommendation services).
      • Enhance test reports with more interactive information filtering features.
      • Expand HINT into a decision-support tool to provide developers with clearer actionable recommendations.

Conclusion

The HINT framework addresses the gap in tools for evaluating user experience and collaboration in AI feature development. By designing dynamic multi-session interaction tests, it enables developers to comprehensively assess AI's user interaction performance before deployment, providing robust support for optimizing and improving human-AI collaboration systems.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/79965/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3490099.3511141
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
AI-Assisted Decision-Making & Automation, Computational Methods in HCI
work
Professions
AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers