Improving Data Quality via Pre-Task Participant Screening in Crowdsourced GUI Experiments
Authors
Paper Title
Improving Data Quality via Pre-Task Participant Screening in Crowdsourced GUI Experiments
Publication Info
- Topic area: Enhancing data quality in crowdsourced GUI experiments through pre-task participant screening.
- Keywords: Crowdsourcing, GUI experiments, data quality, participant screening, size adjustment, movement time, error rate, predictive models, HCI, pointing tasks.
Background and Problem
- Problem / challenge: Crowdsourced GUI experiments often suffer from data quality issues due to inattentive or noncompliant participants, which degrade the validity of performance models.
- Significance: Improving data quality is critical for reliable model evaluation and for drawing valid conclusions in GUI interaction research, particularly in crowdsourced environments.
- Motivation and related work: Previous studies have explored data-quality management techniques such as attention checks and gold-standard tasks but have not specifically addressed GUI-based performance models. This work builds on these efforts by introducing a task-relevant pre-screening mechanism to improve data quality.
Solution
- Proposed approach: A pre-task screening method using a size-adjustment task to identify and exclude nonconforming participants before the main GUI experiment.
- Novelty:
- Introduction of a GUI-based pre-task that generates a continuous error signal to assess participant quality.
- Systematic evaluation of how pre-task thresholds and non-passing participant proportions affect model fit and predictive accuracy.
- Validation of the method across multiple devices (PC and smartphone) and error-handling policies.
- Procedure and key techniques:
- Participants perform a size-adjustment task where they resize an on-screen card to match a physical card.
- Participants with errors exceeding a threshold are excluded from the main pointing task.
- The main task evaluates movement time (MT) and error rate (ER) using established performance models (e.g., Fitts’ law).
- Simulations systematically vary the pre-task threshold, non-passing proportion, and sample size to assess model fit and predictive accuracy.
Results
- Concrete findings:
- Stricter pre-task thresholds and lower proportions of non-passing participants systematically improved model fit (R²) and predictive accuracy.
- For the ER model, R² increased by up to 98% under stricter thresholds and smaller non-passing proportions.
- Smartphone settings showed stronger benefits for both MT and ER models compared to PC settings.
- Advantage over baselines:
- The proposed method effectively screens out inattentive participants, improving model fit and reducing noise in the data.
- It outperforms traditional exclusion methods by providing a task-relevant continuous error signal.
- Experiments / evaluation:
- Three experiments were conducted:
- PC-based mouse interaction with re-aiming after errors.
- Smartphone-based touch interaction with re-aiming.
- Smartphone-based touch interaction without re-aiming.
- Metrics included model fit (R²) and predictive accuracy (via leave-one-out cross-validation).
- Three experiments were conducted:
- Limitations and future work:
- The method may exclude participants with motor or visual impairments, potentially introducing bias.
- Determining the optimal threshold T is nontrivial due to trade-offs with sample size and recruitment.
- Future work will explore alternative pre-tasks, broader task generalization, and improved guidance for threshold selection.
Summary
This paper introduces a pre-task screening method using a size-adjustment task to improve data quality in crowdsourced GUI experiments. By excluding nonconforming participants based on pre-task errors, the method enhances the goodness of fit (R²) and predictive accuracy of GUI performance models, particularly for error rate (ER) and movement time (MT). The approach was validated across three experiments involving PC and smartphone settings, showing consistent effectiveness regardless of device or error-handling policy. While the method is broadly applicable, future work is needed to address potential biases and optimize threshold selection. This screening method provides a practical tool for ensuring reliable data collection in GUI interaction research.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 60%
Creativity on Paid Crowdsourcing Platforms
CHI '20· Crowdsourcing Task Design & Quality Control
- 60%
CrowdSurfer: Seamlessly Integrating Crowd-Feedback Tasks into Everyday Internet Surfing
CHI '23· Crowdsourcing Task Design & Quality Control +1
Based on Jaccard similarity of research subtopics & professions (≥60%)