Normalizing Speed-accuracy Biases in 2D Pointing Tasks with Better Calculation of Effective Target Widths
Paper Title
Normalizing Speed-accuracy Biases in 2D Pointing Tasks with Better Calculation of Effective Target Widths
Publication Info
- Topic area: Human-computer interaction (HCI) and performance evaluation in 2D pointing tasks.
- Keywords: Fitts' law, speed-accuracy tradeoff, effective target width, univariate standard deviation, bivariate standard deviation, throughput, ISO 9241-411, pointing tasks, HCI evaluation, crowdsourcing.
Background and Problem
- Problem / challenge: The calculation of effective target width (We) in 2D pointing tasks remains debated, particularly whether univariate (σx) or bivariate (σxy) standard deviation better normalizes speed-accuracy biases. Prior studies, such as Wobbrock et al. (2011), recommended σxy but tested only under a single bias condition, leaving its effectiveness across multiple biases unverified.
- Significance: Proper normalization of speed-accuracy biases is critical for fair comparisons of devices, interaction techniques, and user groups in HCI studies.
- Motivation and related work: ISO 9241-411 and earlier studies recommend σx for We calculation, but Wobbrock et al. proposed σxy for 2D tasks. However, no prior work systematically compared these methods across multiple speed-accuracy biases or in large-scale crowdsourced settings. This paper addresses these gaps.
Solution
- Proposed approach: A systematic comparison of univariate (σx) and bivariate (σxy) standard deviations for calculating We in 2D pointing tasks under three speed-accuracy biases (accurate, neutral, fast).
- Novelty:
- Demonstrates that σx better normalizes speed-accuracy biases than σxy in 2D pointing tasks.
- Shows robustness of σx through random sampling simulations, addressing small-sample scenarios common in HCI studies.
- Provides practical recommendations for Fitts' law experiments, including task-axis definitions and amplitude choices.
- Procedure and key techniques:
- Conducted a crowdsourced experiment with 346 participants performing ISO-style 2D pointing tasks under three bias conditions.
- Compared model fit (R², AIC, BIC) and throughput (TP) stability across nine ID computation methods, varying σ (univariate vs. bivariate), amplitude (A vs. Ae), and task-axis definitions (target-to-target vs. click-to-target).
- Performed Monte Carlo simulations to evaluate robustness of findings across sample sizes.
Results
- Concrete findings:
- The univariate model ID**xTT achieved the highest model fit (R² = 0.967537) in the mixed-bias condition, outperforming all bivariate models (R² < 0.96).
- Throughput stability was best with IDxTTAe, with TPdiff = 4.74241% and TPcv = 2.77263%.
- Differences between A and Ae, as well as between TT and CT task axes, were minor but favored Ae and TT.
- Advantage over baselines:
- ID**xTT improved R² by ≈0.095 points over nominal ID (R² = 0.872731) in the mixed condition.
- Univariate models consistently outperformed bivariate models in both model fit and TP stability metrics.
- Experiments / evaluation:
- Dataset: 152,226 trials from 342 participants after outlier removal.
- Metrics: Movement time (MT), error rate (ER), throughput (TP), model fit (R², AIC, BIC), and TP stability (TPdiff, TPcv).
- Simulation: Monte Carlo subsets of participants (N = 5, 10, 20, 40, 80, 160) confirmed robustness of univariate models.
- Limitations and future work:
- Limited to mouse input; results may differ for other devices or VR systems.
- Relatively low task difficulty (ID = 2.07–4.70 bits) compared to some prior studies.
- Future work should explore stricter instruction compliance, broader ID ranges, and laboratory replications.
Summary
This study systematically compared univariate (σx) and bivariate (σxy) standard deviations for calculating effective target width (We) in 2D pointing tasks under multiple speed-accuracy biases. Results showed that σx better normalizes bias effects, achieving higher model fit and more stable throughput than σxy. The findings were robust across sample sizes, supporting the use of univariate models (ID**xTT for fit, IDxTTAe for stability) in HCI studies. Practical recommendations include using Ae, the TT task axis, and recruiting at least 20 participants for reliable results. These conclusions provide clarity for standardizing performance evaluations in future Fitts' law experiments.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)