P4.10.4Adding data cannot fix sampling biasdesign

Supplementary data cannot fix the skew of the collection mechanism itself

Aliases: more data · same bias · bias versus variance in data collection · channel-limited correction

What it is

The instinctive response to discovering biased data is "collect more," but adding volume and correcting skew are different things: sampling bias comes from the collection mechanism's systematic filtering of populations, and repeating collection under the same mechanism yields only a more precise version of the same shaped distribution. Statistically this is the bias–variance distinction — more data reduces variance (steadier estimates) while bias (estimates systematically off target) does not move. A collection channel promoted only in major domestic cities enrolls roughly the same rural share with ten thousand participants as with ten million; what is missing is not sample size but a collection path that reaches the absent.

Why it happens

Bias is stubborn at the mechanism level in three ways. Channel lock-in: a channel's population filtering is a property of the channel, not sampling accident, so scaling within it only replicates the filter more finely (rural users stay in single-digit percentages even in a tens-of-millions urban-app sample). Incentive selection: adding rewards to boost volume recruits those responsive to incentives — precisely the already over-represented — so top-up sampling leans further toward them. Aggregation masking: after scaling, aggregate metrics look better (large samples, narrow confidence intervals), and the skew hides behind statistical significance, making the problem harder to see — the more data, the firmer the misplaced confidence. All three point to one conclusion: the fix for skew lies in the collection mechanism (new channels, lower access barriers, new participation modes), not in volume.

Where it stops holding

"More data is useless" is not an argument against all data supplementation: for performance fluctuation caused by random noise (a variance problem), adding volume is exactly the right fix; the conclusion applies only to systematic skew. Switching mechanisms to repair coverage is not plug-and-play either — new channels introduce new behavioral differences and new labeling-quality problems, so after a coverage fix the model usually needs recalibration rather than simple continued training. One more genuine option is deliberate scope contraction: for populations known to be out of reach, an explicit feature boundary plus human fallback is more honest than forced data collection, and often cheaper.

Applying it

  • Diagnose first: before topping up data, run a composition audit to classify the skew as "insufficient volume" or "mechanism filtering" — if group shares barely move across volume increases, it is a mechanism problem and more volume will not help.
  • Spend the de-biasing budget on new collection paths: enter channels the absent populations actually use (SMS for non-smartphone users, community field sites, local-language interfaces) rather than doubling down on the original channel.
  • Set a stop-loss line for scaling: when group shares show no significant change across two consecutive batches of same-channel top-up sampling, freeze further collection spend on that channel and redirect it to mechanism repair.
  • Verify: rerun the coverage audit after every supplementary collection, with "convergence of target group shares toward the external benchmark" as the only pass criterion — growth in total volume does not count as evidence.

Related

  • Same group: P4.10.2 Collection channels decide who never enters the data · P4.10.3 Historical data records past decisions, not facts
  • Adjacent: P4.03.3 Stratified evaluation · P4.08 Environmental cost and volume growth
  • Search terms: sampling bias · bias versus variance · data augmentation limits

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/P4.10.4