A group difference is a treatment effect only after other variables are controlled
Aliases: confound control · incomparable groups · experimental control
What it is
A gap between two groups can be read as a treatment effect only after factors other than the treatment have been controlled—held constant, randomly scattered, or balanced in blocks. Otherwise the gap carries device, time of day, task difficulty, recruitment source, facilitator, or defaults: confounds. Having a second arm answers “is there another group?”; it does not by itself make the groups differ only in treatment. The work here is keeping other variables in check, not picking a significant metric after the fact, and not the slogan that random split in a product A/B is what licenses a causal reading.
Why it happens
A between-group gap is a sum of channels. People who opt into a new version are more fluent; one lab arm always runs in the afternoon on the old lab phone; only the treated arm gets the instruction sheet; the control task is actually harder. Those channels are tangled with treatment, so the observed gap cannot name which channel moved. Experimental control fixes what can be fixed (same task script, same device pool), randomizes or blocks what cannot (time of day, experience strata), and blocks extra help that only one arm receives. Statistically “controlling for a few covariates” repairs dimensions you thought of, not unmeasured channels; it does not replace control in assignment and in the procedure. When an uncontrolled variable still sits in the gap, writing the gap as treatment treats entanglement as separation.
Studying it
Put on an operations checklist every factor that should stay identical and every factor that should be randomized or blocked, and record the realized values in the run (device, slot, facilitator, task version). After assignment, compare those records across arms; on a hole, fix operations before reading results. Pre-register the primary treatment contrast, and declare which covariates are part of the design (blocking variables) and which are only robustness checks. Sensitivity: bring unbalanced factors in one at a time and see whether the treatment estimate moves; a large move means the main claim still carries an uncontrolled confound. Do not replace the record with “the groups looked similar.”
Where it stops holding
Total control can sand the setting until it no longer resembles use; ecological validity and internal control trade, and cleaner is not always better. Some variables are part of the treatment (the new version must change the typeface) and should not be partialled out as confounds, but the claim must say the moved object is the bundle. Matching or regression “control” on observational data still leaves selection and unmeasured confounding; that claim sits below experimental control. Random assignment balances unmeasured factors in expectation; when implementation fails (one source falls entirely into one arm), that expectation is void and the work returns to operational control.
Applying it
- When writing the study, list “what both arms must share” and “what may differ: the treatment only.” Turn the first list into a checklist signed before launch.
- Lock task script, device pool, and facilitator script during the run; one arm must not receive prompts or gifts the other lacks.
- If a covariate splits across arms, stop reading a treatment effect; redo assignment or restrict the population.
- Check: show only the operations log (no outcomes) to someone outside the design and ask what else could explain a group gap. If they can name a factor that was not balanced, the conclusion is not a treatment effect.
Related
- Same group: Q3.18.1 Without a control, treatment and time trend cannot be separated · Q3.18.2 Randomize at the unit where the effect occurs · Q3.18.3 Within-subjects designs need order handled by counterbalancing
- Adjacent: Q3.04 A/B testing · Q3.05 Multivariate testing
- Search terms:
confound·experimental control·ceteris paribus