Multiple comparisons need a corrected threshold or a primary metric named in advance
Aliases: multiplicity correction · Bonferroni · primary-endpoint restriction
What it is
Once an experiment will test several metrics, the error rate is kept honest either by correcting the threshold (Bonferroni, Holm, false-discovery rate, and similar) so each bar tightens, or by collapsing confirmatory right into one primary named before the run, with the rest demoted to secondary or description. Both moves exist so that “significant” still matches an agreed error rate, not to settle which correction formula is best. With neither correction nor a narrowed primary, nominal α holds for one test and not for the table. That the primary must be locked before launch is the next leaf’s timing rule; this leaf is the many-to-one analysis strategy.
Why it happens
Correction splits α across tests: Bonferroni uses α/k, harsh and simple; Holm relaxes step by step; FDR controls the fraction of discoveries that are false, not the probability of at least one false discovery. Naming a primary shrinks the confirmatory family to one test; secondaries that clear a line do not get promoted. The costs differ: correction makes every metric harder to pass; a single primary strips confirmatory status from the rest. With no strategy, a team sits with a table still painted at 0.05 and chooses in the meeting; the error rate is already not 0.05. Correlation among metrics makes some corrections conservative. That is a reason to use a method that matches the dependence (a multivariate test, a gate: secondaries tested only if the primary passes), not a reason to skip correction.
Studying it
The plan picks one or a combination: name the correction and k, or name the unique primary and the secondary list. Gatekeeping or hierarchical tests need the order drawn. Simulate under the planned dependence and check whether family-wise error or FDR sits near the target. The results table splits columns: primary / secondary / exploratory, and marks what was corrected. Uncorrected p on secondaries may be shown; the conclusion sentence may be carried only by the primary or by the set that still stands after correction. Switching correction methods is a protocol deviation, recorded, not a swap because another method looked better.
Where it stops holding
A single pre-registered metric and no other tests need no correction. Descriptive charts that are not read as tests need none. An exploratory stage may skip correction if it drops confirmatory language. FDR fits screening many endpoints of equal status, not a ship/no-ship experiment with one decision—those more often need family-wise control or a single primary. Correction cannot rescue a list whose k was revised after seeing results: the list has to be fixed beforehand.
Applying it
- Before the split, the ticket ticks: correction method + k, or a unique primary. Both blank, no launch.
- Results tables use three identity columns (primary, secondary, exploratory); the exploratory column may not say “confirmed.”
- Do not switch correction methods in the meeting to let a metric through. A new method is a new experiment.
- Check: recompute the whole table under the registered rule. Any metric that does not pass that rule leaves the conclusion.
Related
- Same group: Q3.22.1 Testing many metrics raises false positives · Q3.22.3 Reporting the significant slices is data peeking · Q3.22.4 The primary is locked before launch, not chosen after
- Adjacent: Q3.21 Statistical significance and effect size · Q3.05 Multivariate testing
- Search terms:
multiple-comparison correction·Bonferroni·primary endpoint