Removing a data point is only legitimate if the criterion was set before looking at the outcome
Aliases: exclusion criteria · removal rules
What it is
Removing any point from data is a decision that affects the conclusion, and that decision's legitimacy depends not on the removal itself but on whether the criterion was declared in advance, whether it is independent of the conclusion, and whether it was applied consistently to all the data. "We removed a few outliers" states nothing—how many were removed, by what rule, what the conclusion was before removal—so readers can neither assess the removal's reasonableness nor rule out "the removal happened to make the conclusion look good." Transparency of removal criteria is a foundational component of analytical credibility.
Why it happens
The mechanism making undeclared removals untrustworthy is the abuse potential of researcher degrees of freedom: the same dataset under different removal criteria can yield different or opposite conclusions, and if the criterion is chosen after seeing the results ("dropping these points gets p just under .05"), removal turns from data cleaning into conclusion manufacturing. This is rarely deliberate fraud—people rationalize after the fact, and the intuition "that point just looks wrong" often arises only after seeing the data, with the intuition itself contaminated by the conclusion. Declared criteria work by moving the choice from "post-hoc discretion" to "pre-hoc commitment": a preregistered rule ("remove records with click duration under 1 second") cannot be tuned to a particular conclusion, making the removal not merely acceptable but credibility-enhancing; post-hoc rules can also be assessed for their independence from the conclusion variable and their consistent application. Transparency hands the evaluation key to readers instead of keeping it with the author.
Studying it
Methods for studying removal's impact include many-analysts studies: the same dataset and research question given to independent teams, recording each team's chosen criteria and conclusions to quantify how criterion choice alone shapes the distribution of outcomes—such studies show criterion differences alone can flip conclusions from significantly positive to significantly negative. On the design side, one can test how transparency of removal information (annotating count and rule vs not annotating) affects reader trust and willingness to verify. A methodological caveat: a criterion's "reasonableness" and its "transparency" are independent dimensions—a reasonable but undeclared criterion still cannot be evaluated, and a transparent but unreasonable one ("drop the three smallest values") will be identified as a problem; both must be reported.
Where it stops holding
The requirement for declared criteria scales with the conclusion's stakes: internal exploratory charts can simply annotate "N outliers removed," while externally published conclusions (papers, regulatory reports) need the full criterion definition, counts, and a before/after comparison of conclusions. The "in advance" degree also grades—the ideal is preregistration (locked before analysis), the next best is criteria independent of the outcome variable (data-quality rules like collection failures), and the minimum is post-hoc declaration plus sensitivity analysis. Some fields' standard criteria are themselves contested (psychology's removal of trials under 200ms is convention, but boundary choices affect results), and there, citing the convention's source plus sensitivity checks matters more than citing the convention. Finally, not removing is also a decision—an analysis keeping outliers must explain why (uncertain provenance and conclusions robust to outliers), since transparency requirements run both ways.
Applying it
- Write every removal rule into the analysis document: the precise definition, count removed, proportion, and impact on the conclusion (core metrics before vs after).
- Annotate removal information directly on charts (footnote or tooltip: "37 failed-collection records removed, 1.2%; conclusions identical before and after").
- Prefer outcome-independent quality definitions (collection failures, abnormal durations) over outcome-dependent ones ("3 standard deviations from the mean" is acceptable when defined pre-analysis, unacceptable when chosen post-hoc to fit conclusions).
- Verification: take a report with removals and check whether the removal process can be reproduced from the report alone (which points, why); if not, the declaration is insufficient.
Related
- Same group: U8.05.1 An outlier may be an error or a discovery · U8.05.3 Axis compression hides the body of the distribution
- Nearby: U8.05.1 An outlier may be an error or a discovery · U8.03.2 Uncorrected multiple comparisons inflate significance markers systematically
- Search terms:
exclusion criteria·researcher degrees of freedom·preregistration