Q3.13.3Report effect sizes and confidence intervalsdesignresearch

Results need an effect size and a confidence interval, not a star

Aliases: estimation reporting · effect size · interval estimate

What it is

A quantitative result that offers only a p value or a significant/not tag leaves readers unable to tell how large the difference is or how stable the estimate is. A complete report needs an effect size and a confidence interval: the first states magnitude on a unit or comparable scale, the second gives the range of magnitudes still compatible with the data. Without both, significance cannot be translated into whether to change, for whom, or at what risk. This is a reporting rule, not a separate lecture on interval width or testing philosophy.

Why it happens

A p value collapses magnitude, variability, and sample size into one tail probability; the collapse is not reversible. The same p can come from a large effect in a small sample or a small effect in a large one. Effect size pulls magnitude back out, in raw units (a conversion gap, seconds, people who succeed) or in a standardized metric. The interval shows the uncertain range, so “the point estimate is two percent” and “the interval runs from one percent harm to five percent gain” can be read apart. Omit both, and decision-makers are left to work a binary tag, then hear “significant” as “large” when they retell it.

Studying it

Pre-register the effect-size scale: absolute difference, relative difference, standardized coefficient, or counts, plus the interval’s coverage and model. State the main result in one sentence that carries the point estimate, the interval, and the sample; if p appears, it is an accessory. When several scales are defensible (relative lift looks dramatic on a low base), report the absolute gap as well. Figures hide intervals less easily than tables: points with error, not stars. A reproduction package should recompute the effect and interval from the same model, not merely re-obtain the significance call.

Where it stops holding

Interval coverage is model-dependent: a wrong independence assumption, uncorrected clustering, or selective reporting of which outcomes were tested all make nominal intervals too narrow. Standardized effects ease cross-study comparison and also force alignment of outcomes that are not interchangeable for product decisions; interface choices usually need raw units. Bayesian or bootstrap intervals can replace a classical confidence interval; they cannot replace the requirement to report magnitude and a range of uncertainty. Intervals on exploratory slices that never declared multiplicity are not confirmatory.

Applying it

  • Change the results template to: point estimate + interval + sample and units. Significance may not occupy the first sentence of the abstract on its own.
  • Choose the primary effect’s unit in advance; do not pick whichever of absolute and relative looks better after seeing the data.
  • Send materials back if they lack an interval; discuss shipping only after it is filled in.
  • Check: hide p, leave effect size and interval, and see whether a decision can still be made. If not, the report is still strapped to the significance tag.

Related

  • Same group: Q3.13.1 A significant result is not a reason to ship · Q3.13.2 Large samples make trivial gaps significant
  • Adjacent: Q3.21 Statistical significance and effect size · Q3.18 Experimental design and controls
  • Search terms: effect size · confidence interval · estimation reporting

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Q3.13.3