How wide a confidence interval is tells precision better than whether it includes zero
Aliases: interval width · estimation precision · crossing zero is not precision
What it is
A confidence interval is the range of effect magnitudes still compatible with the data and the model. How wide it is says how imprecise the estimate is. Whether it includes zero only restates the significance test—including zero is roughly “do not reject the null,” excluding zero is roughly “reject.” Two intervals can both exclude zero, one running from 0.1% to 20% and one from 4.8% to 5.2%, with wholly different precision. Two can both include zero, one hugging zero and one stretching from large harm to large gain. Width is the precision statement; crossing zero is a binary replay of the test. Requiring an effect size and an interval in the report is about not leaving decisions with a star; the object here is the width itself.
Why it happens
The endpoints open with the standard error. Little independent information, large variability, or a model that treats dependence as independence all stretch the interval. A wide interval means many magnitudes—harmful, trivial, and substantial—remain compatible, so the point estimate is almost unusable as a basis for action. A narrow interval tightens the compatible magnitudes, whether or not zero happens to sit outside. Reading only “includes zero or not” throws the length away and keeps the one bit that is equivalent to a p threshold. Coverage also depends on the model: a nominal 95% under a false independence claim is systematically too narrow, looking precise while not being precise. Width has to be read with the model, not as decorative error bars.
Studying it
Declare in advance the coverage, the scale (absolute or relative), and a width target for precision, such as “half-width no larger than half the smallest effect of interest.” State the main result in one sentence with the point estimate and both ends; do not say “significant, therefore…” and bolt the interval on later. Compare width with the pre-study target: wider than the target is tagged imprecise, whether or not zero is inside. Draw the interval as a segment, not as a star. Intervals uncorrected for clustering or multiplicity should note that nominal width is too small. Reproduction should redraw the interval from the same model, not merely recover “crossed zero or not.”
Where it stops holding
Bayesian credible intervals and bootstrap intervals also use width to talk about precision, with their own probability talk, and “do not only look at whether zero is inside” still holds. Intervals on standardized effects ease cross-study comparison; decisions often still need width on the original unit. Equivalence tests ask whether the interval falls inside pre-set equivalence bounds—a joint use of width and location, not a new “cross zero” rule. A nominally narrow interval on an exploratory slice that never declared multiplicity looks precise while selection has already broken coverage.
Applying it
- First line of the result: point estimate, both ends, and half-width. “Includes zero?” comes after, never as the heading by itself.
- Write the largest acceptable half-width beforehand; exceed it and tag “not precise enough,” even if zero is outside, and do not publish as confirmed.
- Zero inside and very narrow: write “the effect is precisely limited to a trivial range,” not “the experiment failed.”
- Check: show only the interval segment, hide the includes-zero tag and p. If decision-makers still cannot tell a precise small effect from “anything is possible,” the reading is still strapped to crossing zero.
Related
- Same group: Q3.21.1 A p-value is compatibility, not the chance the effect is real · Q3.21.2 After the fact you cannot tell whether n was enough · Q3.21.3 A one-tailed test lowers the bar and must be declared first
- Adjacent: Q3.13 Significance and practical importance · Q1.08 Sample size
- Search terms:
confidence interval width·precision·compatibility interval