U8.03.5Showing effect sizes with intervals conveys more than significance labelsdesign

An effect size with its confidence interval tells readers far more than a bare significance star

Aliases: estimation display · new statistics

What it is

A significance marker conveys one bit of information ("crossed the threshold / did not"), while an effect size plus its confidence interval conveys a continuous quantity together with its uncertainty range. From an information standpoint, the latter contains everything needed for the former's judgment (whether the interval crosses zero is equivalent to significance) and adds two dimensions that significance markers cannot provide at all: "how big is the effect" and "how precise is the estimate." Decades of methodological discussion in statistics ("the new statistics," the effect-size reporting movement) converge on exactly this claim: estimates and their intervals carry more usable scientific information than binary significance judgments. The direct implication for data visualization: charts should foreground effect sizes with intervals, with significance markers as supplements rather than the main event.

Why it happens

The information advantage becomes clear from the set of questions readers need answered: "Is this difference worth acting on?" requires the effect size. "How much can this estimate be trusted?" requires the interval width. "Will this replicate?" requires the precise estimate with its interval. "Where does this conclusion apply?" requires the interval's position relative to business thresholds. A significance marker answers none of these. Conversely, "is it significant" can be derived from the effect size plus interval (an interval not crossing zero is significant)—the information is one-directionally inclusive. Displaying intervals also has a hidden educational effect: readers who repeatedly see interval widths vary with sample size gradually build the intuition that "all estimates carry uncertainty," while binary star presentation offers no such learning signal. Practical resistance comes from the fact that effect-size-plus-interval asks readers to make one inference step (checking the interval against the threshold) instead of reading a label, raising the cognitive cost for statistically less-trained readers—a source of design compromise, not a reason to reject estimation display.

Where it stops holding

Estimation display does not suit every decision context: when the decision structure is itself binary (a deployment switch is on or off), the judgment the reader ultimately needs is binary, and the interval's continuous information may exceed what the decision requires. The more effective design there is to draw the interval together with the decision threshold (coloring above and below the threshold line) so "does the interval cross the threshold" reads directly—combining information completeness with decision convenience. Another boundary is reader interpretation skill: reading estimates and intervals requires basic statistical literacy, and a business-facing dashboard that drops stars outright may lead readers to misinterpret intervals (reading width as "error size" rather than "plausible range"). A transitional strategy presents both, with the interval as the primary visual, gradually de-emphasizing stars as reader literacy grows.

Applying it

  • Prioritize the difference value and its confidence interval (difference plots) on comparison charts, drawing the reader's needed inference directly.
  • Draw the confidence interval and the decision threshold on the same chart so their relative position reads directly, replacing the inference step that stars demand.
  • Keep stars only as auxiliary markers in secondary positions, with the caption stating the full judgment basis (effect size, interval, threshold).
  • Verification: compare two versions of the same experiment's chart (stars vs interval) and ask business readers whether to ship; the interval version's readers should be able to state both the effect size and the certainty, while the stars version yields only a binary verdict.

Related

  • Same group: U8.03.1 Significance markers do not convey the size of the effect · U8.03.2 Uncorrected multiple comparisons inflate significance markers systematically · U8.03.3 There is no visual cliff between significant and non-significant · U8.03.4 Charts annotating significance must state the test used
  • Nearby: U8.01.1 Point estimates hide interval information · U8.02.2 Overlapping intervals between two groups do not imply a non-significant difference
  • Search terms: estimation display · new statistics · effect size reporting

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/U8.03.5