Q3.21.2A priori power versus post hoc n judgmentdesignresearch

Without a priori power analysis you cannot later judge whether n was enough

Aliases: a priori power · observed power · post hoc sample-size judgment

What it is

An a priori power analysis happens before results are seen. It uses the smallest effect of interest, variability, α, and the design to compute how much independent information is needed. It answers a planning question: does this n have a reasonable chance of detecting that effect. Asking after the results whether “n was enough,” using observed / post hoc power, is almost a monotone transform of p: power looks low when the test was not significant and high when it was, and it does not tell you anything extra about whether the sample was short. Without the beforehand calculation, you cannot judge after the fact whether n was enough—not because small samples cannot be complained about, but because that after-the-fact number carries no independent information. How many people a quantitative study should recruit, planned from effect size, is a separate job finished before launch.

Why it happens

Power is the probability the test will reject if the effect equals some value specified in advance. That value has to come from the plan, not from the point estimate just obtained. Feeding the observed effect back into the formula describes data that have already decided significance: large p makes observed power low, which looks like “underpowered”; small p makes it high, which looks like “powered.” “Non-significant because n was too small” and “significant because n was enough” are then tautologies. The real basis for saying n was short is the gap between the pre-study target for the smallest effect of interest and the realized effective sample, after attrition and clustering. Without that pre-study table, what remains is a restatement of p.

Studying it

Before launch, lock the smallest effect of interest, α, target power, and design (including clustering and the primary outcome), and emit a target n plus a sensitivity range. During the run, track effective sample (independent units, not expanded rows). If it falls short, extend or accept poor precision; do not peek at p to decide whether to stop. The results cite the pre-study table and compare planned n with realized n. After a non-significant test, the account is how wide the interval is and whether it covers the effect of interest—not observed power. If no power calculation was done at the start, what can be added later is a precision statement (interval width) and a calculation of information relative to a pre-stated effect of interest; still do not report observed power as a diagnosis.

Where it stops holding

Sequential or adaptive designs may stop under a prewritten rule; that is not a post hoc power story built after seeing p, and the stopping rule must be written first. If the goal is estimation precision rather than a test, plan n from interval width directly; power language can drop out, but the step “declare a precision target beforehand” cannot. In exploratory work where the measure is still unstable, forcing a two-group power formula produces a spuriously exact headcount; simulate, or run a measurement study first. Huge samples making trivial gaps significant is a different problem: planning power on the smallest effect of interest itself refuses to pile n for trivia.

Applying it

  • An experiment ticket needs an a priori power or precision calculation before traffic is split; no attachment, no launch.
  • Do not write “observed power was only xx%, so the sample was too small” on the results page; write planned n, realized n, and the interval.
  • When the result is not significant and realized n met the pre-study target, close as “information reached the plan for the effect of interest”; do not keep running until significance.
  • Check: hide p, leave the pre-study table and the interval. If it is still impossible to say whether n was enough, the judgment is still tied to observed power.

Related

  • Same group: Q3.21.1 A p-value is compatibility, not the chance the effect is real · Q3.21.3 A one-tailed test lowers the bar and must be declared first · Q3.21.4 Width of the interval tells precision better than whether it includes zero
  • Adjacent: Q3.13 Significance and practical importance · Q1.08 Sample size
  • Search terms: a priori power · observed power · sample size planning

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Q3.21.2