Why Post-hoc Power Analysis Is Meaningless
After a study produces a non-significant result, a tempting response is to calculate how much statistical power the study had — using the observed data — and to report: "Although the result was not significant, our study was underpowered, so a real effect might still exist." This practice, known as post-hoc power analysis (also called observed-power or retrospective power analysis), is widespread. It is also scientifically indefensible.
The Mathematical Tautology
The fundamental problem is that post-hoc power calculated from observed data is entirely determined by the $p$-value and the sample size. There is a one-to-one mathematical relationship between an observed $p$-value and the observed power when that $p$-value is used as the basis for the power calculation. For a $z$-test or any large-sample test, a result with $p = 0.05$ will always yield an observed power of exactly 50%. For tests based on the $t$-, $F$-, or $\chi^2$-distribution the exact value depends on the degrees of freedom, but the fundamental relationship remains: observed power is entirely determined by the $p$-value and the sample size — it carries no independent information. A result with $p = 0.20$ will always yield an observed power well below 50%, regardless of the study design, the outcome, or the clinical field. The post-hoc power calculation adds no new information whatsoever — it is simply a rescaling of the $p$-value into a different number.
Reporting "our study had only 34% power to detect the observed difference" is therefore completely circular: it says nothing more than "our $p$-value was 0.20." The reader already knew that from the results section.
It Answers the Wrong Question
Power analysis is a planning tool. Before a study begins, it helps answer: "Given a true effect of a specified size, how likely am I to detect it with this design?" That is a meaningful prospective question that influences whether the study should be done at all, and how.
After the study is complete, the question "was my study powered?" is no longer answerable in any useful sense. The study ran; the data exist; the result is what it is. Calculating retrospective power from the study's own data tells you only what you already know from the $p$-value. The underlying question — whether there really is an effect, and whether the study was large enough to have a reasonable chance of finding it — cannot be answered by reprocessing the same data in a different form.
The Right Tool for Non-Significant Results: Confidence Intervals
When a study returns a non-significant result, the correct question is not "was the study underpowered?" but rather: "does the confidence interval allow us to rule out a clinically important effect?"
A wide confidence interval that spans from a large negative effect to a large positive effect genuinely tells us that the study was too small to be informative — but we can see that directly from the interval width, without any power calculation. The constructive response is to note that a larger study would be needed.
A narrow confidence interval that excludes the minimum clinically important difference (see Minimal important difference) tells a completely different story: the true effect, if it exists at all, is too small to matter. This is genuine evidence of absence — not the same as the absence of evidence. In this case, the non-significant result is informative and the study is not "underpowered" in any meaningful sense.
In summary: Power analysis belongs before the study begins, not after. Post-hoc power analysis from observed data is a mathematical tautology that tells you nothing beyond the $p$-value. When confronted with a non-significant result, report and interpret the confidence interval rather than calculating retrospective power.
The prospective framework for power and sample size planning is covered in the Sample size and statistical power chapter. For the correct interpretation of confidence intervals, see the P-values and confidence intervals chapter.