Clinical versus Statistical Significance
One of the most consequential misunderstandings in medical research is treating a statistically significant result as clinically important, and a non-significant result as evidence of no effect. A $p$-value below 0.05 says something precise and limited: that the observed data would be unlikely if the null hypothesis were exactly true. It says nothing about whether the effect is large enough to matter to patients or to clinical practice. Conversely, a non-significant result from a small study may reflect inadequate power rather than a genuinely absent effect. Disentangling these three concepts — statistical significance, clinical significance, and the minimal important difference — is essential for both designing and reading clinical research.
Statistical Significance
The mechanics of $p$-values and confidence intervals — how they are calculated, what they measure, and their mathematical relationship — are covered in the P-values and confidence intervals chapter. Here we focus on what statistical significance does and does not imply in clinical context.
The threshold $p < 0.05$ has no special scientific status. It was adopted by convention following R.A. Fisher's suggestion in the 1920s and has persisted largely through inertia. Treating it as a bright line — splitting results into "significant" and "non-significant" as if separated by a conceptual chasm — discards the information carried by the continuous scale of evidence that a $p$-value represents. A result with $p = 0.049$ and one with $p = 0.051$ are effectively identical in evidential terms; reporting one as a positive finding and the other as negative is scientifically incoherent.
A confidence interval is almost always more informative than a $p$-value alone: it communicates the of the effect, its , and the with which it has been estimated — all of which the $p$-value discards. These are the quantities that actually matter for clinical interpretation.
A further complication arises when many tests are conducted in the same study. If 20 independent tests are performed at the 5% significance level and all null hypotheses are true, one spuriously significant result is expected on average. This inflation of the Type I error rate by multiple testing demands either adjustment of the significance threshold (the Bonferroni correction divides α by the number of tests) or, in exploratory settings, control of the false discovery rate. The pre-specification of primary and secondary outcomes in a registered trial protocol is the most transparent way to limit the scope for post-hoc data-dredging.
Effect Size and Clinical Significance
A statistically significant result answers the question "is the effect non-zero?" — which is almost never the scientifically interesting question. What matters clinically is whether the effect is large enough to be worth acting on. This is the domain of clinical significance, and it requires a different set of tools: effect size measures and pre-specified thresholds of importance.
Effect size quantifies the magnitude of a difference or association independently of sample size. Common measures include Cohen's $d$ (the difference between two means expressed in standard deviation units), the odds ratio and relative risk for binary outcomes, and the hazard ratio for time-to-event outcomes. These allow comparisons across studies and outcomes in a way that raw differences in original units sometimes do not.
The contrast between statistical and clinical significance is sharpest at the extremes of sample size. Consider two examples. A trial enrolling 50 000 hypertensive patients finds a mean reduction in systolic blood pressure of 0.5 mmHg in the treatment arm, with $p = 0.003$. The effect is statistically significant — it is real and almost certainly not due to chance — but no cardiologist would alter prescribing on the strength of a 0.5 mmHg reduction, which is smaller than the measurement error of most sphygmomanometers. Now consider a trial of 20 patients with the same condition that finds a mean reduction of 12 mmHg, with $p = 0.12$. The result is non-significant — but the point estimate is clinically meaningful, the confidence interval is wide, and the most likely explanation is that the trial was simply too small to detect a genuine effect reliably. The first trial is unambiguous but useless; the second is inconclusive but informative.
The Minimal Important Difference
The concept that formalises clinical significance is the minimal important difference (MID) — also called the minimal clinically important difference (MCID). It is defined as the smallest change in an outcome that a patient or clinician would recognise as genuinely meaningful. The MID sets the clinical standard against which to judge whether a statistically significant effect clears the bar of practical relevance.
Establishing the MID for a given outcome is an empirical and clinical undertaking, not a statistical one. The most direct approach is anchor-based: patients or clinicians are asked to rate whether they have improved, and the MID is estimated as the average change in the outcome score among those who report a small but definite improvement. A global rating of change scale — "compared to before treatment, how are you now?" — is a typical anchor. An alternative is to anchor the outcome change against an external clinical milestone, such as the point at which a patient can perform a specific physical task that they could not before treatment.
When patient-level data are unavailable, distribution-based methods offer a statistical approximation. The most common is half a standard deviation of the baseline score, which roughly corresponds to a "small" effect in Cohen's taxonomy. One standard error of measurement is another candidate. These methods are computationally straightforward but carry the limitation that they capture statistical dispersion, not patient perspective; an outcome that is highly variable between patients may have a large half-SD that exceeds what any individual patient would actually notice.
For some widely used instruments, MIDs have been established through formal research and are available in the literature: approximately 13 mm on a 100-mm visual analogue scale for pain, meaningful changes on the WOMAC score in osteoarthritis, and condition-specific thresholds for the SF-36 health survey. For novel outcomes or instruments, a formal consensus exercise — such as a Delphi panel — may be needed to establish an agreed threshold before the study begins. What is not acceptable is to define the MID seeing the results: that is simply rationalising whichever effect happened to be observed.
Sample Size, Power, and the Distinction
The relationship between sample size and statistical significance is mechanical: with a sufficiently large sample, any non-zero effect will eventually reach statistical significance, no matter how small. This is an arithmetic inevitability, not a statement about clinical importance. It means that in large studies — nationwide registries, mega-trials — statistical significance is nearly guaranteed for any effect that is not precisely zero, and the $p$-value conveys almost no information about clinical relevance. The appropriate question in such settings is not "is the effect significant?" but "is the effect larger than the MID?"
At the other extreme, a small study may fail to detect an effect that is clinically important, because the study had insufficient statistical power — the probability of detecting a true effect of a specified size. Power should be calculated before a study begins, targeting the MID rather than an arbitrarily small effect, so that a non-significant result is genuinely informative and not merely a failure of scale. The details of power and sample size calculation are covered in the Sample size and statistical power chapter, including the important point that post-hoc power analysis — calculating power from the same data that produced the non-significant result — is circular and adds no information beyond the $p$-value itself.
Some trial designs make the MID structurally explicit. Non-inferiority and equivalence trials pre-specify a margin $\Delta$ — the largest difference that would still be considered clinically acceptable — and test whether the true effect falls within that boundary. A non-inferiority trial asks: "is the new treatment no worse than the standard by more than $\Delta$?" rather than "is the new treatment better?" This framing formalises the role of clinical judgement in the design stage and produces conclusions that are directly relevant to practice, without the ambiguity of interpreting a conventional $p$-value.
A Practical Framework
When interpreting any result, placing it in a two-by-two framework clarifies what the finding does and does not imply:
| Effect ≥ MID (clinically meaningful) | Effect < MID (clinically trivial) | |
|---|---|---|
| Statistically significant | Positive finding. The effect is real and large enough to matter. This is the ideal outcome of a well-powered confirmatory trial. | Significant but trivial. The effect is real but too small to be clinically relevant. Common in large studies. Statistical significance here is uninformative; the CI should be compared to the MID. |
| Not statistically significant | Inconclusive. The point estimate is meaningful but the study was probably too small. A wider CI spans both clinically important and trivial values. Larger studies or meta-analysis are needed. | Genuinely null. If the CI is narrow enough to exclude the MID, the study has provided genuine evidence of absence: the effect, if it exists, is too small to matter. This is not "absence of evidence" — it is evidence of absence of a clinically important effect. |
The bottom-right cell deserves emphasis. It is widely misunderstood: a non-significant result is routinely dismissed as "inconclusive" when it is in fact highly informative. If a large, well-powered trial finds $p = 0.40$ and a 95% CI of −1.2 to +1.8 mmHg — entirely below any plausible MID for blood pressure — this is strong evidence that the treatment has no clinically meaningful effect. The correct interpretation is not "we failed to prove an effect" but "we have ruled out a clinically important effect."
Throughout all of this, confidence intervals are the right tool for communication. They report the effect size alongside its uncertainty, allow direct comparison with the MID, and make transparent whether a significant result is also meaningful and whether a non-significant result is informative or merely underpowered. Reporting $p$-values alone — without effect sizes and CIs — is a practice that clinical journals have been discouraging for decades, for good reason.