Sample Size and Statistical Power
Every clinical study begins with a question, and the quality of the answer depends heavily on a decision that must be made before a single patient is enrolled: how large does the study need to be? Too few participants and the study cannot reliably detect the effect it is looking for, even if that effect genuinely exists. Too many participants and resources are wasted, the study runs longer than necessary, and — since participants accept some degree of risk or inconvenience — more people are exposed to the experimental condition than strictly needed. Getting the sample size right is therefore both a scientific and an ethical obligation.
Calculating the required sample size is not a mechanical exercise. It requires the investigator to make explicit, upfront commitments about what they hope to detect, how certain they need to be, and how much risk of error they are willing to accept. This page explains the conceptual framework behind those decisions.
The Two Types of Error
Every hypothesis test is a decision made under uncertainty. The test examines sample data and returns one of two verdicts: either the null hypothesis is rejected (the result is declared statistically significant) or it is not rejected (no significant difference is found). Since neither the investigator nor the test knows the unknowable truth — whether the null hypothesis is actually true or false — two kinds of mistake are possible.
| H₀ is TRUE (no real effect exists) |
H₀ is FALSE (a real effect exists) |
|
|---|---|---|
| Test is significant (H₀ rejected) |
Type I errorFalse positive Probability = α (significance level) |
Correct decisionTrue positive Probability = 1 − β = Power |
| Test is not significant (H₀ not rejected) |
Correct decisionTrue negative Probability = 1 − α |
Type II errorFalse negative Probability = β |
Type I Error — the False Positive
A Type I error occurs when the null hypothesis is true — there is genuinely no effect — but the statistical test returns a significant result anyway. The investigator concludes that a difference exists when in reality it does not. This is a false alarm.
The probability of committing a Type I error is denoted α and is set directly by the investigator through the choice of significance level. If a significance level of 5% ($p < 0.05$) is adopted, then in a world where the null hypothesis is true, 1 in 20 tests will produce a spuriously significant result simply through sampling variation. Choosing a stricter significance level — say $p < 0.01$ — reduces α to 1 in 100, making false positives rarer but harder to achieve. The $p = 0.05$ threshold has no special scientific status; it is a historical convention, and its limitations are discussed in the Clinical versus statistical significance chapter.
Type II Error — the False Negative
A Type II error occurs when the null hypothesis is false — a real effect genuinely exists — but the test fails to detect it. The investigator concludes that no significant difference was found when in reality there is one. This is a missed discovery.
The probability of a Type II error is denoted β. A study with β = 0.20 will fail to detect a real effect 20% of the time, even when it is truly present. Unlike α, β is not directly set by the investigator; it depends on α, the sample size, and the size of the effect being sought. The only ways to reduce β — to make false negatives rarer — are to increase the sample size, to choose a larger target effect size, or (with some cost) to relax the significance threshold.
There is a deliberate asymmetry in the conventional thresholds for these two error types. The Type I error rate is typically set at 5%, while the Type II error rate is accepted at up to 20%. This reflects a judgement — not always appropriate for every field — that falsely claiming an effect exists (Type I) is more damaging than failing to detect a real effect (Type II). In diagnostic medicine the balance is often reversed: failing to diagnose a serious disease is more harmful than a false alarm that leads to further investigation.
Statistical Power
Statistical power is the probability that a study will correctly detect an effect that genuinely exists. It is the complement of the Type II error rate: power = 1 − β. A study with 80% power will detect a true effect 4 times out of 5 and miss it 1 time in 5.
Power is the central concept in sample size planning. A study that is underpowered is not merely a waste of resources; it is arguably unethical, because it subjects participants to risks and inconveniences for a study that was never realistically capable of answering the question it set out to address.
The convention of 80% power is widely accepted in medical research, meaning investigators are willing to accept a 20% chance of missing a real effect. For confirmatory trials or for decisions with serious clinical consequences — approving a new drug, changing first-line treatment — 90% power (β = 0.10) is often considered more appropriate, accepting that the extra precision requires a considerably larger sample.
An important subtlety: power is not a fixed property of a study design. It depends on the true size of the effect, which is unknown in advance. A study designed with 80% power to detect a difference of 10 mmHg in blood pressure will have higher power if the true difference turns out to be 15 mmHg, and lower power if it is only 6 mmHg. The power calculation specifies how much power the study will have to detect a effect of a size — not power to detect any and all effects.
What Determines the Required Sample Size?
Four inputs are needed to calculate a required sample size. The sample size grows when any of the first three increase, or when the fourth decreases.
The Significance Level (α)
A stricter significance threshold — a lower α — means the test demands more evidence before declaring a result significant. To achieve that stricter standard, a larger sample is needed. Setting α at 1% rather than 5% roughly doubles the required sample in many settings.
The Desired Power (1 − β)
Higher desired power means a lower tolerance for missing a real effect. Increasing power from 80% to 90% typically requires increasing the sample size by roughly a third. The relationship is non-linear: pushing from 90% to 95% requires a further substantial increase.
The Minimum Clinically Important Difference
This is the most consequential input and, in practice, the hardest to specify correctly. The investigator must decide: what is the effect that would have genuine clinical or practical importance — the threshold below which the finding would not change practice, prescribing, or policy? The temptation to specify a very small difference in order to justify a large study should be resisted; the clinically important difference must be set on clinical, not statistical, grounds. How this threshold is defined — anchor-based methods, distribution-based methods, and expert consensus — is covered in the Minimal important difference section of the Clinical versus statistical significance chapter.
Variability
For continuous outcomes, the variability of the measurements — typically expressed as a standard deviation — directly affects the required sample size. Greater variability means individual measurements are more scattered around their true mean, making it harder to see a real difference through the noise. For binary outcomes, the baseline event rate in the control group plays a similar role.
Estimates of variability are usually taken from published literature, pilot studies, or expert opinion. When the estimate turns out to be wrong — variability higher than assumed — the study will be underpowered; when it is lower, the study will exceed its target power. This is one reason why interim analyses and adaptive designs, which allow sample size re-estimation based on accumulating data, have grown in popularity.
One-Tailed versus Two-Tailed Tests
Most sample size calculations assume a two-tailed test: the investigator wants to detect an effect in either direction (the new treatment could be better worse than control). This is the correct choice in the vast majority of clinical research, where an unexpected harmful effect is just as important to detect as a beneficial one.
A one-tailed test is appropriate only when there is a strong, pre-specified scientific reason to believe the effect can only go in one direction, and where a difference in the opposite direction would be treated identically to no difference at all. Such situations are rare in clinical medicine. Using a one-tailed test to reduce the required sample size when a two-tailed test is really needed is a form of bias.
The common mistake of calculating power retrospectively from the study's own data — and why it is scientifically indefensible — is covered in the Post-hoc power analysis chapter.
Sample Size Calculators in MedCalc
MedCalc provides dedicated sample size calculators for the most common study designs. Each calculator takes the significance level, desired power, and the appropriate measure of effect size as inputs, and returns the required sample size per group or in total.
- Sample size: Single mean — for estimating a population mean or testing it against a reference value.
- Sample size: Single proportion — for estimating a prevalence or event rate.
- Sample size: Comparison of two means — for a two-group parallel trial with a continuous outcome.
- Sample size: Paired samples t-test — for before–after or matched-pairs designs.
- Sample size: Comparison of two proportions — for a two-group trial with a binary outcome.
- Sample size: McNemar test — for paired binary outcomes (e.g. crossover designs).
- Sample size: Correlation coefficient — for detecting a specified correlation between two continuous variables.
- Sample size: Survival analysis (log-rank test) — for time-to-event outcomes comparing two survival curves.
- Sample size: Bland–Altman plot — for method comparison studies assessing limits of agreement.
- Sample size: Area under ROC curve — for diagnostic accuracy studies evaluating a single test.
- Sample size: Comparison of two ROC curves — for studies comparing the accuracy of two diagnostic tests.
Note that sample size calculation for confidence interval width — rather than hypothesis testing power — is covered separately; see for example the Sample size for a confidence interval of required width, which was also discussed in the Univariate statistics chapter.