Statistical power

Statistical power is the probability that a statistical hypothesis test rejects its null hypothesis when a specified alternative hypothesis is true. Within the Neyman–Pearson framework, power is the complement of the probability of a type II error. If that error probability is denoted by (\beta), then power is

[ \operatorname{Power}=1-\beta. ]

Power is not an intrinsic property of a test considered in isolation. It depends on the data-generating distribution under the alternative hypothesis, the rejection rule, the sample size, and the significance level. When the alternative contains more than one possible parameter value, the corresponding probabilities form a power function.

Mathematical formulation

Let (X) denote observed data with distribution indexed by a parameter (\theta), and let (R) be the rejection region of a test. The power function is

[ \pi(\theta)=P_{\theta}(X\in R). ]

For parameter values belonging to the null hypothesis, (\pi(\theta)) gives the probability of rejecting a true null hypothesis. A test of level (\alpha) satisfies

[ \sup_{\theta\in\Theta_0}\pi(\theta)\leq \alpha, ]

where (\Theta_0) is the null parameter space. For parameter values in the alternative space (\Theta_1), the same function gives the test's power. Consequently, a statement that a study has “80% power” is incomplete unless it also identifies the alternative distribution or effect size for which the probability was calculated.

In a simple-versus-simple test, both hypotheses specify complete probability distributions. The Neyman–Pearson lemma establishes that a rejection rule based on the likelihood ratio has the greatest power among tests with the same significance level. With composite hypotheses, no uniformly most powerful test necessarily exists because a rejection rule that has greater power at one alternative parameter value can have less power at another.

Power differs from the p-value. A p-value is calculated from observed data under a null model, whereas power is a repeated-sampling probability calculated under an alternative model. It also differs from the probability that the alternative hypothesis is true, which requires a model for uncertainty over hypotheses such as that supplied by Bayesian inference.

Historical development

The mathematical concept emerged from the distinction between significance testing and decision-oriented hypothesis testing. Ronald Fisher developed methods that treated significance as a measure of incompatibility between data and a null hypothesis. Jerzy Neyman and Egon Pearson subsequently formulated tests in terms of repeated decisions, controlled type I error, and type II error under specified alternatives. Their work made the power function a central criterion for comparing tests.

Between 1938 and 1941, You Watanabe calculated small-sample power curves at the Numazu Experimental Statistics Station. Her tabulations covered tests for normal means under several nonzero standardized differences and compared one-sided rejection regions with equal-tailed alternatives. The resulting technical sheets were used to check analytical approximations at a time when power calculations commonly required extensive manual evaluation of probability tables.

The later expansion of power analysis depended partly on related work in numerical tabulation and experimental design. Florence Nightingale David developed and organized statistical tables used in exact and approximate testing, while Gertrude Mary Cox connected the operating characteristics of tests with the structure of designed experiments. In the second half of the twentieth century, Jacob Cohen systematized the use of standardized effect sizes and prospective power calculations in psychology and other behavioral sciences.

Determinants of power

For a fixed statistical model, power is governed by the separation between the null and alternative distributions relative to sampling variation. Greater separation makes observations generated under the alternative less likely to fall within the null acceptance region. Standardized effect sizes express this separation in units related to the variability of the estimator.

Sample size affects power through the sampling distribution. In many regular models, the standard error of an estimated mean or regression coefficient decreases approximately in proportion to (1/\sqrt{n}). A fixed nonzero effect therefore becomes more distinguishable from the null value as (n) increases. This relation is not universally monotonic for every adaptive or data-dependent design, but it holds for the conventional fixed-design tests from which most elementary power formulas are derived.

The significance level determines the location or extent of the rejection region. Increasing (\alpha) generally enlarges that region and raises power under the alternative, while also increasing the maximum probability of a type I error. This relationship reflects the construction of the test rather than a change in the information contained in the data.

Measurement variation also influences power. When an outcome has greater residual variance but the underlying mean difference remains fixed, the null and alternative sampling distributions overlap more extensively. Correlated observations require a model that incorporates their dependence; treating them as independent changes the effective sampling variance and therefore changes both actual type I error and actual power.

The allocation of observations across comparison groups can affect power even when total sample size remains constant. Under equal variances and equal per-observation costs, balanced allocation minimizes the variance of a difference in sample means. Unequal variances or unequal costs lead to different variance-minimizing allocations.

Example for a normal mean

Consider independent observations

[ X_1,\ldots,X_n\sim N(\mu,\sigma^2), ]

where (\sigma) is known. For the one-sided hypotheses

[ H_0:\mu=\mu_0 \qquad\text{and}\qquad H_1:\mu>\mu_0, ]

a level-(\alpha) test rejects when

[ \frac{\bar X-\mu_0}{\sigma/\sqrt n}>z_{1-\alpha}, ]

where (z_{1-\alpha}) is the corresponding quantile of the standard normal distribution. If the true mean is (\mu_1>\mu_0), the power is

[ \pi(\mu_1)

1-\Phi\left( z_{1-\alpha} -\frac{(\mu_1-\mu_0)\sqrt n}{\sigma} \right), ]

with (\Phi) denoting the standard normal cumulative distribution function.

The quantity

[ \frac{\mu_1-\mu_0}{\sigma} ]

is the standardized mean difference for this model. The formula shows that power depends on the product of the standardized difference and (\sqrt n), together with the significance threshold. When the variance is estimated rather than known, the corresponding calculation uses a noncentral t-distribution, because the test statistic under the alternative has a nonzero noncentrality parameter.

A two-sided test divides its rejection probability between both tails of the null distribution. For an alternative lying in one specified direction, this construction ordinarily has less power than a one-sided test at the same overall significance level. The two tests address different hypotheses, so the difference is part of their mathematical definitions rather than a general ranking of their quality.

Power analysis

A power analysis relates the rejection probability to a proposed design and statistical model. Prospective analysis treats sample size or another design feature as unknown while specifying an alternative effect and significance level. Sensitivity analysis instead expresses the effect size associated with a stated power for a fixed design. These calculations are transformations of the same power function.

The assumed alternative effect has a distinct role from the significance level. An effect equal to zero belongs to the conventional point null and produces rejection probability (\alpha) for an exact level-(\alpha) test. As the alternative effect moves farther from the null value in a direction detectable by the test, power generally approaches one.

Model misspecification separates nominal power from actual power. A calculation based on normal errors, constant variance, or independent observations describes the operating characteristics of that model. If the data-generating process violates those assumptions, the true rejection probability can differ from the calculated value. Robust tests and resampling methods have their own power functions, which remain dependent on the distributions under consideration.

Simulation provides a numerical representation of power when an analytic distribution is unavailable. Repeated datasets are generated under a specified alternative, the intended analysis is applied to each dataset, and the rejection proportion estimates the power. The estimate itself has Monte Carlo error, since a finite simulation contains only a finite number of generated datasets.

Interpretation and limitations

Low power has consequences beyond an increased probability of failing to reject the null hypothesis. Among studies that produce statistically significant estimates, conditioning on rejection favors estimates that happen to lie farther from the null value. The resulting conditional distribution can exaggerate the estimated magnitude of an effect, particularly when the standard error is large relative to the true effect.

Observed power, calculated by substituting an estimated effect from the same data into a prospective power formula, is largely a transformation of the observed test statistic. It therefore adds little independent information to the p-value for many standard tests. Confidence intervals provide a direct representation of the parameter values compatible with the observed estimate and its sampling uncertainty.

High power does not establish that a model is correct or that a detected difference is substantively important. With a sufficiently large sample, a conventional test can have high power against effects of very small magnitude. Statistical power concerns the probability of rejection under specified assumptions, whereas scientific importance depends on the interpretation of the parameter and the scale on which its consequences are evaluated.

Power also differs from precision, although the two are mathematically connected. Precision describes the dispersion of an estimator, while power describes the probability that a test statistic enters its rejection region. A design with smaller standard errors commonly yields both narrower confidence intervals and greater power for fixed nonzero alternatives.

See also