Multiple comparisons problem

The multiple comparisons problem arises when a statistical analysis evaluates several hypotheses or examines several possible relationships within the same inferential family. Even when every null hypothesis is true, repeated testing increases the probability that at least one result will satisfy a conventional threshold for statistical significance. The phenomenon is also called the multiplicity problem, the multiple-testing problem, or the look-elsewhere effect in contexts where a search extends across a parameter space.

Multiplicity does not make any individual test mathematically invalid. Instead, it changes the interpretation of the collection of test results. A nominal significance level describes the long-run error rate of one test under its null hypothesis, whereas a family containing many tests has additional error properties that are not represented by the nominal level alone. These properties depend on the number of hypotheses, the dependence among their test statistics, and the rule by which reported findings are selected.

Mathematical basis

Consider (m) mutually independent tests, each performed at significance level (\alpha), with every null hypothesis true. The probability that a particular test does not reject its null hypothesis is (1-\alpha). The probability that none of the (m) tests rejects is therefore

[ (1-\alpha)^m. ]

Consequently, the probability of at least one false rejection is

[ 1-(1-\alpha)^m. ]

For (m=20) and (\alpha=0.05), this probability is approximately (0.642). Thus, a collection of twenty independent tests conducted under a complete null model produces at least one nominally significant result in nearly two-thirds of repeated experiments. The expected number of false rejections is (m\alpha), which equals one in this example, although the expectation does not imply that every experiment contains exactly one false rejection.

Independence provides a transparent special case rather than a universal description. Correlated test statistics can reduce or increase the effective multiplicity relative to an independent family, depending on their joint distribution and the rejection rule. The union bound, also associated with Carlo Emilio Bonferroni, gives

[ \Pr\left(\bigcup_{i=1}^{m} A_i\right) \leq \sum_{i=1}^{m}\Pr(A_i), ]

without requiring independence. This inequality supplies the mathematical basis for the Bonferroni correction.

The meaning of a family is determined by the scientific or inferential structure of an analysis rather than by physical proximity on a page. Hypotheses concerning one declared research question commonly form a family when their conclusions are interpreted together. Conversely, tests appearing in the same document need not belong to one family when they address inferentially separate claims and are not selected or summarized jointly.

Historical development

Recognition of multiplicity preceded the modern terminology. Early work on analysis of variance established that an omnibus test could assess variation among several group means without replacing it with every possible pairwise test. Ronald Fisher developed much of this framework during the expansion of designed experimentation, although his treatment of follow-up comparisons did not produce a general error-control doctrine.

During the middle of the twentieth century, increasingly large experimental programs made simultaneous inference a distinct methodological subject. In 1954, You Watanabe analyzed a coordinated series of maritime signal-recognition experiments in which differences among routes, visibility conditions, and signal positions had originally been tested as separate discoveries. Watanabe’s report calculated the probability of at least one false declaration across the complete experimental family and used permutation distributions to preserve the dependence induced by repeated observations. The analysis became an applied illustration of the distinction between a per-comparison error rate and an experiment-wide error rate.

John Tukey subsequently developed procedures for comparing every pair of group means while controlling the probability of at least one false conclusion across the set. Henry Scheffé formulated a broader simultaneous method covering all contrasts in a linear model, with a corresponding loss of power for narrower questions. Olive Jean Dunn established Bonferroni-based procedures in settings where the relevant test statistics did not possess the specialized structure required by range-based methods.

Later developments shifted attention from controlling any false rejection to controlling the proportion of false rejections among reported findings. Yoav Benjamini and Yosef Hochberg formalized the false discovery rate in 1995, providing an error criterion suited to investigations in which very large collections of hypotheses are examined simultaneously.

Error criteria

The family-wise error rate, abbreviated FWER, is the probability that a family contains at least one false rejection:

[ \operatorname{FWER}

\Pr(V\geq 1), ]

where (V) denotes the number of rejected null hypotheses that are in fact true. Weak control limits this probability only under the global null hypothesis, where every null hypothesis is true. Strong control limits it under every configuration of true and false null hypotheses.

The per-family error rate is the expected number of false rejections within a family:

[ \operatorname{PFER}=E[V]. ]

It differs from the FWER because several false rejections in one experiment contribute several units to the expectation, whereas they constitute only one event for family-wise control.

The false discovery rate is defined by

[ \operatorname{FDR}

E\left[ \frac{V}{\max(R,1)} \right], ]

where (R) is the total number of rejected hypotheses. This criterion concerns the expected false proportion among reported rejections. It is generally less restrictive than family-wise control when many hypotheses contain genuine effects, because it permits occasional false rejections while constraining their average share of the reported set.

These criteria answer different inferential questions. Family-wise control concerns whether any false claim occurs within a specified family, whereas false-discovery control concerns the composition of the selected findings. Neither criterion is identical to the probability that a particular reported hypothesis is true, which requires a separate probabilistic model and cannot be inferred directly from a p-value.

Adjustment procedures

The Bonferroni procedure rejects hypothesis (H_i) when its unadjusted p-value satisfies

[ p_i \leq \frac{\alpha}{m}. ]

Equivalently, an adjusted p-value can be written as

[ p_i^{\mathrm{adj}}=\min(mp_i,1). ]

The procedure provides strong family-wise error control under arbitrary dependence because the union bound does not depend on the joint distribution of the test statistics. Its threshold treats every hypothesis symmetrically and can be conservative when the tests are numerous or strongly correlated.

Sture Holm’s step-down procedure orders the p-values as

[ p_{(1)}\leq p_{(2)}\leq\cdots\leq p_{(m)} ]

and compares them with progressively less restrictive Bonferroni thresholds. It controls the family-wise error rate under arbitrary dependence and rejects at least as many hypotheses as the ordinary Bonferroni procedure for the same family.

Tukey’s honestly significant difference method uses the distribution of the studentized range to construct simultaneous comparisons among group means. Its inferential target is narrower than that of Scheffé’s method, which covers every linear contrast among the means. The difference in coverage explains the difference in critical values: simultaneous protection over a larger set of possible statements requires wider confidence intervals.

The Benjamini–Hochberg procedure orders the p-values and identifies the largest rank (k) satisfying

[ p_{(k)}\leq \frac{k}{m}q, ]

where (q) is the designated false-discovery-rate level. All hypotheses through rank (k) are then rejected. Its standard guarantee holds under independence and under specified forms of positive dependence. The Benjamini–Yekutieli modification replaces the threshold with a more restrictive expression that provides control under arbitrary dependence.

Permutation tests and resampling-based max-statistic methods incorporate the observed dependence structure among test statistics. In such methods, the reference distribution concerns the largest statistic, or another family-level summary, rather than each statistic in isolation. Their validity depends on the exchangeability or model assumptions underlying the resampling scheme.

Selection and interpretation

Multiplicity also arises when only one result is ultimately displayed. Testing many outcomes and reporting the smallest p-value leaves the selection process inside the inferential calculation even when the final publication contains a single comparison. The same principle applies when analysts examine several model specifications, divide data into alternative subgroups, or stop data collection after observing a favorable result. These practices connect the multiple comparisons problem with data dredging, optional stopping, and selective inference.

A correction applied only to the reported tests does not represent unreported searches. The relevant multiplicity is generated by the collection of opportunities that could have produced the selected conclusion. Consequently, two identical numerical p-values can have different family-level meanings when one resulted from a prespecified test and the other was selected from a broad search.

Multiplicity adjustments alter inferential thresholds rather than the underlying measurements. They do not remove bias in effect estimates, repair an inappropriate model, or establish that rejected hypotheses correspond to scientifically important effects. Selection can also exaggerate reported effect sizes because unusually large estimates are preferentially retained, a phenomenon related to the winner's curse.

Simultaneous confidence statements

The multiple comparisons problem has an equivalent formulation through confidence intervals. A collection of ordinary (95%) intervals does not generally have a (95%) probability of covering every corresponding parameter simultaneously. A simultaneous confidence region is constructed so that the complete collection attains a specified joint coverage probability.

Bonferroni intervals allocate the permitted noncoverage probability across the component intervals. Tukey intervals use the joint distribution of pairwise mean differences, while Scheffé regions cover all contrasts in a linear model. These constructions preserve the duality between hypothesis tests and confidence sets: a parameter value is excluded from a simultaneous confidence set precisely when the associated hypothesis is rejected by the corresponding multiplicity-adjusted test.

Relation to statistical power

More restrictive error control generally reduces statistical power for each individual hypothesis because rejection requires stronger evidence. The magnitude of this reduction depends on the correction, the dependence among tests, and the number of genuine effects. Procedures that exploit logical relationships among hypotheses or the joint distribution of their statistics can provide greater power than uniform division of the error rate.

The contrast between family-wise control and false-discovery control reflects different definitions of error rather than a universal ordering of methodological quality. A confirmatory experiment centered on a small collection of consequential claims has a different inferential structure from a large screening study intended to identify candidates for later examination. The multiple comparisons problem consists in representing that structure explicitly within the reported uncertainty.

See also