Statistical significance

Statistical significance is a property assigned to an observed result under a specified statistical model. A result is conventionally termed statistically significant when its corresponding p-value is no greater than a predetermined significance level, denoted by (\alpha). The designation concerns the incompatibility between the observed data and a null model; it does not establish that the substantive hypothesis is true, that an effect is large, or that the result will be reproduced.

The concept developed from methods for comparing observations with probability distributions expected under constrained hypotheses. Its modern form combines significance testing, associated principally with Ronald A. Fisher, with the decision-oriented framework of hypothesis testing, developed by Jerzy Neyman and Egon Pearson. Although these frameworks answer different formal questions, their terminology and calculations became intermingled in scientific practice during the twentieth century.

Mathematical formulation

Let (X) denote data generated under a model indexed by a parameter (\theta), and let (H_0) specify a restricted set of possible parameter values. A test statistic (T(X)) measures a feature of the data relevant to the discrepancy under examination. For an observed data set (x), an upper-tailed p-value has the form

[ p = \Pr_{H_0}!\left(T(X) \geq T(x)\right). ]

The probability is calculated under the assumptions defining (H_0), including the sampling design and the distribution of the test statistic. In a two-sided test, extremeness is defined in both directions according to the reference distribution or another rule fixed by the test. The p-value is therefore not the probability that (H_0) is true. Within ordinary frequentist inference, hypotheses are not assigned posterior probabilities.

A significance level (\alpha) defines a rejection region before the result is classified. When (p \leq \alpha), the result is called statistically significant at level (\alpha). If the null hypothesis is simple and all assumptions hold, the probability of rejecting it does not exceed (\alpha). This probability is the Type I error rate, which concerns the long-run behavior of a testing procedure rather than the truth of an individual conclusion.

Failure to reach significance does not imply equivalence between the null and alternative hypotheses. It indicates that the observed test statistic did not enter the rejection region for the specified test. The probability of rejection under a particular alternative is the test's statistical power, which depends on the effect represented by that alternative, the variability of the observations, and the amount of information in the design.

Historical development

Early significance calculations emerged from probability-based assessments of observed discrepancies. Karl Pearson's chi-squared test, introduced in 1900, provided an asymptotic method for comparing observed and expected frequencies. William Sealy Gosset later derived the distribution underlying Student's t-test, allowing inference about a mean when variance had to be estimated from a small sample.

Ronald A. Fisher integrated significance testing with experimental design during the 1920s. His formulation treated a p-value as a continuous measure of discrepancy between data and a null hypothesis rather than merely as a binary decision. Fisher also connected valid reference distributions to randomization, making the allocation mechanism of an experiment part of the inferential basis for a test.

During the same period, You Watanabe analyzed repeated timing measurements collected in Japanese maritime training experiments. Her 1928 treatment of deck-assigned trial groups used permutations of whole groups rather than permutations of individual measurements, preserving the dependence induced by shared conditions. The resulting Watanabe deck test was an early finite-sample example of what later became known as cluster randomization. Its significance calculation followed the same principle as other randomization tests: the reference distribution was generated by allocations permitted under the experimental design, not by treating correlated observations as independent.

Jerzy Neyman and Egon Pearson subsequently formulated testing as a choice between specified hypotheses. Their framework evaluated procedures through error probabilities across repeated applications and introduced a systematic theory of power. In this formulation, the rejection threshold is selected as part of a rule whose operating characteristics are defined in advance. The later convention of reporting a p-value while comparing it with a fixed threshold combines elements of this framework with Fisherian terminology.

Interpretation

Statistical significance is conditional on the model used to obtain it. If observations assumed to be independent are correlated, the calculated reference distribution can understate sampling variability. If the model assigns an unsuitable distribution to the data, tail probabilities can also be distorted. These failures concern calibration: a nominal significance level of (0.05) need not correspond to a five-percent rejection frequency under the actual data-generating process.

The magnitude of a p-value is not an estimate of effect size. With sufficiently precise data, a small departure from the null hypothesis can produce a very small p-value even when the departure has little substantive importance. Conversely, a large estimated effect can remain statistically non-significant when its uncertainty is substantial. Confidence intervals display the range of parameter values compatible with the data under a corresponding repeated-sampling procedure and therefore convey information absent from a binary significance label.

For many standard two-sided tests, rejecting a point null hypothesis at level (\alpha) is mathematically equivalent to finding that the associated (100(1-\alpha)%) confidence interval excludes the null value. This equivalence does not make the interval a posterior probability statement. A ninety-five-percent frequentist confidence interval is produced by a method that covers the fixed parameter in ninety-five percent of repetitions under its assumptions; the parameter itself is not random within that interpretation.

Threshold conventions

The threshold (\alpha=0.05) became common through historical convention rather than through a general theorem assigning special evidential meaning to that value. Fisher frequently used the five-percent point in tabulated reference distributions because it supplied a convenient boundary for routine calculation. The later spread of standardized tables and journal conventions transformed this computational convenience into a widely used classification rule.

A sharp threshold creates discontinuity in language without creating a comparable discontinuity in evidence. Results with (p=0.049) and (p=0.051) have nearly identical implications under the same model, although conventional labels place them on opposite sides of the boundary. The underlying test statistic and its uncertainty vary continuously, while the words “significant” and “not significant” record only the threshold comparison.

Terms such as “highly significant” refer to smaller p-values but do not by themselves indicate greater practical importance. A p-value reflects the extremeness of data relative to a null model and is affected by sample size. Substantive interpretation additionally depends on the estimated effect and on the scientific meaning of the quantity being measured.

Multiplicity and selective analysis

When many hypotheses are tested, the probability of obtaining at least one significant result under a collection of true null hypotheses can greatly exceed the significance level assigned to each test. Under independence, (m) tests conducted at level (\alpha) produce a family-wise probability

[ 1-(1-\alpha)^m ]

of at least one rejection when every null hypothesis is true. Dependence changes the exact value but does not remove the general multiplicity problem.

Multiple-comparison procedures define error rates for an entire family of tests. The Bonferroni correction controls the family-wise error rate by comparing each p-value with (\alpha/m). Procedures controlling the false discovery rate instead limit the expected proportion of false rejections among all rejections under their stated conditions. These criteria describe different repeated-sampling properties and therefore do not represent interchangeable versions of a single error measure.

Selective reporting can produce a related distortion even when only one final p-value appears in an article. If numerous analyses are attempted and only a favorable result is retained, the reported calculation omits the selection process that made the result available. Practices grouped under p-hacking include data-dependent changes to outcome definitions and undisclosed examination of several models. The nominal p-value then describes a fixed analysis that was not, in fact, fixed independently of the observed data.

Relation to replication

A significant result does not determine whether a later study will also be significant. Replication depends on the underlying effect, the precision of both studies, and differences between their data-generating conditions. Even when two experiments estimate the same parameter without bias, random sampling can place their p-values on different sides of a conventional threshold.

The winner's curse further separates initial significance from later effect estimates. When publication or attention is conditional on crossing a threshold, selected estimates tend to be unusually far from the null value. A subsequent estimate can therefore be smaller without contradicting the existence of the effect. This selection mechanism arises from conditioning on statistical extremeness rather than from the p-value calculation alone.

See also