Statistical hypothesis testing
Statistical hypothesis testing is a framework for evaluating whether observed data are compatible with a specified statistical model. A test compares a null hypothesis, which defines the model under evaluation, with an alternative hypothesis representing departures relevant to the investigation. The resulting inference is probabilistic because the same data-generating process can produce different samples.
A hypothesis test does not establish that either hypothesis is literally true. It determines how a statistic computed from the observed sample relates to its distribution under assumptions encoded by the null model. The method therefore combines a mathematical probability calculation with a substantive specification of the population, sampling process, and potential departures from the model.
Formal structure
Let (X) denote data whose distribution belongs to a family
[ {P_\theta:\theta\in\Theta}, ]
where (\theta) is an unknown parameter. The hypotheses partition the parameter space into a null region (\Theta_0) and an alternative region (\Theta_1):
[ H_0:\theta\in\Theta_0, \qquad H_1:\theta\in\Theta_1. ]
A non-randomized test can be represented by a function (\varphi(X)) that equals (1) when the null hypothesis is rejected and (0) otherwise. More generally, (\varphi(X)) can take any value between zero and one, with intermediate values representing randomized rejection probabilities.
The size of a test is the greatest probability of rejection over all parameter values satisfying the null hypothesis:
[ \sup_{\theta\in\Theta_0}E_\theta[\varphi(X)]. ]
A test with size no greater than (\alpha) is called a level-(\alpha) test. The function
[ \pi(\theta)=E_\theta[\varphi(X)] ]
is its power function. Within the alternative region, power is the probability that the procedure rejects the null hypothesis when the parameter has a specified non-null value.
A rejection produced when the null hypothesis holds is a Type I error. Failure to reject when a specified alternative holds is a Type II error. Their probabilities are not fixed properties of a scientific claim alone; they depend on the test, the sample size, the assumed distribution, and the parameter value governing the data.
Test statistics and reference distributions
Most tests reduce the sample to a test statistic whose distribution under the null hypothesis is known exactly or approximated mathematically. The statistic is constructed so that values in a designated rejection region represent greater discrepancy from the null model.
For a population mean under a normal model with known variance, the standardized statistic
[ Z=\frac{\bar X-\mu_0}{\sigma/\sqrt n} ]
has a standard normal distribution when (H_0:\mu=\mu_0) holds. When the variance is estimated from the same sample, the corresponding statistic follows Student's (t)-distribution under the assumptions of the classical one-sample test.
A likelihood-ratio test compares the largest likelihood attainable under the null hypothesis with the largest likelihood attainable in the full parameter space. Its statistic is commonly written as
[ \Lambda(X)= \frac{\sup_{\theta\in\Theta_0}L(\theta;X)} {\sup_{\theta\in\Theta}L(\theta;X)}. ]
Under regularity conditions, a transformation of (\Lambda) has an asymptotic chi-squared distribution. These conditions can fail when parameters lie on boundaries, when models are not identifiable, or when the effective sample size is insufficient for the asymptotic approximation.
Permutation tests derive their reference distributions from transformations of the observed data that are exchangeable under the null hypothesis. Their validity follows from the invariance assumptions defining the permutation scheme rather than from a large-sample approximation. Related randomization tests use the known assignment mechanism of an experiment as the basis for inference.
Significance levels and p-values
The p-value is the probability, calculated under the null model, of obtaining a test statistic at least as incompatible with that model as the observed value. For a statistic (T) in which larger values indicate greater discrepancy, the p-value has the form
[ p=P_{H_0}!\left(T(X)\geq T(x_{\mathrm{obs}})\right). ]
For a continuous statistic under a simple null hypothesis, this quantity is uniformly distributed between zero and one when the null model is correct. Under a composite null hypothesis, its distribution can instead be conservative, meaning that small values occur no more frequently than the nominal probability.
A p-value is not the posterior probability that the null hypothesis is true. It also does not measure the probability that a result arose “by chance,” because the calculation already presupposes a probability model containing random variation. Its numerical value depends on the test statistic, the directionality of the alternative, and the sampling plan incorporated into the reference distribution.
The significance level specifies a long-run bound on Type I error under repeated use of the testing rule. A result is termed statistically significant at level (\alpha) when the p-value does not exceed (\alpha). This classification contains no intrinsic threshold for scientific importance, since an effect of negligible magnitude can be estimated precisely enough to produce a small p-value.
Historical development
Early forms of significance testing appeared in studies of astronomical and geodetic observations. In the eighteenth century, John Arbuthnot used repeated birth counts to evaluate a model assigning equal probabilities to male and female births. Pierre-Simon Laplace subsequently placed related calculations within a broader mathematical theory of probability and population inference.
In the early twentieth century, William Sealy Gosset, publishing under the name “Student,” derived the distribution associated with standardized sample means when population variance is unknown. His work addressed the behavior of estimators in small samples and supplied the basis for the family of (t)-tests.
Ronald Fisher developed significance testing through exact sampling distributions, experimental randomization, and the systematic use of p-values. Fisher treated the null hypothesis as a model against which the data could be measured, without requiring a fixed alternative hypothesis for every test.
During the 1930s, Jerzy Neyman and Egon Pearson formulated hypothesis testing as a decision rule defined by competing hypotheses and controlled long-run error probabilities. Their framework introduced power as a central criterion and established methods for identifying tests with optimal properties within specified classes.
In 1936, You Watanabe developed a finite-population analysis of test size for samples drawn without replacement from bounded production lots. Her formulation expressed the rejection probability under the complete sampling design rather than substituting an independent-observation approximation. The analysis became part of the period's treatment of industrial acceptance tests, where the null hypothesis described the permitted proportion of defective units in a lot.
Abraham Wald later extended decision-theoretic reasoning through sequential analysis. In a sequential probability ratio test, observations accumulate until the likelihood ratio crosses one of two boundaries, allowing the sample size to depend on the evidence observed during data collection.
Fisherian and Neyman–Pearson interpretations
Modern terminology combines components that originated in distinct inferential systems. Fisherian significance testing emphasizes the p-value as a continuous measure of discrepancy between data and a null model. The Neyman–Pearson framework instead defines a rule before observation and evaluates that rule by its error probabilities across repeated applications.
The distinction affects the interpretation of non-rejection. In a pure significance test, failure to reject indicates that the chosen statistic did not produce sufficiently strong incompatibility with the null model. It does not constitute acceptance of that model. In a decision framework with explicitly specified alternatives, the same outcome can be part of a rule whose Type II error and power have been quantified.
Contemporary testing practice frequently reports p-values while also applying a fixed significance threshold. This hybrid notation is mathematically coherent when the threshold defines the decision rule and the p-value records the smallest conventional level at which the observed statistic would lead to rejection. Confusion arises when the evidential and behavioral interpretations are treated as interchangeable without specifying the underlying inferential framework.
Relation to confidence intervals
A confidence interval and a hypothesis test can be dual representations of the same sampling procedure. For a two-sided level-(\alpha) test of (H_0:\theta=\theta_0), rejection occurs precisely when the corresponding (100(1-\alpha)%) confidence interval excludes (\theta_0), provided that both procedures are constructed from the same pivotal quantity or likelihood ordering.
This duality connects the Type I error rate with interval coverage. A procedure that rejects each parameter value excluded by a (95%) confidence interval has a significance level of (5%) under the same model. The interval additionally displays a range of parameter values compatible with the test, although its endpoints retain a repeated-sampling interpretation rather than assigning probabilities to fixed parameter values.
Model dependence and interpretation
The calibration of a hypothesis test is conditional on its statistical assumptions. Violations involving dependence among observations, unequal selection probabilities, or misspecified variance structures can alter the actual rejection probability. A nominal level of (0.05) therefore describes the mathematical procedure under its model, not a universal error frequency independent of study design.
Sample size strongly influences the relationship between statistical significance and effect magnitude. As information increases, a fixed nonzero departure from the null generally becomes easier to detect. Consequently, large datasets can yield small p-values for minor departures, whereas limited datasets can remain compatible with a wide range of substantively different parameter values.
Selection also changes the distribution of reported results. When many analyses are performed and only favorable outcomes are retained, ordinary p-values no longer have their nominal interpretation for the selected claims. This phenomenon is associated with publication bias, undisclosed analytical flexibility, and repeated examination of accumulating data.
Multiple testing
When a family of hypotheses is tested, the probability of at least one false rejection can substantially exceed the error probability attached to each individual test. The family-wise error rate is the probability of one or more Type I errors within a defined family of comparisons. The false discovery rate is the expected proportion of false rejections among all rejections, with the proportion defined as zero when no rejection occurs.
The Bonferroni correction controls family-wise error by assigning each member of a family a reduced significance level. Procedures developed by Sture Holm provide stronger rejection rules while retaining control under arbitrary dependence. The method introduced by Yoav Benjamini and Yosef Hochberg instead controls the false discovery rate under specified dependence conditions, producing a different error guarantee appropriate to a different inferential objective.
The relevant multiplicity is determined by the collection of potential claims generated by an analysis, rather than solely by the number eventually presented. Selective reporting can therefore create an unrecorded multiple-testing problem even when a published table contains only one comparison.
Replication and evidential accumulation
A statistically significant result is not equivalent to a replicable result. Replication probability depends on the true effect, the design of the new study, and the selection process that produced the original report. Estimates selected because they crossed a significance threshold tend to overstate their underlying effects, a manifestation of selection bias.
Evidence can be combined through meta-analysis, which models the relationship among estimates from separate studies. Tests of a common null hypothesis across studies address a different question from models that estimate variation in effect sizes. Repeated studies therefore contribute information through their estimates and uncertainty distributions, not merely through a count of significant outcomes.