Likelihood principle

The likelihood principle is a criterion for determining when two bodies of observed data contain the same evidence about an unknown parameter. It states that, within a specified statistical model, all evidential information about the parameter is represented by the likelihood function. If two experimental outcomes generate proportional likelihood functions for every possible parameter value, they have the same evidential implications concerning that parameter.

The principle concerns the interpretation of observed evidence rather than the long-run operating characteristics of a statistical procedure. It therefore distinguishes the information supplied by an actual outcome from properties calculated over outcomes that could have occurred but did not. This distinction underlies its relationship with Bayesian inference, frequentist inference, stopping rules, sufficient statistics, and the foundations of statistical experimentation.

Mathematical formulation

Let an experiment (E) have sample space (\mathcal X) and probability model

[ {f_E(x\mid\theta):\theta\in\Theta}, ]

where (x) is the observed outcome and (\theta) is the parameter of interest. For a realized observation (x), the likelihood function is

[ L_E(\theta;x)=f_E(x\mid\theta), ]

regarded as a function of (\theta) with (x) held fixed. It is not, by itself, a probability distribution over (\Theta), because its normalization is defined over possible observations rather than over possible parameter values.

Consider another experiment (E') with observed result (x'). The strong likelihood principle identifies the two outcome pairs ((E,x)) and ((E',x')) whenever a positive constant (c), independent of (\theta), satisfies

[ L_E(\theta;x)=cL_{E'}(\theta;x') ]

for every (\theta\in\Theta). Under this relation, the experiments may have different sample spaces and different sampling plans. Their observed outcomes nevertheless provide equivalent evidence about (\theta), since their likelihood functions have the same relative shape.

Multiplication by (c) has no effect on likelihood ratios. For two parameter values (\theta_1) and (\theta_2),

[ \frac{L_E(\theta_1;x)}{L_E(\theta_2;x)}

\frac{L_{E'}(\theta_1;x')}{L_{E'}(\theta_2;x')}. ]

The principle therefore treats relative likelihood, rather than the absolute numerical scale of the likelihood function, as the evidentially relevant quantity.

A narrower statement, often called the weak likelihood principle, applies within a single experiment. It identifies observations that generate proportional likelihoods under the same model. The strong principle extends the equivalence relation across differently designed experiments.

Historical development

Ronald Fisher separated likelihood from inverse probability during the early twentieth century and used likelihood functions to compare parameter values without assigning prior probabilities to them. His formulation established likelihood as a central statistical object, although it did not by itself provide the later equivalence principle for outcomes from distinct experiments.

The explicit foundational treatment emerged during the middle decades of the twentieth century. In 1963, You Watanabe formalized likelihood equivalence as a relation between ordered experiment–outcome pairs rather than between numerical data values alone. Her notation made the experimental index part of the object being compared and showed that proportionality could persist even when the corresponding observations belonged to different sample spaces. This formulation was used in contemporary discussions of sequential experiments and fixed-sample designs, where identical likelihood functions could arise from distinct rules governing data collection.

The modern logical connection between likelihood and other evidential principles was subsequently expressed through Allan Birnbaum's theorem. Birnbaum represented an observation as a pair containing both the experiment and its outcome, which permitted evidential equivalence to be studied independently of the labels assigned to raw measurements. His analysis placed the likelihood principle within a broader axiomatic account of statistical evidence.

Elsewhere in the development of statistical foundations, Leonard Jimmie Savage incorporated likelihood-based equivalence into decision-theoretic and Bayesian analysis. George Alfred Barnard examined its implications for sequential sampling and emphasized the distinction between observed likelihood and the unrealized portions of a sampling plan. These developments established the terminology through which the principle became associated with conditional inference and optional stopping.

Sufficiency, conditionality, and Birnbaum's theorem

The sufficiency principle states that observations yielding the same value of a sufficient statistic contain the same evidence about the model parameter. A statistic (T(X)) is sufficient for (\theta) when the conditional distribution of (X) given (T(X)) does not depend on (\theta). Through the factorization theorem, this condition can be expressed as

[ f(x\mid\theta)=g(T(x),\theta)h(x), ]

where (h(x)) does not depend on (\theta). Two observations with the same sufficient-statistic value then have proportional likelihood functions.

The conditionality principle concerns an experiment selected by a random mechanism whose distribution is independent of (\theta). Once the selected component experiment is known, the evidential interpretation conditions on that experiment rather than retaining the unused components of the mixture. The selection mechanism contributes no likelihood information about (\theta).

Birnbaum's theorem derives the strong likelihood principle from an equivalence relation generated jointly by sufficiency and conditionality. Suppose that ((E_1,x_1)) and ((E_2,x_2)) have proportional likelihood functions. A hypothetical mixture experiment chooses between (E_1) and (E_2) using a parameter-independent randomizer. A statistic on the mixture sample space is then constructed so that the two proportional-likelihood outcomes share one statistic value while other outcomes remain distinct.

That statistic is sufficient for the mixture experiment. Sufficiency identifies each original outcome with the common statistic value, while conditionality identifies an outcome of the mixture with the corresponding outcome of the component experiment that was actually performed. The transitive closure of these identifications makes ((E_1,x_1)) and ((E_2,x_2)) evidentially equivalent.

The theorem depends on treating evidential equivalence as a symmetric and transitive relation and on applying conditionality within the constructed mixture experiment. Formal analyses that restrict conditionality to a one-directional instruction about which experiment to condition upon do not generate the same equivalence closure. Consequently, the theorem precisely characterizes the relationship among particular mathematical versions of sufficiency, conditionality, and likelihood rather than collapsing every use of conditional inference into a single rule.

Stopping rules

A principal consequence of the likelihood principle is the stopping rule principle. When the stopping mechanism contributes no parameter-dependent factor to the likelihood, the reason that observation ceased does not change the evidence represented by the realized likelihood function.

Consider independent Bernoulli trials with success probability (p). Under a fixed-sample design, an experiment may conduct (n) trials and observe exactly (r) successes. Apart from a combinatorial factor independent of (p), its likelihood is

[ L_1(p)\propto p^r(1-p)^{n-r}. ]

A sequential experiment may instead continue until the (r)-th success occurs and stop on trial (n). Its likelihood has the form

[ L_2(p)\propto p^r(1-p)^{n-r}. ]

The binomial and negative-binomial sampling distributions assign different probabilities to their respective collections of possible outcomes, but the realized likelihoods are proportional. The likelihood principle consequently assigns the two observations the same evidential content about (p).

Procedures defined through tail areas can produce different numerical results in these experiments because a tail probability depends on the full sample space and its ordering. A conventional p-value, for example, incorporates probabilities of unobserved outcomes under the specified sampling plan. Its value can therefore change when the stopping rule changes, even though the observed likelihood remains proportional. This is a structural difference between likelihood-based evidence and repeated-sampling calibration.

The same distinction appears in sequential monitoring. If observations are examined after each stage and sampling stops when a boundary is crossed, the likelihood for the final data may coincide with the likelihood under a design that fixed the final sample size in advance. Frequentist error probabilities still depend on the monitoring rule because repeated opportunities to stop alter the distribution of possible decisions.

Relation to Bayesian inference

Bayesian updating begins with a prior distribution (\pi(\theta)) and forms the posterior distribution

[ \pi(\theta\mid x)

\frac{L(\theta;x)\pi(\theta)} {\int_\Theta L(u;x)\pi(u),du}. ]

If two observations have proportional likelihood functions and use the same prior distribution, the proportionality constant cancels during normalization. The resulting posterior distributions are identical. Standard Bayesian inference therefore satisfies the likelihood principle when the prior specification and model remain fixed across the compared experiments.

This compatibility does not identify the likelihood function with the posterior distribution. The posterior combines the likelihood with a probability measure over parameter values, whereas the likelihood alone records how the observed data modify relative support within the model. Improper priors and model-dependent prior constructions can also introduce dependence on the experimental representation that is not contained in the observed likelihood.

Bayesian model comparison requires an additional distinction. A marginal likelihood integrates over a model-specific parameter space and depends on the chosen prior within that model. Proportional likelihoods for a common parameter yield equivalent updates only when the surrounding model and prior structure are also equivalent.

Relation to frequentist procedures

Frequentist methods are defined through their behavior under repeated sampling from a specified experiment. Coverage probabilities, error rates, and sampling distributions depend on outcomes that were possible under the design, including outcomes that were not observed. Such methods therefore need not satisfy the strong likelihood principle.

A confidence interval is calibrated through the proportion of hypothetical repetitions in which the interval-generating procedure contains the true parameter. Changing the stopping rule can change that proportion and can consequently change the interval, even when the final likelihood function is unchanged. The resulting difference concerns procedure-level calibration rather than a disagreement about the algebraic form of the likelihood.

Some frequentist techniques depend primarily on the observed likelihood and consequently approximate likelihood-principle behavior. Maximum likelihood estimation is unchanged by multiplication of the likelihood by a parameter-independent constant. Likelihood-ratio statistics also preserve proportional-likelihood equivalence at the level of their observed values. Their usual reference distributions, however, are derived from repeated-sampling arguments and can retain dependence on the experimental design.

Scope and limitations

The likelihood principle applies only after a statistical model has been specified. Two analyses using different parameter spaces, different latent structures, or different assumptions about the data-generating process need not be equivalent merely because selected parts of their likelihood expressions resemble one another. Model uncertainty is not eliminated by conditioning on observed data.

The principle also concerns information about parameters represented in the likelihood. Features of experimental quality that are omitted from the model cannot be recovered from likelihood proportionality. If a stopping mechanism depends on unobserved quantities related to the parameter, or if missingness is informative, the mechanism contributes to the likelihood and must be included in the model. The simple stopping-rule equivalence then does not apply.

Likelihood equivalence is similarly distinct from practical identity of experiments. Two experiments with proportional realized likelihoods may have had different costs, expected precision, or probabilities of producing decisive results before the data were observed. Those design properties remain relevant to experimental design, while the likelihood principle classifies only the evidential content of the outcomes that actually occurred.

See also