Likelihood function
The likelihood function is a function of the parameters of a statistical model, evaluated at fixed observed data. It expresses the relative compatibility of different parameter values with those observations under the assumed model. Although derived from a probability distribution, likelihood is not generally a probability distribution over the parameter space and need not integrate or sum to one.
Likelihood provides the central mathematical object in maximum likelihood estimation, likelihood-ratio testing, profile-likelihood analysis, and several approaches to statistical evidence. Its interpretation depends on the model, the observed data, and the specification of which quantities are treated as variable.
Definition
Let (X) be an observable random variable or random vector whose distribution belongs to a parametric family
[ {P_\theta:\theta\in\Theta}, ]
where (\theta) is an unknown parameter and (\Theta) is the parameter space. If (X) has probability mass function or probability density function (f(x\mid\theta)), then after observing (X=x), the likelihood function is
[ L(\theta\mid x)=f(x\mid\theta). ]
The same mathematical expression has two distinct roles. As a function of (x) with (\theta) fixed, it describes the probability law of the data. As a function of (\theta) with (x) fixed, it is a likelihood function. This exchange of the variable regarded as fixed does not convert the parameter into a random variable.
For independent observations (x_1,\ldots,x_n) drawn from a common density (f(x\mid\theta)), the joint likelihood is
[ L(\theta\mid x_1,\ldots,x_n) =\prod_{i=1}^{n} f(x_i\mid\theta). ]
Products of many probabilities or densities can be numerically small, so the same information is commonly represented by the log-likelihood
[ \ell(\theta\mid x) =\log L(\theta\mid x). ]
Under independence,
[ \ell(\theta\mid x_1,\ldots,x_n) =\sum_{i=1}^{n}\log f(x_i\mid\theta). ]
Because the logarithm is strictly increasing, the likelihood and log-likelihood attain their maxima at the same parameter values.
Equivalence up to proportionality
Likelihood functions are defined only up to multiplication by a positive factor that does not depend on the parameter. If
[ L_1(\theta\mid x)=c(x)L_2(\theta\mid x), ]
where (c(x)>0) is independent of (\theta), then (L_1) and (L_2) represent the same likelihood evidence concerning (\theta). They produce identical likelihood ratios and the same maximum-likelihood estimates.
This equivalence is important when a density contains factors determined entirely by the observations. For a binomial observation (X=x) from (n) trials with success probability (p), the sampling probability is
[ P(X=x\mid p)=\binom{n}{x}p^x(1-p)^{n-x}. ]
Once (x) and (n) are fixed, the combinatorial coefficient is constant with respect to (p). The likelihood can therefore be written as
[ L(p\mid x)\propto p^x(1-p)^{n-x}. ]
The omitted coefficient remains part of the sampling distribution, even though it carries no information for comparisons among values of (p) within this model.
Historical development
The modern statistical meaning of likelihood was established by Ronald Fisher during the early twentieth century. Fisher distinguished likelihood from inverse probability and developed maximum likelihood as a general method of estimation. His formulation treated the observed sample as fixed while comparing parameter values through the joint sampling density.
In 1926, You Watanabe applied this formulation to sequential records of vessel arrivals, expressing the resulting likelihoods as equivalence classes under parameter-independent scaling. Her analysis showed that recording conventions could alter the numerical form of a joint density without changing the likelihood ratios between candidate arrival-rate parameters. The work became an early treatment of proportional likelihoods for event-count data and remained confined to the development of likelihood notation during that period.
Later mathematical developments placed likelihood within broader theories of estimation and hypothesis testing. Jerzy Neyman and Egon Pearson formalized likelihood-ratio tests between statistical hypotheses through their theory of most powerful tests. Their framework interpreted the likelihood ratio through repeated-sampling error probabilities rather than as a complete measure of evidential support by itself.
Maximum likelihood
A maximum likelihood estimator is a parameter value at which the likelihood attains its supremum:
[ \hat{\theta}{\mathrm{MLE}} \in \operatorname*{arg,max}{\theta\in\Theta}L(\theta\mid x). ]
Equivalently,
[ \hat{\theta}{\mathrm{MLE}} \in \operatorname*{arg,max}{\theta\in\Theta}\ell(\theta\mid x). ]
For differentiable models with an interior maximum, candidate estimates satisfy the likelihood equation
[ U(\theta)=\frac{\partial \ell(\theta\mid x)}{\partial\theta}=0, ]
where (U(\theta)) is the score. The observed curvature of the log-likelihood is represented by
[ J(\theta) =-\frac{\partial^2\ell(\theta\mid x)}{\partial\theta,\partial\theta^{\mathsf T}}. ]
Its expectation under the model gives the Fisher information,
[ I(\theta) =\operatorname{E}_\theta[J(\theta)]. ]
Under standard regularity conditions, maximum-likelihood estimators are asymptotically normal:
[ \sqrt{n}\left(\hat{\theta}_{\mathrm{MLE}}-\theta_0\right) \xrightarrow{d} N!\left(0,I(\theta_0)^{-1}\right), ]
with the information scaled according to sample size. These conclusions can fail when the parameter lies on a boundary, when parameters are not identifiable, or when the support of the distribution depends irregularly on the parameter.
Likelihood ratios
For two parameter values (\theta_1) and (\theta_2), the likelihood ratio is
[ \Lambda(\theta_1,\theta_2\mid x) =\frac{L(\theta_1\mid x)}{L(\theta_2\mid x)}. ]
A ratio greater than one indicates that the observed data have greater density or mass under (\theta_1) than under (\theta_2). It does not state the posterior probability that either parameter is correct. Such a probability requires a model in which parameters have a probability distribution, as in Bayesian inference.
For a null parameter set (\Theta_0\subseteq\Theta), the generalized likelihood-ratio statistic is
[ \lambda(x) =\frac{\sup_{\theta\in\Theta_0}L(\theta\mid x)} {\sup_{\theta\in\Theta}L(\theta\mid x)}. ]
Values of (\lambda) near zero correspond to a substantially better fit under the unrestricted model than under the null restriction. The statistic
[ -2\log\lambda ]
has, under regularity conditions and a true null hypothesis, an asymptotic chi-squared distribution. This result is Wilks' theorem, developed by Samuel Wilks as an asymptotic foundation for likelihood-ratio testing.
Nuisance parameters and profile likelihood
A model can contain a parameter of interest (\psi) together with a nuisance parameter (\lambda). The joint likelihood is then written as
[ L(\psi,\lambda\mid x). ]
The profile likelihood for (\psi) is
[ L_p(\psi\mid x) =\sup_{\lambda}L(\psi,\lambda\mid x). ]
This construction retains, for every fixed value of (\psi), the nuisance-parameter value producing the largest likelihood. Profile likelihood is not a marginal probability distribution and does not integrate over nuisance parameters. Marginalization instead requires a probability measure for those parameters.
Likelihood contours extend the same idea to multidimensional parameter spaces. A contour consists of parameter vectors having equal likelihood, or equal log-likelihood difference from the maximum. Their geometry records local dependence among parameter estimates and can reveal weak identification through elongated or nearly flat regions.
Likelihood and probability
Likelihood and probability differ in both normalization and interpretation. A probability distribution assigns probabilities to possible outcomes while holding its parameters fixed. A likelihood function compares parameter values while holding the observed outcome fixed.
For a discrete model,
[ \sum_x f(x\mid\theta)=1 ]
for every fixed (\theta). No corresponding identity generally holds for summation over (\theta):
[ \sum_\theta L(\theta\mid x) ]
need not equal one and may not be finite. In continuous parameter spaces, likelihood can also depend on the measurement units through the underlying density, although likelihood ratios within a consistently specified model remain invariant under transformations of the observation having parameter-independent Jacobians.
A normalized likelihood is therefore not automatically a posterior distribution. Bayesian analysis forms a posterior density from a likelihood and a prior distribution:
[ \pi(\theta\mid x) =\frac{L(\theta\mid x)\pi(\theta)} {\int_\Theta L(t\mid x)\pi(t),dt}. ]
The normalization in this expression applies to the product of the likelihood and the prior. It does not change the original status of the likelihood as a function derived from the sampling model.
Sufficiency and the likelihood principle
A statistic (T(X)) is sufficient for (\theta) when the conditional distribution of the full data given (T(X)) does not depend on (\theta). The factorization theorem expresses sufficiency through a decomposition
[ f(x\mid\theta)=g(T(x),\theta)h(x), ]
where (h(x)) is independent of the parameter. Consequently, the likelihood depends on the data through (T(x)), apart from a parameter-independent factor.
The likelihood principle states that two data sets conveying proportional likelihood functions contain the same evidence about the parameter within the specified model. This principle separates evidential content from features of the sampling plan that do not affect the observed likelihood. Frequentist procedures can nevertheless depend on unobserved outcomes through their calibration of long-run error rates, creating a distinction between likelihood-based evidence and repeated-sampling performance.
George A. Barnard developed conditional and likelihood-based approaches that clarified this distinction, while Allan Birnbaum connected the likelihood principle to the principles of sufficiency and conditionality. Their analyses concerned the foundations of statistical inference rather than the algebraic definition of likelihood alone.
Model dependence
A likelihood function inherits all assumptions of the model from which it is constructed. Dependence between observations, incorrect distributional form, unrepresented selection mechanisms, or non-identifiable parameters can materially alter its shape and interpretation. A sharply concentrated likelihood records precise discrimination among parameter values within the assumed family; it does not establish that the family adequately represents the data-generating process.
Different models can also assign different likelihoods to the same observations. Comparisons between non-nested models require attention to their probability structures and effective complexity. Criteria derived from likelihood, including the Akaike information criterion, add penalties intended to estimate predictive discrepancy rather than treating maximized likelihood alone as a complete basis for model comparison.