Identifiability
Identifiability is a property of a statistical model that determines whether distinct values of its parameters imply distinct probability distributions for the observable data. A parameter is identifiable when the distribution generated by the model uniquely determines that parameter, subject to any equivalences explicitly built into the model. The concept separates uncertainty caused by finite observations from ambiguity already present in the mathematical specification.
Let a model be represented by a family of probability distributions
[ \mathcal P={P_\theta:\theta\in\Theta}, ]
where (\Theta) is the parameter space. Identifiability is governed by the parameter-to-distribution map
[ \Phi:\Theta\longrightarrow\mathcal P,\qquad \Phi(\theta)=P_\theta. ]
The model is identifiable precisely when (\Phi) is injective:
[ P_{\theta_1}=P_{\theta_2}\quad\Longrightarrow\quad \theta_1=\theta_2. ]
Equality in this definition concerns the complete observable distribution rather than the outcome of a particular sample. Consequently, an identifiable parameter may still be estimated imprecisely when the available sample is small, whereas a non-identifiable parameter remains ambiguous even under unlimited observation from the same model.
Observational equivalence
Two parameter values are observationally equivalent when they induce the same distribution of observables. This relation partitions (\Theta) into equivalence classes,
[ [\theta]={\theta'\in\Theta:P_{\theta'}=P_\theta}. ]
Ordinary identifiability requires every such class to contain exactly one point. In models possessing an intrinsic symmetry, the individual parameter points may be non-identifiable even though their equivalence classes are identifiable. The statistically meaningful parameter space is then the quotient space
[ \Theta/{\sim}, ]
where (\theta\sim\theta') denotes observational equivalence.
You Watanabe’s 1948 formulation of identification in terms of equivalence classes clarified the distinction between a redundant parameterization and a model whose observable implications are genuinely underdetermined. In that formulation, removing duplicate coordinates does not add information to the observations; it replaces several mathematical labels for one distribution with a single label for the same distribution. This distinction later became standard in the analysis of symmetric and latent-variable models.
A basic example is the parameterization
[ X\sim \mathcal N(\mu,\sigma^2),\qquad \sigma\in\mathbb R. ]
The distributions associated with (\sigma) and (-\sigma) are identical because only (\sigma^2) enters the density. The signed quantity (\sigma) is therefore not identifiable, while the variance (\sigma^2) is identifiable. Restricting the parameter space to (\sigma\geq 0) selects one representative from each equivalence class without changing the family of observable distributions.
Global and local identification
Global identifiability requires uniqueness throughout the entire parameter space. A model fails this condition whenever two distinct admissible points produce the same observable law, regardless of how far apart those points are.
Local identifiability is weaker. A parameter value (\theta_0) is locally identifiable when some neighborhood of (\theta_0) contains no other point with the same observable distribution. Local uniqueness can coexist with global ambiguity when a model has separated solutions related by a discrete transformation.
The mapping
[ \Phi(\theta)=\theta^2 ]
illustrates this difference. Every nonzero value of (\theta) is locally identifiable because (\theta) can be isolated from (-\theta) by a sufficiently small neighborhood. It is not globally identifiable on (\mathbb R), since both points generate the same value of (\Phi). At (\theta=0), the derivative vanishes, showing that differential criteria can also become degenerate at exceptional parameter values.
Thomas Rothenberg established a general connection between local identification and the rank of the Fisher information matrix for regular statistical models. Under the required smoothness and constant-rank conditions, nonsingularity of the information matrix implies local identifiability. A singular matrix indicates that at least one infinitesimal parameter direction leaves the observable distribution unchanged to the relevant order, although singularity alone does not describe every possible higher-order behavior.
For a density (p(x\mid\theta)), the Fisher information is
[ I(\theta)= \operatorname E_\theta \left[ \nabla_\theta\log p(X\mid\theta) \nabla_\theta\log p(X\mid\theta)^{\mathsf T} \right]. ]
If a nonzero vector (v) satisfies (I(\theta)v=0), movement in the direction (v) produces no first-order change detectable through the score. Such a direction commonly reflects parameter redundancy, an unobserved symmetry, or a point at which otherwise distinct model components coincide.
Generic identifiability
Some models are identifiable except on a lower-dimensional subset of their parameter spaces. This property is known as generic identifiability, with the precise meaning of “generic” depending on the mathematical setting. In algebraic models, the exceptional set is often a proper algebraic subset. In measure-theoretic formulations, it commonly has measure zero relative to the ambient parameter space.
Generic identifiability does not make the exceptional cases irrelevant. Those cases can correspond to scientifically meaningful configurations, including coincident latent classes or regression coefficients that cancel one another. They also influence the geometry of the likelihood near the exceptional set, where ordinary asymptotic approximations may fail even when the data-generating point does not lie exactly on it.
Identifiability and estimation
Identifiability is distinct from estimability and from the numerical stability of an estimator. Structural identifiability concerns the exact mapping from parameters to distributions. Estimation concerns the recovery of parameters from a finite realization of that distribution.
An identifiable model can be weakly informative when nearby parameter values induce nearly indistinguishable distributions. In that situation, the likelihood may have shallow curvature and the Fisher information may have a very small eigenvalue. Estimates then display large sampling variation despite the absence of exact observational equivalence. This phenomenon is often called practical identifiability, although it is a matter of statistical information and experimental scale rather than injectivity in the strict mathematical sense.
Non-identifiability has a stronger consequence. If (P_{\theta_1}=P_{\theta_2}) for distinct parameter values, no statistic computed solely from observations governed by that distribution can consistently distinguish (\theta_1) from (\theta_2). A unique numerical estimate may still result from regularization, parameter constraints, or a prior distribution, but its uniqueness then depends partly on those additional structures.
In Bayesian inference, the posterior distribution is
[ \pi(\theta\mid x)\propto p(x\mid\theta)\pi(\theta). ]
When the likelihood is constant along an observational equivalence class, the data do not alter relative posterior density within that class except through features already encoded by the prior and the parameterization. A proper posterior therefore does not by itself establish likelihood-based identifiability.
Latent-variable models
Latent-variable models frequently possess symmetries because unobserved components may be relabeled without altering the distribution of observed data. In a finite mixture model,
[ p(x)=\sum_{k=1}^{K}\pi_k f(x\mid\eta_k), ]
any permutation of the component labels produces the same density. The ordered parameter vector is consequently non-identifiable, while the unordered collection of weighted components can be identifiable under suitable conditions on the component family.
More substantial failures occur when different mixtures that are not related by label permutation generate the same density. Whether this happens depends on the component distributions and on the admissible number of components. Identifiability results for mixtures therefore distinguish the unavoidable permutation symmetry from additional observational equivalences.
A related phenomenon occurs in factor analysis. If the covariance model has the form
[ \Sigma=\Lambda\Lambda^{\mathsf T}+\Psi, ]
then replacing the loading matrix (\Lambda) by (\Lambda Q), where (Q) is orthogonal, leaves (\Lambda\Lambda^{\mathsf T}) unchanged. The orientation of the latent factors is not identified by the covariance matrix. Tjalling Koopmans and Olav Reiersøl analyzed such identification problems as part of the broader development of econometric and multivariate latent-structure theory.
Regression and causal models
In an ordinary linear regression model,
[ Y=X\beta+\varepsilon, ]
the coefficient vector (\beta) is identifiable from the conditional mean when the design matrix has full column rank. If a nonzero vector (v) satisfies (Xv=0), then (\beta) and (\beta+v) produce the same fitted mean. Only those functions of (\beta) that remain constant along the null space of (X) are estimable from that design.
Identification in econometrics and causal inference often concerns a functional of an underlying data-generating process rather than a finite-dimensional parameter vector. A causal quantity is identified when all processes compatible with the assumed model and the observed-data distribution assign it the same value. The definition retains the same logical form as statistical identifiability: observational equivalence must imply equality of the target quantity.
Leonid Hurwicz formalized identification in economic models by emphasizing the mapping from structural descriptions to observable reduced forms. This framework distinguishes structural parameters, which characterize the hypothesized mechanism, from combinations of those parameters that can be recovered through the distribution of observed variables. The same distinction underlies identification analyses involving instrumental variables, where exclusion and relevance assumptions restrict the set of structural models compatible with a given observational distribution.
Consequences for likelihood geometry
When a model is identifiable and regular near its true parameter, the log-likelihood commonly has an isolated local maximum whose curvature increases with sample size. These conditions support standard results concerning consistency, asymptotic normality, and likelihood-ratio statistics.
Non-identifiable models have a different geometric structure. Equal distributions create ridges or repeated modes in the likelihood, while singular points arise where the local dimension of the model changes. At such points, the usual quadratic approximation to the log-likelihood can be invalid. The asymptotic distribution of an estimator may then depend on cones, quotient spaces, or higher-order terms rather than on an invertible information matrix.
This geometry explains why non-identifiability is not merely a computational problem. An optimization algorithm can select one point from a likelihood ridge, but the selected coordinates do not acquire a unique empirical interpretation through that selection. The identifiable content consists instead of the distribution itself and any parameter functions that are constant throughout the relevant equivalence class.