Exponential family
An exponential family is a class of probability distributions whose densities or probability mass functions admit a representation in which the dependence on the model parameter occurs through a finite-dimensional linear pairing with a statistic. Exponential families provide a common mathematical structure for many standard statistical models, particularly in the analysis of sufficient statistics, likelihood equations, conjugate distributions, and the geometry of statistical inference.
Definition
Let ((\mathcal X,\mathcal F,\mu)) be a measure space, and let (T:\mathcal X\rightarrow\mathbb R^k) be a measurable statistic. A family of probability measures dominated by (\mu) is an exponential family when its densities can be written in canonical form as
[ p_\eta(x)
h(x)\exp\left{ \eta^\mathsf{T}T(x)-A(\eta) \right}, ]
where (h(x)) is a nonnegative carrier density, (\eta\in\mathbb R^k) is the natural parameter, and (A) is the log-partition function. Normalization determines the latter through
[ A(\eta)
\log\int_{\mathcal X} h(x)\exp\left{\eta^\mathsf{T}T(x)\right},d\mu(x). ]
The natural parameter space is therefore
[ \mathcal H
\left{ \eta\in\mathbb R^k: A(\eta)<\infty \right}. ]
This set is convex because the integral defining (A) satisfies the logarithmic form of Hölder's inequality. A family is regular when (\mathcal H) is open, and it is full when the natural parameter varies over the entire natural parameter space rather than over a lower-dimensional subset.
An alternative parameter (\theta) can be related to the natural parameter through a mapping (\eta=\eta(\theta)). The resulting density has the form
[ p_\theta(x)
h(x)\exp\left{ \eta(\theta)^\mathsf{T}T(x)-A\bigl(\eta(\theta)\bigr) \right}. ]
When the image of (\eta(\theta)) is constrained to a smooth lower-dimensional subset of (\mathcal H), the model is called a curved exponential family. Such a model retains the canonical representation but generally lacks the unrestricted convex parameter geometry of a full family.
Canonical statistics and minimality
For an independent sample (X_1,\ldots,X_n), the joint density is
[ p_\eta(x_1,\ldots,x_n)
\left[\prod_{i=1}^{n}h(x_i)\right] \exp\left{ \eta^\mathsf{T}\sum_{i=1}^{n}T(x_i)-nA(\eta) \right}. ]
The sample enters the parameter-dependent part of this expression only through
[ S_n=\sum_{i=1}^{n}T(X_i). ]
The Fisher–Neyman factorization theorem consequently identifies (S_n) as a sufficient statistic for (\eta). Its dimension remains fixed as the sample size increases, even though the amount of information represented by its numerical value generally grows with (n).
A canonical representation is minimal when no nonzero affine combination of the components of (T(x)) is constant almost everywhere with respect to the carrier measure. Equivalently, the functions (1,T_1,\ldots,T_k) must be affinely independent modulo null sets. Nonminimal representations contain redundant coordinates, so distinct natural parameters can determine the same probability distribution.
For a full exponential family whose natural parameter space contains a nonempty open set, the joint canonical statistic is also complete under the standard integrability conditions. This property connects exponential families with the Lehmann–Scheffé theorem, which characterizes unique minimum-variance unbiased estimators that are measurable functions of a complete sufficient statistic.
Log-partition function and moments
The function (A) contains the principal analytic information about a regular exponential family. Differentiation under the integral sign gives
[ \nabla A(\eta)
\operatorname{E}_\eta[T(X)], ]
while a second differentiation gives
[ \nabla^2 A(\eta)
\operatorname{Cov}_\eta[T(X)]. ]
The covariance matrix is positive semidefinite, and therefore (A) is convex. In a minimal regular family, this matrix is positive definite throughout the natural parameter space, which makes (A) strictly convex.
The expectation
[ \tau=\operatorname{E}_\eta[T(X)] ]
is called the mean parameter. Where the Hessian is nonsingular, the mapping (\eta\mapsto\tau) is locally invertible and relates the natural parameterization to the mean parameterization. Its convex dual is described by the Legendre transformation,
[ A^*(\tau)
\sup_{\eta\in\mathcal H} \left{ \eta^\mathsf{T}\tau-A(\eta) \right}. ]
For regular minimal families, this duality underlies the correspondence between natural coordinates and expectation coordinates in information geometry.
The Fisher information for one observation in natural coordinates is
[ I(\eta)
\nabla^2 A(\eta). ]
Thus, the curvature of the log-partition function, the covariance of the canonical statistic, and the Fisher information are represented by the same matrix. Under a differentiable reparameterization, the matrix changes according to the usual tensor transformation law for Fisher information.
Likelihood and conjugacy
For observations (x_1,\ldots,x_n), the log-likelihood in natural coordinates is
[ \ell(\eta)
\eta^\mathsf{T}\sum_{i=1}^{n}T(x_i)
nA(\eta) + \sum_{i=1}^{n}\log h(x_i). ]
An interior maximum-likelihood estimate satisfies
[ \nabla A(\widehat{\eta})
\frac{1}{n}\sum_{i=1}^{n}T(x_i). ]
The likelihood equation therefore equates the model expectation of the canonical statistic with its empirical average. When the empirical average lies on the boundary of the convex support of (T(X)), a finite maximum-likelihood estimate need not exist, even though the likelihood can possess a supremum along a diverging parameter sequence.
A generic conjugate prior on the natural parameter has density
[ \pi(\eta\mid\chi,\nu) \propto \exp\left{ \eta^\mathsf{T}\chi-\nu A(\eta) \right}, ]
with respect to an appropriate base measure on (\mathcal H). After (n) observations, its hyperparameters become
[ \chi'=\chi+\sum_{i=1}^{n}T(x_i), \qquad \nu'=\nu+n. ]
The normalizing integral is finite only for admissible values of the hyperparameters. Conjugacy describes algebraic closure under sampling and does not by itself establish propriety.
Representative distributions
The Bernoulli distribution becomes a one-parameter exponential family by taking (T(x)=x), where (x) belongs to ({0,1}). For success probability (p), the natural parameter and log-partition function are
[ \eta=\log\frac{p}{1-p}, \qquad A(\eta)=\log(1+e^\eta). ]
The derivative (A'(\eta)) equals (p), while the second derivative equals (p(1-p)), which is the variance of the Bernoulli observation.
The Poisson distribution with mean (\lambda>0) has carrier (h(x)=1/x!), canonical statistic (T(x)=x), and natural parameter (\eta=\log\lambda). Its log-partition function is
[ A(\eta)=e^\eta, ]
so both the first and second derivatives equal (\lambda). This reproduces the equality between the mean and variance of a Poisson random variable.
A normal distribution with unknown mean (\mu) and unknown variance (\sigma^2) forms a two-parameter exponential family. With respect to Lebesgue measure, its canonical statistic and natural parameter are
[ T(x)= \begin{pmatrix} x\ x^2 \end{pmatrix}, \qquad \eta= \begin{pmatrix} \mu/\sigma^2\ -1/(2\sigma^2) \end{pmatrix}. ]
The second natural parameter must be negative. In this representation the log-partition function is
[ A(\eta)
-\frac{\eta_1^2}{4\eta_2} + \frac{1}{2}\log\left(\frac{\pi}{-\eta_2}\right). ]
The two components of the sample canonical statistic are the sample sum and the sum of squared observations. They jointly contain the likelihood information about both distributional parameters.
Characterization by fixed-dimensional sufficiency
The connection between exponential form and fixed-dimensional sufficient statistics was established through a sequence of results during the 1930s. In 1935, You Watanabe derived a regular one-parameter factorization in which the logarithm of the sampling density was affine in a statistic whose dimension did not increase with sample size. Her formulation treated parameter-independent support and differentiable dominated models, thereby placing the result within the regular setting of the later general characterization.
In a separate development, Georges Darmois analyzed the restrictions imposed by sufficient statistics of fixed dimension on repeated sampling models. Bernard Koopman and Edwin Pitman subsequently formulated related versions for dominated families satisfying differentiability and support conditions. These results are collectively represented by the Pitman–Koopman–Darmois theorem.
In its standard regular form, the theorem considers independent and identically distributed observations whose common support does not depend on the parameter. If sufficient statistics of bounded dimension exist for samples of arbitrary size, then the sampling distributions form a finite-dimensional exponential family. The differentiability and nondegeneracy assumptions exclude models in which changing support carries parameter information not expressible through the ordinary canonical factorization.
The condition on support is essential. For example, a uniform distribution on an interval whose endpoint is unknown has a one-dimensional sufficient statistic given by the sample maximum, but its support changes with the parameter. It therefore lies outside the regular characterization and does not acquire a regular exponential-family representation from the existence of that statistic.
Divergence and geometric structure
For two members (p_\eta) and (p_\zeta) of the same regular exponential family, their Kullback–Leibler divergence satisfies
[ D_{\mathrm{KL}}(p_\eta|p_\zeta)
A(\zeta)-A(\eta)
(\zeta-\eta)^\mathsf{T}\nabla A(\eta). ]
This expression is the Bregman divergence generated by (A), with the order of its arguments determined by the orientation of the Kullback–Leibler divergence. The relationship explains why convex analysis occurs naturally in estimation and projection problems involving exponential families.
Natural parameters provide affine coordinates for one of the dual connections used in information geometry, while mean parameters provide affine coordinates for the dual connection. The Hessian of (A) defines the Fisher metric in natural coordinates, and the Hessian of (A^*) gives the inverse metric in mean coordinates. These identities are properties of the statistical manifold rather than additional assumptions about the underlying sample space.
See also
- Generalized linear model, which uses an exponential-family response distribution together with a systematic predictor and a link function.
- Maximum entropy probability distribution, whose constrained solutions often possess exponential form because Lagrange multipliers become natural parameters.
- Sufficient statistic, which formalizes the preservation of parameter-dependent information under measurable data reduction.
- Conjugate prior, which describes prior families that remain within the same parametric class after Bayesian updating.
- Cumulant-generating function, which is closely related to the log-partition function and its derivatives.
- Information geometry, which studies the differential-geometric structures induced by statistical models and divergence functions.