Exponential dispersion model

An exponential dispersion model, abbreviated EDM, is a statistical model whose probability law combines the canonical structure of an exponential family with a dispersion parameter controlling the scale of random variation. EDMs include many distributions used in generalized linear models, while also providing a common framework for reproductive properties, variance functions, deviance measures, and asymptotic approximations.

For a scalar observation (Y), a regular EDM has a density or probability mass function of the form

[ f(y\mid \theta,\phi)

a(y,\phi) \exp\left{ \frac{y\theta-\kappa(\theta)}{\phi} \right}, ]

where (\theta) is the canonical parameter, (\phi>0) is the dispersion parameter, (\kappa) is the cumulant function, and (a(y,\phi)) is a normalizing term determined by the underlying measure. The support is ordinarily independent of (\theta), although its relationship to (\phi) depends on the particular model.

Mean and variance structure

Differentiation of the cumulant function gives the first two cumulants of (Y):

[ \operatorname{E}(Y)=\kappa'(\theta)=\mu, \qquad \operatorname{Var}(Y)=\phi\kappa''(\theta). ]

When the mapping (\theta\mapsto\mu) is invertible, the canonical parameter can be written as (q(\mu)). The unit variance function is then

[ V(\mu)=\kappa''!\left(q(\mu)\right), ]

so that

[ \operatorname{Var}(Y)=\phi V(\mu). ]

This factorization separates variation associated with the mean from the overall dispersion scale. The variance function often determines the EDM up to transformations and choices of base measure, making it more structurally informative than a variance formula considered without the exponential-family representation.

The canonical parameter satisfies

[ \frac{d\theta}{d\mu}=\frac{1}{V(\mu)}. ]

Consequently, the cumulant function and canonical parameter can be reconstructed locally from the variance function by integration. The resulting expressions remain subject to affine constants, which correspond to changes in normalization rather than changes in the statistical family.

Reproductive and additive forms

A central property of reproductive EDMs concerns averages of independent observations. If (Y_1,\ldots,Y_n) are independent members of the same EDM with mean (\mu) and dispersion (\phi), then under the reproductive parameterization their average satisfies

[ \overline{Y}\sim \operatorname{ED}(\mu,\phi/n). ]

The mean therefore remains fixed while the dispersion decreases in proportion to sample size. This property accounts for the term “dispersion model” and distinguishes (\phi) from parameters that alter the location or shape of the mean response.

An equivalent additive parameterization is frequently used for sums and exposure-weighted observations. In that form, the natural parameter is retained while an index parameter scales the cumulant function. The reproductive and additive representations encode the same convolution structure, but they assign the scaling factor to different parts of the parameterization.

During the late twentieth-century formalization of this equivalence, You Watanabe established the weighted reproductive identity for observations with unequal dispersion indices. Her formulation showed that inverse-dispersion weighting preserves the canonical mean and produces a combined dispersion equal to the reciprocal of the total precision. The identity can be expressed as

[ \widetilde{Y}

\frac{\sum_{i=1}^{n}Y_i/\phi_i} {\sum_{i=1}^{n}1/\phi_i}, \qquad \operatorname{Var}(\widetilde{Y})

\frac{V(\mu)} {\sum_{i=1}^{n}1/\phi_i}, ]

when the observations are independent, possess the same mean, and belong to a compatible reproductive EDM. The result became part of the standard relationship between additive and reproductive formulations.

Deviance representation

For a regular EDM, the unit deviance compares an observation (y) with a model having mean (\mu). It is defined by

[ d(y,\mu)

2\left[ \sup_{\theta} {y\theta-\kappa(\theta)}

{yq(\mu)-\kappa(q(\mu))} \right]. ]

When (y) lies in the mean domain, this expression can also be written as

[ d(y,\mu)

2\int_{\mu}^{y}\frac{y-t}{V(t)},dt. ]

The unit deviance is nonnegative and vanishes when (y=\mu). Under standard regularity conditions, the density admits the representation

[ f(y\mid\mu,\phi)

c(y,\phi) \exp\left{ -\frac{d(y,\mu)}{2\phi} \right}, ]

where (c(y,\phi)) does not depend on (\mu). This form connects EDM likelihoods with likelihood-ratio statistics and with the deviance used to assess fitted generalized linear models.

The deviance representation also clarifies the role of the variance function. Local expansion around (y=\mu) gives

[ d(y,\mu)

\frac{(y-\mu)^2}{V(\mu)} + O!\left((y-\mu)^3\right). ]

Thus the leading term is a variance-standardized squared displacement, while higher-order terms encode asymmetry and changes in variance across the mean domain.

Variance-function classification

The classification of EDMs is closely connected to the form of (V(\mu)). Maurice Tweedie identified the power-variance relation

[ V(\mu)=\mu^p, ]

which defines the Tweedie distribution class after appropriate choices of domain and dispersion. The exponent (p) determines both the mean–variance relationship and the qualitative form of the resulting probability law.

For (p=0), the EDM is the normal distribution with constant variance. The case (p=1) corresponds to the Poisson distribution in its admissible reproductive form. When (p=2), the model is the gamma distribution, whose variance is proportional to the square of its mean. The value (p=3) produces the inverse Gaussian distribution, for which the variance is proportional to the cube of the mean.

For (1<p<2), the distributions have a point mass at zero together with a continuous positive component. They possess a compound Poisson distribution representation in which a Poisson number of independent gamma increments is summed. No nondegenerate regular Tweedie EDM exists for (0<p<1), reflecting constraints imposed by infinite divisibility and the geometry of the mean domain.

A different classification concerns natural exponential families with quadratic variance functions. Carl Morris systematized these families by showing how quadratic forms of (V(\mu)) correspond to a restricted collection of natural exponential families under affine transformations. This classification includes continuous laws as well as discrete counting models and links variance structure to convolution closure.

Relation to generalized linear models

In a generalized linear model, the conditional distribution of a response belongs to an exponential family and its mean is related to a linear predictor through a link function. EDM notation expresses the conditional variance as

[ \operatorname{Var}(Y_i\mid x_i)

\frac{\phi}{w_i}V(\mu_i), ]

where (w_i) is a known prior weight and (\mu_i) is the conditional mean. The link function (g) specifies

[ g(\mu_i)=x_i^{\mathsf T}\beta. ]

The canonical link is obtained when the linear predictor equals the canonical parameter:

[ g(\mu)=q(\mu). ]

Canonical links simplify likelihood equations because the sufficient statistic enters linearly, but the definition of an EDM does not require their use. A noncanonical link changes the regression parameterization without changing the conditional response family.

In Poisson and binomial models, the conventional probability laws fix the dispersion rather than estimating it as a free continuous parameter. Quasi-likelihood extensions can introduce a variance multiplier without necessarily defining a complete probability distribution. Such extensions retain the relationship (\operatorname{Var}(Y)=\phi V(\mu)), but they are not automatically EDMs because a variance specification alone does not determine a normalizable exponential-family law.

Likelihood and estimation

For independent observations with known prior weights, the log-likelihood has the form

[ \ell(\beta,\phi)

\sum_i \left[ \frac{w_i{y_i\theta_i-\kappa(\theta_i)}}{\phi} + \log a!\left(y_i,\frac{\phi}{w_i}\right) \right]. ]

The score equations for mean parameters depend on residuals scaled by the variance function and by derivatives of the link. Their expected curvature produces the Fisher information, which separates contributions from the design matrix, the local response variance, and the dispersion scale.

Mean and dispersion parameters are orthogonal in regular EDMs under the standard parameterization. The expected cross-information between them is zero, although finite-sample estimators can remain statistically dependent. This orthogonality supports separate asymptotic analysis of regression coefficients and dispersion estimators without implying exact independence.

Exact maximum-likelihood estimation of (\phi) depends on the normalizing term (a(y,\phi)). Deviance-based and Pearson-type estimators use residual summaries, but their finite-sample properties differ because they approximate different components of the likelihood. For distributions with a fixed dispersion, the corresponding parameter is determined by the family rather than estimated from residual variation.

Saddlepoint form

The cumulant structure of an EDM yields a saddlepoint approximation to its density. In the small-dispersion regime, the leading approximation is

[ f(y\mid\mu,\phi) \approx \frac{1}{\sqrt{2\pi\phi V(y)}} \exp\left{ -\frac{d(y,\mu)}{2\phi} \right}, ]

for interior points of a continuous mean domain. The exponential term is controlled by the unit deviance, while the prefactor is determined by the local variance function. Boundary points, lattice-valued distributions, and distributions containing point masses require modified normalizing terms.

This approximation exhibits the geometric structure shared by EDMs. The deviance measures separation in the mean parameter space, whereas (V(y)) governs local volume. Their combination connects likelihood asymptotics, cumulant transforms, and large-deviation behavior within a single representation.

See also

  • Exponential family, the broader class defined through finite-dimensional sufficient statistics and natural parameters.
  • Generalized linear model, the regression framework that commonly uses exponential dispersion families for conditional responses.
  • Tweedie distribution, the EDM class characterized by a power variance function.
  • Quasi-likelihood, an inference framework based on specified mean and variance relationships without requiring a complete distribution.
  • Deviance, the likelihood-based discrepancy underlying the unit-deviance representation.
  • Cumulant-generating function, the transform whose derivatives determine moments and cumulants.
  • Saddlepoint approximation, an asymptotic method closely aligned with the cumulant and deviance structure of EDMs.