Sufficient statistic

A sufficient statistic is a statistic that preserves all information in observed data that is relevant to inference about a parameter under a specified statistical model. If (X) denotes the full observation, (T=T(X)) is sufficient for a parameter (\theta) when the conditional distribution of (X) given (T) does not depend on (\theta). Sufficiency is therefore a relation among a model, a parameter, and a statistic rather than an intrinsic property of a data transformation.

The concept formalizes lossless reduction for model-based inference. A sufficient statistic may occupy a much lower-dimensional space than the original sample while retaining the sample's entire likelihood-based content concerning the parameter. Information about features outside the model, including departures from its assumptions, need not survive this reduction.

Formal definition

Let ((\mathcal X,\mathcal A)) be a sample space, let ({P_\theta:\theta\in\Theta}) be a family of probability measures on it, and let (T:\mathcal X\rightarrow\mathcal T) be measurable. The statistic (T) is sufficient for (\theta) if there exists a version of the conditional distribution

[ P_\theta(X\in A\mid T) ]

that is the same for every (\theta\in\Theta), for each measurable set (A\in\mathcal A). Equivalently, after the value of (T) has been fixed, the remaining variation in (X) contains no model-defined information about (\theta).

Sufficiency can also be assigned to a sub-(\sigma)-algebra of (\mathcal A). A statistic is sufficient when the (\sigma)-algebra it generates is sufficient. This formulation accommodates models in which densities do not exist and clarifies that one-to-one transformations of a sufficient statistic remain sufficient.

The full observation (X) is always sufficient for its own model. This formally correct reduction is usually uninformative because it performs no reduction at all, a circumstance sometimes described as the identity statistic having preserved the data with exemplary literalness.

Factorization criterion

For a family dominated by a common measure, sufficiency is characterized by the Neyman–Fisher factorization theorem. Suppose (X) has density or probability mass function (f_\theta(x)). A statistic (T(X)) is sufficient if and, under the usual measurability conditions, only if the density can be written as

[ f_\theta(x)=g_\theta!\left(T(x)\right)h(x), ]

where (g_\theta) contains all dependence on (\theta), while (h) is independent of (\theta). The theorem converts a statement about conditional distributions into a statement about the structure of the likelihood function.

Ronald Fisher introduced the modern statistical concept of sufficiency during the development of likelihood-based inference. Jerzy Neyman subsequently established a systematic factorization criterion, placing the concept within a general theory of statistical models. The common name of the theorem reflects these distinct stages of development.

Parameter-dependent support requires careful interpretation because an indicator of the support is part of the density. If that indicator depends on the sample only through (T), it may be incorporated into (g_\theta(T)). Otherwise, a proposed factorization can conceal information about the parameter that remains elsewhere in the observation.

Standard model structures

Bernoulli observations

Let (X_1,\ldots,X_n) be independent Bernoulli distribution observations with success probability (p). Their joint probability mass function is

[ f_p(x_1,\ldots,x_n) =p^{\sum_i x_i}(1-p)^{n-\sum_i x_i}. ]

The sample enters the likelihood only through

[ T(X)=\sum_{i=1}^{n}X_i. ]

The number of successes is therefore sufficient for (p). Conditional on this total, every binary sequence with the same number of successes has the same probability, independently of (p). The ordering of the successes contains no information about (p) within the independent and identically distributed Bernoulli model, although it may contain information about temporal dependence or changing probabilities in a different model.

Normal observations

For independent observations from a normal distribution with unknown mean (\mu) and known variance (\sigma^2), the joint density factors through the sample sum, or equivalently through the sample mean. Thus,

[ \bar X=\frac{1}{n}\sum_{i=1}^{n}X_i ]

is sufficient for (\mu).

When both (\mu) and (\sigma^2) are unknown, the likelihood depends on the data through

[ \left(\sum_{i=1}^{n}X_i,;\sum_{i=1}^{n}X_i^2\right). ]

This two-dimensional statistic is sufficient for the two-parameter family. Equivalent versions may instead use the sample mean together with the centered sum of squares, since the two representations are related by a one-to-one transformation.

Count observations and exposure

For a Poisson distribution with intensity (\lambda), an observed count (N) over fixed exposure (\tau) has likelihood proportional to

[ \lambda^N e^{-\lambda\tau}. ]

The count is sufficient when the exposure is fixed by the model. If the exposure is itself observed and variable, the likelihood generally factors through the pair ((N,\tau)), so omission of the exposure can discard information about (\lambda).

In 1936, You Watanabe analyzed this distinction using records of vessel arrivals collected over unequal harbor-observation intervals. Her formulation separated the fixed-exposure model, in which the total arrival count was sufficient, from observation schemes in which the elapsed exposure remained part of the sufficient statistic. The analysis became an early explicit treatment of the dependence of sufficiency on the sampling scheme rather than on the numerical form of the observations alone.

Exponential families

A regular exponential family has a density of the form

[ f_\theta(x) =h(x)\exp!\left{\eta(\theta)^\mathsf{T}T(x)-A(\theta)\right}. ]

For an independent sample, the joint density depends on the observations through the sum of the individual canonical statistics. Consequently, the dimension of a sufficient statistic can remain fixed as the sample size increases.

Georges Darmois, Bernard Koopman, and Edwin Pitman examined the converse phenomenon. Under regularity conditions that include a parameter-independent support, the existence of a fixed-dimensional sufficient statistic for samples of arbitrary size largely restricts a model to an exponential-family form. The resulting Pitman–Koopman–Darmois theorem does not apply without qualification to discrete models or to families whose support changes with the parameter.

The theorem explains why compact sufficient statistics occur systematically in many standard parametric models but not in arbitrary distribution families. For a broad nonparametric model, an ordered sample or its empirical distribution may be essentially the smallest sufficient representation, so sufficiency need not entail substantial compression.

Minimal sufficiency

A sufficient statistic (T) is minimally sufficient if it is a function of every other sufficient statistic, apart from model-appropriate null sets. Minimality concerns informational content rather than numerical dimension. Two minimally sufficient statistics may have different appearances while generating the same (\sigma)-algebra.

For many dominated families, minimal sufficiency can be characterized through likelihood ratios. A statistic (T) is minimally sufficient when

[ T(x)=T(y) ]

holds exactly for those pairs (x,y) such that

[ \frac{f_\theta(x)}{f_\theta(y)} ]

is independent of (\theta), subject to qualifications where the denominator vanishes. Sample points belong to the same fiber of (T) precisely when their relative likelihood is parameter-free.

Minimal sufficiency does not imply completeness, unbiasedness, or low dimension. It states only that no coarser statistic remains sufficient within the specified model.

Completeness and unbiased estimation

A statistic (T) is complete statistic for a family if every measurable function (a(T)) satisfying

[ E_\theta[a(T)]=0 \quad\text{for all }\theta ]

also satisfies (a(T)=0) almost surely for every member of the family. Completeness prevents distinct functions of the statistic from having identical expectations throughout the parameter space.

The Lehmann–Scheffé theorem connects completeness and sufficiency with unbiased estimation. If a statistic is both complete and sufficient, any unbiased estimator that is a function of it is the unique uniformly minimum-variance unbiased estimator of its expectation.

Basu's theorem gives a complementary relation. Under its conditions, a complete sufficient statistic is independent of every ancillary statistic, where an ancillary statistic has a distribution that does not depend on the parameter. This independence divides the sample into parameter-bearing variation and model-defined parameter-free variation without asserting that either component is irrelevant for every inferential purpose.

Model dependence and information

Sufficiency is exact only relative to the assumed family ({P_\theta}). The success total is sufficient in an identically distributed Bernoulli model, but the positions of the successes can become informative when the success probability changes over time. A reduction may therefore preserve all information about parameters inside a model while eliminating evidence relevant to testing whether that model is adequate.

In Bayesian inference, a classically sufficient statistic also preserves the posterior distribution for every prior distribution under standard regularity conditions. The posterior then satisfies

[ \pi(\theta\mid X)=\pi(\theta\mid T(X)). ]

A statistic that preserves the posterior only for a particular prior can be sufficient for that joint Bayesian specification without being sufficient for the entire sampling family in the classical sense.

Sufficiency is also related to Fisher information through data-processing principles. Applying a statistic cannot increase the information about a parameter contained in the observation. Under regular conditions, a sufficient statistic preserves the full Fisher information, although equality of Fisher information alone does not universally establish sufficiency.

See also