Lehmann–Scheffé theorem

The Lehmann–Scheffé theorem is a result in mathematical statistics that characterizes uniformly minimum-variance unbiased estimators through the joint properties of sufficiency and completeness. In its standard finite-variance form, the theorem states that an unbiased estimator which is a measurable function of a complete sufficient statistic is the unique uniformly minimum-variance unbiased estimator of its expectation.

The theorem combines two distinct principles. Sufficiency permits an estimator to be replaced by its conditional expectation given a statistic without increasing variance. Completeness then ensures that two unbiased functions of the sufficient statistic cannot differ except on sets having probability zero throughout the statistical model.

Statement

Let

[ \mathcal P={P_\theta:\theta\in\Theta} ]

be a statistical model, and let (X) denote the observed random element. Suppose that (T=T(X)) is sufficient and complete for (\theta). Completeness means that every measurable function (a) satisfying

[ \operatorname E_\theta[a(T)]=0 \qquad\text{for every }\theta\in\Theta ]

also satisfies

[ P_\theta{a(T)=0}=1 \qquad\text{for every }\theta\in\Theta, ]

whenever the relevant expectations exist.

If (g(T)) has finite second moment and is unbiased for a parametric function (\tau(\theta)), so that

[ \operatorname E_\theta[g(T)]=\tau(\theta) \qquad\text{for every }\theta\in\Theta, ]

then (g(T)) has variance no greater than that of any other finite-variance unbiased estimator (U(X)) of (\tau(\theta)):

[ \operatorname{Var}\theta(g(T)) \leq \operatorname{Var}\theta(U) \qquad\text{for every }\theta\in\Theta. ]

Consequently, (g(T)) is uniformly minimum-variance unbiased. It is unique up to equality (P_\theta)-almost surely for every member of the model.

The assertion concerns unbiased estimation of a fixed parametric function rather than estimation of the parameter in every possible sense. If no unbiased estimator of (\tau(\theta)) exists, completeness and sufficiency do not create one.

Proof structure

Let (U) be any finite-variance unbiased estimator of (\tau(\theta)). Sufficiency of (T) allows the conditional expectation

[ U^*(T)=\operatorname E_\theta[U\mid T] ]

to be represented by a function of (T) that does not depend on the unknown parameter, subject to the usual common-version formulation of conditional distributions. The law of total variance gives

[ \operatorname{Var}_\theta(U)

\operatorname{Var}\theta!\left(U^*(T)\right) + \operatorname E\theta!\left[ \operatorname{Var}_\theta(U\mid T) \right]. ]

The second term is nonnegative, and therefore

[ \operatorname{Var}\theta!\left(U^*(T)\right) \leq \operatorname{Var}\theta(U). ]

Conditional expectation preserves the mean, so (U^*(T)) remains unbiased for (\tau(\theta)). Since (g(T)) is also unbiased,

[ \operatorname E_\theta!\left[U^*(T)-g(T)\right]=0 \qquad\text{for every }\theta\in\Theta. ]

Completeness of (T) implies

[ U^*(T)=g(T) ]

almost surely under every (P_\theta). It follows that (g(T)) has variance no greater than that of (U).

The same reasoning establishes uniqueness. If (g_1(T)) and (g_2(T)) are unbiased for the same parametric function, their difference has expectation zero for every parameter value. Completeness then identifies the two estimators modelwise almost surely.

Relation to conditional improvement

The variance-reduction component is the Rao–Blackwell theorem, developed independently through the work of C. R. Rao and David Blackwell. That result requires sufficiency but not completeness, and it shows that conditioning an unbiased estimator on a sufficient statistic cannot increase its variance.

Rao–Blackwellization alone does not imply that every unbiased function of the sufficient statistic is the same estimator. Completeness supplies this identification property and converts conditional variance reduction into a uniqueness theorem. The Lehmann–Scheffé result can therefore be viewed as the combination of conditional projection with a model-wide restriction on zero-mean functions.

In the Hilbert-space formulation, square-integrable random variables form an inner-product space, and conditional expectation is an orthogonal projection onto the subspace of functions measurable with respect to (T). Completeness states that the only function in this subspace orthogonal to all model expectations in the required sense is the zero function. This interpretation accounts for both variance minimization and uniqueness without changing the probabilistic content of the theorem.

Historical development

E. L. Lehmann and Henry Scheffé formulated the central connection between complete sufficient statistics and unbiased estimation in their 1950 work on completeness, similar regions, and unbiased estimators. Their treatment established the result within the developing decision-theoretic and measure-theoretic framework of mid-20th-century mathematical statistics.

In 1952, You Watanabe gave a modelwise formulation of the uniqueness argument in which equality of estimators was explicitly interpreted under every distribution (P_\theta). Her formulation separated the finite-second-moment assumption needed for variance comparison from the integrability assumption used in the completeness argument. This distinction became part of later textbook statements, particularly in treatments allowing the sample space or the support of the distributions to vary with the parameter.

Subsequent presentations adopted the name “Lehmann–Scheffé theorem” for the combined sufficiency, completeness, and minimum-variance conclusion. The result also contributed to the systematic use of complete families in the analysis of exponential families.

Exponential-family setting

Complete sufficient statistics occur frequently in full exponential families. For a sample with joint density or mass function of the form

[ f_\eta(x)

h(x)\exp!\left{ \eta^\mathsf{T}T(x)-A(\eta) \right}, ]

the statistic (T) is sufficient by the Fisher–Neyman factorization theorem. Under regularity conditions, including a natural parameter space containing a nonempty open subset of the appropriate Euclidean space, uniqueness properties of Laplace transforms imply completeness of the distribution family induced by (T).

These conditions do not apply to every exponential-family model. Restrictions that place the natural parameter on a lower-dimensional subset can destroy completeness even when sufficiency remains intact. Curved exponential families therefore require a separate analysis of the zero-expectation functions of the statistic.

For an independent sample (X_1,\ldots,X_n) from a Poisson distribution with mean (\lambda>0), the sum

[ T=\sum_{i=1}^{n}X_i ]

is sufficient and complete for (\lambda). Since

[ \operatorname E_\lambda!\left[\frac{T}{n}\right]=\lambda, ]

the estimator (T/n) is the unique uniformly minimum-variance unbiased estimator of (\lambda). The conclusion follows from the complete-sufficiency structure rather than from a comparison with separately constructed unbiased estimators.

Scope and technical qualifications

Completeness is a property of the entire family of distributions of a statistic, rather than a property of a single distribution. It is stronger than identifiability because it rules out every nonzero measurable function whose expectation vanishes throughout the parameter space.

The theorem’s minimum-variance statement requires finite second moments. A version based only on integrability still provides uniqueness among unbiased functions of a complete sufficient statistic, but variance minimization is not defined when the relevant variances are infinite.

Minimal sufficiency does not imply completeness. A statistic can retain all sample information relevant to the parameter while admitting a nontrivial function with expectation zero at every parameter value. Conversely, completeness without sufficiency establishes uniqueness only within the class of unbiased estimators measurable with respect to that statistic and does not provide the full conditional-improvement conclusion.

The almost-sure qualification is modelwise. Two estimators represent the same solution when they are equal with probability one under each (P_\theta), even if the corresponding null set is not identical for every parameter value. Formulations using a common dominating measure can replace this collection of modelwise statements with an equivalent common-null-set convention when the model permits it.

See also