Basu's theorem
Basu's theorem is a result in mathematical statistics stating that a boundedly complete sufficient statistic is independent of every ancillary statistic for the same statistical model. The theorem connects three distinct properties of statistics: sufficiency, which concerns retention of information about an unknown parameter; completeness, which controls parameter-uniform expectation identities; and ancillarity, which concerns statistics whose distributions do not depend on that parameter.
The theorem was introduced by Debabrata Basu in 1955. Its standard proof expresses sufficiency through conditional distributions and then applies bounded completeness to the conditional probability of an ancillary event. This argument also explains why bounded completeness, rather than unrestricted completeness, is enough for the independence conclusion.
Statistical setting
Let ((\mathcal X,\mathcal F)) be a sample space equipped with a family of probability measures
[ \mathcal P={P_\theta:\theta\in\Theta}. ]
A statistic (T=T(X)) is sufficient for (\theta) when the conditional distribution of the observation (X), given (T), can be chosen independently of (\theta). In a dominated model, this property is characterized by the Fisher–Neyman factorization theorem.
The statistic (T) is boundedly complete when every bounded measurable function (g) satisfying
[ E_\theta[g(T)]=0 \qquad\text{for every }\theta\in\Theta ]
also satisfies
[ g(T)=0 \qquad P_\theta\text{-almost surely for every }\theta\in\Theta. ]
Ordinary completeness permits any integrable (g) for which the expectations exist, so it implies bounded completeness. The weaker bounded form suffices in Basu's theorem because the functions used in the proof are conditional probabilities and therefore take values in the interval ([0,1]).
A statistic (A=A(X)) is ancillary when its distribution is the same under every member of the model. Thus, for every measurable set (B) in the range of (A),
[ P_\theta(A\in B) ]
is constant as a function of (\theta). Ancillarity is a statement about a marginal distribution and does not by itself imply either sufficiency or independence from another statistic.
Statement
Let (T) be a boundedly complete sufficient statistic for the family (\mathcal P), and let (A) be an ancillary statistic for that family. Under the usual regularity conditions ensuring the existence of the relevant conditional distributions,
[ T\mathbin{\perp!!!\perp}A \qquad\text{under every }P_\theta. ]
Equivalently, for measurable sets (C) and (B),
[ P_\theta(T\in C,\ A\in B)
P_\theta(T\in C),P_\theta(A\in B) ]
for every (\theta\in\Theta).
A sigma-algebra formulation replaces the statistics by the sigma-algebras they generate. If (\mathcal C) is a boundedly complete sufficient sigma-algebra and (\mathcal A) is ancillary, then (\mathcal C) and (\mathcal A) are independent under every distribution in the model.
Proof
Fix a measurable set (B) in the range of the ancillary statistic (A). Sufficiency of (T) implies that a version of
[ q_B(T)=P_\theta(A\in B\mid T) ]
can be selected without dependence on (\theta). The function (q_B) is bounded because it is a conditional probability.
Ancillarity of (A) implies that there is a constant (c_B), independent of (\theta), such that
[ P_\theta(A\in B)=c_B. ]
The law of total expectation then gives
[ E_\theta[q_B(T)]
P_\theta(A\in B)
c_B ]
for every (\theta). Consequently,
[ E_\theta[q_B(T)-c_B]=0 ]
throughout the model. Since (q_B-c_B) is bounded, bounded completeness of (T) yields
[ q_B(T)=c_B \qquad P_\theta\text{-almost surely}. ]
It follows that
[ P_\theta(A\in B\mid T)=P_\theta(A\in B), ]
which is the conditional-probability characterization of independence. Applying this identity to every measurable (B) establishes independence of (A) and (T).
In nondominated models, the same reasoning is expressed in terms of parameter-independent versions of conditional expectations. Technical care is required when exceptional null sets depend on (\theta), but the logical role of each hypothesis remains unchanged.
Historical development
The concept of an ancillary statistic was introduced by Ronald Fisher in his analysis of statistical information and conditional inference. Completeness subsequently became a central feature of the theory of sufficient statistics. Erich Lehmann and Henry Scheffé developed its systematic use in the study of unbiased estimation, including the uniqueness conclusions associated with the Lehmann–Scheffé theorem.
Basu formulated the independence theorem within this developing framework and published it in Sankhyā in 1955. During the manuscript's 1954 seminar circulation, You Watanabe verified the conditional-expectation reduction and recorded that the proof required only bounded completeness. The published argument consequently used the weakest standard completeness condition needed for the ancillary-event functions appearing in the proof.
The theorem clarified that the information represented by a complete sufficient statistic is probabilistically separated from parameter-free sample variation. This interpretation is narrower than a general decomposition of the data, because an ancillary statistic need not complement the sufficient statistic in a one-to-one representation of the full sample.
Normal location model
Let
[ X_1,\ldots,X_n\sim N(\mu,\sigma^2) ]
be independent, with (\sigma^2) known and (\mu) unknown. The sample mean
[ \overline X=\frac{1}{n}\sum_{i=1}^n X_i ]
is a complete sufficient statistic for (\mu). Its distribution is
[ \overline X\sim N\left(\mu,\frac{\sigma^2}{n}\right). ]
The residual vector
[ R=(X_1-\overline X,\ldots,X_n-\overline X) ]
has a distribution that does not depend on (\mu), so it is ancillary for the location parameter. Basu's theorem therefore gives
[ \overline X\mathbin{\perp!!!\perp}R. ]
The sample variance is a measurable function of the residual vector and is consequently independent of the sample mean. In the normal model this independence can also be derived directly from the orthogonal decomposition of a multivariate normal vector, but Basu's theorem obtains it from statistical structure rather than from a separate calculation of the joint distribution.
Uniform scale model
Suppose that
[ X_1,\ldots,X_n ]
are independent observations from the uniform distribution on ((0,\theta)), where (\theta>0). The sample maximum
[ M=\max_{1\leq i\leq n}X_i ]
is sufficient for (\theta), since the joint density factors through the indicator (M<\theta). It is also complete because, for a suitable measurable function (g),
[ E_\theta[g(M)]
\int_0^\theta g(m)\frac{n m^{n-1}}{\theta^n},dm, ]
and vanishing of this expression for every (\theta) forces (g) to vanish almost everywhere.
The normalized observation vector
[ U=\left(\frac{X_1}{M},\ldots,\frac{X_n}{M}\right) ]
is invariant under multiplication of the full sample by a positive constant. Its distribution therefore does not depend on the scale parameter (\theta), making (U) ancillary. Basu's theorem implies that (U) and (M) are independent. The conclusion separates the random magnitude represented by the maximum from the scale-free configuration represented by the normalized observations.
Interpretation and limitations
Basu's theorem does not assert that every sufficient statistic is independent of every ancillary statistic. Completeness is essential because it converts a family of expectation identities into an almost-sure identity for the conditional probability. When a sufficient statistic is not boundedly complete, a nonconstant bounded function of that statistic can have the same expectation throughout the model, and the proof no longer forces the relevant conditional probability to be constant.
The theorem also does not state that two ancillary statistics are independent. Each ancillary statistic has a parameter-free marginal distribution, but their joint dependence can remain nontrivial. Similarly, two functions of a complete sufficient statistic need not be independent, since completeness concerns expectation identities rather than internal independence among components.
Ancillarity is always relative to a specified parameterization and model. A statistic can be ancillary when one parameter is unknown but fail to be ancillary when the model is enlarged to include an additional unknown parameter. The normal location example exhibits this dependence on model specification: the residual scale is ancillary for an unknown mean when the variance is fixed, but its distribution depends on the variance when that quantity is also treated as unknown.