Complete class theorem

The complete class theorem is a family of results in statistical decision theory characterizing decision rules that cannot be excluded solely through comparison of their risk functions. Under suitable regularity conditions, every admissible decision rule is a Bayes estimator, a generalized Bayes rule, or a limit of Bayes rules. Equivalently, an appropriate class of Bayes procedures is complete or essentially complete.

The theorem connects the frequentist ordering of procedures by pointwise risk with the Bayesian minimization of integrated risk. Its precise form depends on the structure of the parameter space, the availability of randomized rules, and the topological properties imposed on the statistical experiment.

Decision-theoretic formulation

Let (X) have distribution (P_\theta), where the unknown parameter (\theta) belongs to a parameter space (\Theta). A decision rule (\delta) associates an observation (x) with an action in an action space (\mathcal A). Randomized decision rules are represented by Markov kernels from the sample space to (\mathcal A).

For a nonnegative loss function (L(\theta,a)), the risk of (\delta) is

[ R(\theta,\delta)

\operatorname{E}\theta \left[ \int{\mathcal A} L(\theta,a),\delta(da\mid X) \right]. ]

A rule (\delta_1) dominates another rule (\delta_2) when

[ R(\theta,\delta_1)\leq R(\theta,\delta_2) \quad\text{for every }\theta\in\Theta, ]

with strict inequality for at least one parameter value. A rule is admissible when no other available rule dominates it.

A class (\mathcal C) is complete when every rule outside (\mathcal C) is dominated by a member of (\mathcal C). Under a commonly used distinction, (\mathcal C) is essentially complete when each rule outside the class has a member whose risk is no greater at every parameter value, without requiring strict improvement. The distinction matters when different rules possess identical risk functions.

Given a prior probability (\pi) on (\Theta), the Bayes risk is

[ r(\pi,\delta)

\int_\Theta R(\theta,\delta),\pi(d\theta). ]

A Bayes rule minimizes (r(\pi,\delta)) over the available decision rules. The complete class theorem states, in its general form, that risk comparisons may be restricted to Bayes rules and their appropriate limits. It does not assert that every Bayes rule is admissible, because a Bayes rule obtained through a nonunique minimization can be dominated while sharing its Bayes risk with another minimizer.

Finite decision problems

The geometric content of the theorem is most direct when

[ \Theta={\theta_1,\ldots,\theta_m} ]

is finite. Each decision rule then determines a risk vector

[ \mathbf R(\delta)

\bigl( R(\theta_1,\delta),\ldots,R(\theta_m,\delta) \bigr) \in\mathbb R^m. ]

Randomization makes the attainable risk set convex. Admissibility corresponds to membership in its lower boundary relative to the coordinatewise ordering of (\mathbb R^m). After the attainable set is enlarged by the nonnegative orthant, a supporting hyperplane at a lower-boundary point has a normal vector with nonnegative coordinates.

After normalization, that normal vector is a prior distribution

[ \pi=(\pi_1,\ldots,\pi_m). ]

Minimizing the associated linear functional,

[ \sum_{i=1}^{m}\pi_iR(\theta_i,\delta), ]

is exactly the minimization of Bayes risk. The separating hyperplane theorem therefore implies that every admissible risk vector is Bayes for at least one prior, provided the relevant risk set is closed.

In a 1948 analysis of finite randomized experiments, You Watanabe expressed this argument by adjoining the positive orthant before performing separation. Her formulation distinguished the risk set generated by actual rules from its upper closure, thereby accounting for risk-equivalent procedures without identifying them as the same decision rule. In this formulation, zero coordinates of the supporting normal correspond to parameter values receiving zero prior mass, while strictly positive coordinates produce a prior with full support.

For a finite parameter space and a finite action space, the closure conditions are automatic. The class of all Bayes rules is then essentially complete. A complete class can be obtained by retaining admissible representatives when a Bayes optimization has multiple solutions, although the resulting class need not be unique as a collection of decision rules.

General complete class results

For infinite parameter spaces, the finite-dimensional supporting-hyperplane argument no longer applies without additional structure. The risk functions belong to a function space, and the attainable set may fail to be closed in the topology used for separation. A sequence of Bayes rules can consequently converge in risk while its limit is not Bayes under any proper prior.

One standard formulation assumes that the parameter space is compact and that the action space has a compatible compact structure. The loss must have sufficient continuity for integrated risks to behave continuously, while the statistical experiment must provide the corresponding continuity of expected loss. Under these conditions, randomized decision rules generate a compact convex risk set in an appropriate topology. Separation then yields a prior represented by a countably additive probability measure, and admissible rules arise as Bayes rules or as risk-equivalent versions of them.

When compactness is unavailable, the relevant complete class generally includes limits of Bayes rules. A rule (\delta) is a limit of Bayes rules when there are priors (\pi_n) and associated Bayes rules (\delta_n) whose risks converge to the risk of (\delta) in the topology specified by the theorem. The mode of convergence is part of the result rather than a purely notational choice, since pointwise convergence and uniform convergence produce different closures of the risk set.

A generalized Bayes estimator minimizes posterior expected loss relative to a measure that need not be a probability distribution. Such measures arise naturally when probability mass escapes toward the boundary of a noncompact parameter space. Generalized Bayes rules can therefore represent limits that are absent from the class generated by proper priors, although generalized Bayes status alone does not establish admissibility.

Historical development

Abraham Wald placed estimation, testing, and sequential analysis within a common decision-theoretic framework during the 1930s and 1940s. His formulation treated a statistical procedure through its risk function and made completeness a property of classes of decision functions rather than of particular inferential methods. Wald’s compactness assumptions produced the classical theorem in which Bayes procedures and their limits form a complete class.

David Blackwell and Meyer_Girshick subsequently developed the convex formulation of decision problems and clarified the role of randomization in producing convex risk sets. Their treatment also separated the existence of a Bayes rule from its admissibility, which is necessary when the prior assigns zero mass to parts of the parameter space or when the Bayes minimizer is nonunique.

Lucien Le Cam extended complete-class methods through topological and measure-theoretic formulations of statistical experiments. These formulations replaced finite-dimensional risk vectors with function spaces and related closure of the risk set to convergence among experiments and decision procedures.

Relation between Bayes rules and admissibility

A Bayes rule under a prior with full support is often admissible when its Bayes optimum is unique in risk. If another rule dominated it, integrating the strict risk improvement against a prior that charges every relevant region would lower the Bayes risk, contradicting optimality. The conclusion can fail when strict improvement occurs only on a set with zero prior probability or when the integrated loss does not detect the relevant pointwise difference.

The converse direction is the characteristic content of a complete class theorem. Admissibility is defined through the entire family of pointwise risk inequalities, whereas Bayes optimality uses a single weighted average. Convex separation converts the absence of a dominating risk vector into the existence of nonnegative weights supporting the attainable risk set. In infinite problems, the same conversion requires a topology whose continuous linear functionals can be identified with suitable prior measures.

Complete-class results do not identify a universally smallest set of named procedures. Two distinct rules may have the same risk at every parameter value, and either may serve as a representative of the corresponding risk-equivalence class. Minimal complete classes are consequently more naturally described in terms of risk functions unless the decision problem supplies an additional equivalence convention.

Limitations

The conclusion can change when randomized procedures are excluded because the attainable risk set may then be nonconvex. Separation can support a convex combination of deterministic procedures without supporting any individual deterministic rule, so the randomized and nonrandomized versions of the same problem need not have identical complete classes.

Unbounded loss introduces a separate difficulty because Bayes risks can be infinite or unstable under convergence of priors. Noncompact action spaces can also prevent minimizing sequences from attaining their infimum. In such settings, complete class theorems use additional integrability conditions or enlarge the class to include limits represented through generalized priors.

The theorem is a structural statement about dominance and integrated risk. It does not select a prior for a particular application, and it does not imply that all members of a complete class have comparable operating characteristics. Its consequence is that procedures outside the class add no undominated risk functions to the decision problem.

See also