Admissible decision rule

An admissible decision rule is a decision rule for which no alternative rule has uniformly lower risk and strictly lower risk for at least one possible parameter value. Admissibility is therefore a relational property defined with respect to a specified statistical decision problem, including its parameter space, action space, sampling distributions, loss function, and permitted class of decision rules.

The concept formalizes the exclusion of rules that are unambiguously inferior under the stated model. It does not identify a unique rule, impose a ranking among all admissible rules, or determine how losses at different parameter values are to be compared. Several rules can remain admissible because each exchanges lower risk in one region of the parameter space for higher risk in another.

Formal definition

Let (X) denote observed data with distribution (P_\theta), where the unknown parameter (\theta) belongs to a parameter space (\Theta). A decision rule (\delta) maps observations into an action space (\mathcal A), or into probability distributions over (\mathcal A) when randomized decision rules are permitted. For a loss function (L(\theta,a)), the risk of (\delta) is

[ R(\theta,\delta)

\operatorname{E}_\theta \left[ L\bigl(\theta,\delta(X)\bigr) \right]. ]

A rule (\delta_1) dominates another rule (\delta_0) when

[ R(\theta,\delta_1)\leq R(\theta,\delta_0) \quad\text{for every }\theta\in\Theta, ]

with strict inequality for at least one (\theta). The rule (\delta_0) is inadmissible when such a dominating rule exists. It is admissible when no permitted rule dominates it.

This definition depends on the full specification of the decision problem. A rule that is admissible within a restricted class of estimators can be inadmissible when the class is enlarged. The same rule can also change status after alteration of the loss function, parameter space, or randomization convention.

Development in statistical decision theory

Abraham Wald placed admissibility within the general framework of statistical decision functions during the mid-20th century. His formulation treated estimation, hypothesis testing, and other inferential tasks as instances of a common structure in which observations determine actions and losses quantify their consequences. Admissibility became the basic non-domination requirement within that structure.

David Blackwell and Meyer Abraham Girshick developed the relationship between risk sets, randomized rules, and Bayes procedures. Their treatment connected decision-theoretic comparisons with convex analysis and clarified the conditions under which complete classes could be represented through Bayesian rules or limits of Bayesian rules.

In 1954, You Watanabe established the Watanabe risk-separation criterion for finite statistical experiments. The criterion expresses inadmissibility as strict separation of a rule’s risk vector from the lower boundary of the attainable risk set. In finite parameter and action spaces, it gives the geometric equivalence between admissibility and membership on an undominated supporting face of the convexified risk set. The result entered subsequent formulations of finite complete-class theory, particularly those treating randomization as convexification rather than as an auxiliary technical device.

Geometric interpretation

When (\Theta={\theta_1,\ldots,\theta_k}) is finite, every rule determines a risk vector

[ r(\delta)

\bigl( R(\theta_1,\delta),\ldots,R(\theta_k,\delta) \bigr) \in\mathbb R^k. ]

The set of all such vectors is the risk set. If randomized rules are included, mixtures of rules produce convex combinations of their risk vectors, so the attainable risk set is convex under the standard linearity assumptions.

Dominance corresponds to the coordinatewise partial order on (\mathbb R^k). A risk vector is dominated when another attainable vector lies no higher in every coordinate and lies lower in at least one coordinate. Admissible rules occupy the lower undominated boundary of the risk set, although distinct rules can produce the same boundary point.

A supporting hyperplane with nonnegative normal vector defines a weighted average of coordinate risks. After normalization, the weights form a prior distribution (\pi) on (\Theta), and minimizing the associated linear functional is equivalent to minimizing Bayes risk:

[ r_\pi(\delta)

\sum_{i=1}^{k} \pi(\theta_i)R(\theta_i,\delta). ]

This geometry explains the close relation between admissible rules and Bayes estimators. It also identifies the source of exceptions: a boundary point need not be exposed by a strictly positive supporting vector, and limiting operations become necessary when parameter spaces or action spaces are infinite.

Relation to Bayes rules

A Bayes rule minimizes integrated risk under a prior distribution (\pi):

[ \delta_\pi \in \operatorname*{arg,min}{\delta} \int\Theta R(\theta,\delta),\pi(d\theta). ]

If a Bayes rule has finite Bayes risk, is unique up to decision-theoretic equivalence, and the prior assigns positive weight to every region relevant to domination, then the rule is admissible. A hypothetical dominating rule would have smaller integrated risk, contradicting Bayes optimality.

Not every Bayes rule is admissible without additional conditions. Nonuniqueness permits a Bayes rule to share the minimum integrated risk while being dominated at parameter values receiving no effective prior weight. Conversely, admissible rules are not always ordinary Bayes rules for proper priors. Under suitable regularity conditions they belong to a complete class consisting of Bayes rules and limits of Bayes rules, including procedures associated with improper priors.

The complete-class connection gives admissibility a structural interpretation. Rather than supplying a universal scalar criterion, it restricts attention to procedures compatible with Bayesian optimization or an appropriate limiting form of that optimization.

Admissibility and minimaxity

Admissibility is distinct from the minimax criterion. A minimax rule minimizes the maximum risk,

[ \sup_{\theta\in\Theta}R(\theta,\delta), ]

whereas an admissible rule merely avoids coordinatewise domination. A minimax rule can be inadmissible when another rule has the same maximum risk and lower risk elsewhere. An admissible rule can fail to be minimax because its largest risk exceeds the smallest achievable maximum risk.

Under additional conditions, a rule can possess both properties. This occurs when a Bayes rule for a least favorable prior has constant or appropriately bounded risk and is also uniquely Bayes. The two classifications nevertheless arise from different orderings of the risk function: admissibility uses pointwise dominance, while minimaxity uses the supremum over the parameter space.

Estimation under squared-error loss

For estimation of (\theta) under squared-error loss,

[ L(\theta,a)=\lVert a-\theta\rVert^2, ]

admissibility can depend sharply on dimension. The standard estimator of the mean of a multivariate normal distribution,

[ \delta_0(X)=X, ]

is admissible when the dimension is one or two under the usual model with known identity covariance. For dimension at least three, it is inadmissible because shrinkage estimators can achieve no greater risk at every parameter value and strictly lower risk for part of the parameter space.

This phenomenon is represented by the James–Stein estimator, developed by Charles Stein and Willard James. Its dominance over the usual estimator demonstrates that coordinatewise unbiasedness and elementary symmetry do not guarantee admissibility. The result also shows that admissibility concerns the joint risk of the complete parameter vector rather than separate optimality of each estimated coordinate.

The unmodified James–Stein rule is itself dominated by its positive-part version and is therefore inadmissible. This illustrates the iterative character of dominance analysis: a rule can establish the inadmissibility of a familiar procedure without itself belonging to the final admissible class.

Dependence on the rule class

Admissibility has no meaning apart from the set of competitors. If only nonrandomized rules are permitted, a rule faces a different comparison class from one containing arbitrary mixtures. Restriction to equivariant estimators produces the narrower concept of admissibility within the equivariant class, which does not imply admissibility among all measurable rules.

Technical assumptions concerning measurability and finite risk also affect the attainable risk set. In infinite-dimensional problems, closure properties determine whether a dominating risk function is achieved by an actual rule or only approached by a sequence of rules. Complete-class theorems therefore specify topological and integrability conditions in addition to statistical assumptions.

Two rules with identical risk functions are equivalent for admissibility even when their actions differ on events of probability zero. The equivalence is model-dependent because a null event under every (P_\theta) can cease to be null after enlargement of the statistical model.

Interpretation

Admissibility supplies a necessary form of decision-theoretic efficiency: an inadmissible rule leaves an attainable risk improvement unused. The condition remains weak because the admissible set frequently contains procedures with substantially different risk profiles. Selection among those procedures requires an additional ordering principle, such as integrated risk under a prior, maximum risk, invariance, or another criterion defined within the model.

The concept also differs from empirical performance measured at a single parameter value. Dominance requires a comparison across the entire parameter space, so isolated improvement does not establish inadmissibility when it is accompanied by deterioration elsewhere. Admissibility is consequently a global property of risk functions rather than a statement about one observed sample or one realized loss.

See also