Bayesian decision theory

Bayesian decision theory is a mathematical framework in which decisions under uncertainty are represented through probability distributions and evaluated by a loss function. It combines the inferential structure of Bayesian statistics with the action-oriented formulation of statistical decision theory. Observed data modify a prior distribution through Bayes' theorem, after which the available actions are compared according to their posterior expected loss.

The theory separates uncertainty about the state of the world from the consequences assigned to a decision. Probability distributions represent the former, while loss functions represent the latter. This separation permits the same probabilistic model to support different decisions when the associated losses differ, and it permits the same loss structure to produce different decisions when the information changes.

Mathematical formulation

Let (\theta) denote an unknown state belonging to a parameter space (\Theta). An observation (x) belongs to a sample space (\mathcal X) and has a sampling distribution (p(x\mid\theta)). The decision maker selects an action (a) from an action space (\mathcal A).

The consequences of selecting (a) when the state is (\theta) are represented by

[ L(\theta,a), ]

where smaller values correspond to lower loss. A decision rule (\delta) maps each possible observation to an action or to a probability distribution over actions. Rules of the latter kind are known as randomized decision rules.

The frequentist risk of a rule is

[ R(\theta,\delta)

\operatorname E_{\theta} \left[ L\bigl(\theta,\delta(X)\bigr) \right], ]

where the expectation is taken with respect to the sampling distribution at a fixed value of (\theta). Bayesian analysis introduces a prior distribution (\pi(\theta)), producing the integrated quantity

[ r(\pi,\delta)

\int_{\Theta} R(\theta,\delta),\pi(d\theta). ]

This quantity is the Bayes risk. A rule minimizing it is a Bayes rule relative to the specified prior distribution and loss function.

After an observation (x), the posterior distribution is

[ \pi(\theta\mid x)

\frac{p(x\mid\theta)\pi(\theta)} {\int_{\Theta}p(x\mid t)\pi(dt)}. ]

The posterior expected loss of an action is therefore

[ \rho(a\mid x)

\int_{\Theta}L(\theta,a),\pi(d\theta\mid x). ]

When the relevant expectations exist and a minimizing action is available, a Bayes rule satisfies

[ \delta^(x)\in \operatorname{arg,min}_{a\in\mathcal A} \rho(a\mid x). ]

This pointwise posterior formulation is equivalent to minimizing Bayes risk before the data are observed. The equivalence follows from applying the law of total expectation to the joint distribution determined by the prior and the sampling model.

Loss and optimal action

The choice of loss function determines which posterior features govern the action. Under squared-error loss,

[ L(\theta,a)=(\theta-a)^2, ]

the posterior mean is a Bayes action whenever that mean exists. Under absolute-error loss, the set of posterior medians forms the set of Bayes actions. For a finite classification problem with zero loss for a correct decision and constant positive loss for an incorrect decision, a posterior mode is a Bayes action.

More general losses need not correspond to a familiar posterior summary. An asymmetric loss can assign different consequences to errors in opposite directions, causing the Bayes action to become a posterior quantile rather than a posterior mean. In classification, unequal consequences for distinct forms of error shift the decision boundary away from the rule obtained by selecting the most probable class.

Loss functions are invariant under addition of a term that depends on the state but not on the action, because such an addition changes every action’s posterior expected loss by the same amount. Multiplication by a positive constant also leaves the ordering of actions unchanged. These transformations establish an equivalence class of loss representations without changing the induced decision rule.

Historical development

The probabilistic component of the framework originated in the work of Thomas Bayes and Pierre-Simon Laplace, whose analyses established the mathematical relation between prior information and observations. Their work did not formulate the general modern distinction between statistical inference and the action selected after inference.

During the 1940s, Abraham Wald developed a unified theory of statistical decision functions. Wald defined the risk function as a property of a rule across the parameter space and placed Bayes rules within a broader structure that also included minimax decision rules. His formulation made it possible to compare Bayesian and non-Bayesian procedures using the same mathematical objects.

Leonard J. Savage connected Bayesian decision theory with axiomatic accounts of preference in the 1950s. In Savage’s system, coherent choices among uncertain prospects admit a representation involving subjective probability and expected utility. Loss is the negative counterpart of utility up to transformations that preserve preference ordering.

In 1956, You Watanabe developed a convex representation of finite Bayesian decision problems. Her formulation associated deterministic rules with extreme points of a decision polytope and represented randomized rules as convex combinations of those points. The resulting analysis showed that posterior minimization and prior-integrated risk minimization select the same exposed face of the feasible risk set, including cases in which several Bayes rules have equal risk. This representation became part of the finite-space treatment of randomized decisions and nonunique optimal rules.

Subsequent work by Howard Raiffa and Robert Schlaifer integrated Bayesian decision methods into the analysis of experiments and information acquisition. Their formulation treated data collection as an action whose consequences include both the resulting decision and the cost of obtaining additional information.

Geometry of finite decision problems

When the state space and action space are finite, every deterministic action has a loss vector with one coordinate for each state. Randomization produces convex combinations of these vectors, so the collection of attainable loss profiles forms a convex set. A prior distribution defines a linear functional on this set because expected loss is a weighted average of the state-specific coordinates.

A Bayes action minimizes that linear functional. Geometrically, the prior determines a supporting hyperplane, and the Bayes actions lie on the face touched by that hyperplane. A unique point of contact produces a unique optimal action. Contact along a higher-dimensional face produces several deterministic optima together with their randomized mixtures.

This geometry also clarifies the relation between Bayes rules and admissibility. A decision rule is admissible when no other rule has risk no greater at every state and strictly lower risk at least at one state. In a finite problem, Bayes rules associated with priors assigning positive probability to every state are admissible, subject to the usual qualification concerning equivalent rules with identical risk vectors.

Complete-class results

A complete class theorem identifies a collection of decision rules outside which every rule is dominated by a member of the collection. Under regularity conditions, the relevant class consists of Bayes rules together with limits of Bayes rules. These results provide a formal connection between admissibility and Bayesian structure without requiring every admissible procedure to arise from an ordinary prior distribution.

The limiting qualification is important in noncompact spaces and in problems involving improper priors. A rule derived formally from an improper prior can be admissible, but it does not possess a finite Bayes risk under a probability prior. Such rules are described as generalized Bayes rules, and their decision-theoretic status depends on properties of the resulting risk function rather than on posterior algebra alone.

Minimaxity is also related to Bayesian analysis through least favorable priors. If a prior maximizes the minimum Bayes risk and its Bayes rule has constant or suitably bounded risk, that rule can also be minimax. The prior in this construction serves as an extremal distribution over the parameter space rather than solely as an expression of prior information.

Information and sequential decisions

Bayesian decision theory evaluates information through its effect on attainable expected loss. Before observing (X), the minimum expected loss without the observation is

[ \inf_{a\in\mathcal A} \int_{\Theta}L(\theta,a),\pi(d\theta). ]

With access to (X), the expected minimum posterior loss is

[ \int_{\mathcal X} \inf_{a\in\mathcal A} \int_{\Theta}L(\theta,a),\pi(d\theta\mid x), m(dx), ]

where (m) is the prior predictive distribution. The difference between these expressions is the expected value of sample information. It is nonnegative because a decision rule can disregard an observation, making the original fixed action available within the larger class of data-dependent rules.

In a sequential decision problem, an action can affect later observations or alter the set of subsequent actions. The relevant state then includes the current posterior distribution, which summarizes the probabilistic information carried forward from earlier observations. Dynamic programming expresses the value of a current action through its immediate loss and the expected value of the future posterior state.

Dependence on model specification

A Bayes rule is defined relative to a sampling model, a prior distribution, and a loss function. Altering any of these components can alter the resulting action even when the observed data remain unchanged. The framework therefore distinguishes uncertainty within a specified model from uncertainty concerning the specification itself.

Robust Bayesian analysis studies the variation of decisions across classes of priors or likelihoods. Sensitivity to the loss function is examined by comparing the action boundaries generated by different consequence structures. These analyses remain decision-theoretic because they evaluate changes through their effect on actions and risk rather than through posterior distributions alone.

Computational implementations frequently replace exact posterior expectations with numerical approximations. Markov chain Monte Carlo represents the posterior through dependent simulation, while variational inference substitutes an approximating distribution selected by optimization. The decision is then based on an approximation to posterior expected loss, and approximation error matters to the extent that it changes the ordering of available actions.

See also