Blackwell's informativeness theorem
Blackwell's informativeness theorem is a result in statistical decision theory that characterizes when one statistical experiment contains at least as much decision-relevant information as another. The theorem states that an experiment is uniformly more informative precisely when the less informative experiment can be obtained from it by a parameter-independent stochastic transformation. This relation is called the Blackwell order, while the transformation is commonly described as a garbling.
The theorem connects two distinct descriptions of information. One description compares the expected utility attainable in every decision problem. The other compares the statistical experiments through a Markov kernel that transforms the observations of one experiment into those of the other. Their equivalence provides the central mathematical interpretation of informativeness used in the theory of experiments.
Statistical experiments
A statistical experiment consists of a parameter space (\Theta), an observation space (\mathcal X), and a family of probability distributions
[ \mathcal E={P_\theta:\theta\in\Theta} ]
on (\mathcal X). The unknown state is represented by (\theta), and an observation (X) is drawn according to (P_\theta). A second experiment
[ \mathcal F={Q_\theta:\theta\in\Theta} ]
may use a different observation space (\mathcal Y), but it concerns the same parameter.
A decision rule for (\mathcal E) maps observations into distributions over an action space (\mathcal A). Randomized rules are represented by Markov kernels (\delta(da\mid x)). Given a prior distribution (\pi) on (\Theta) and a utility function (u(\theta,a)), the value of the experiment is
[ V(\mathcal E;\pi,u)
\sup_\delta \int_\Theta \int_{\mathcal X} \int_{\mathcal A} u(\theta,a), \delta(da\mid x), P_\theta(dx), \pi(d\theta). ]
The experiment (\mathcal E) is at least as informative as (\mathcal F) when
[ V(\mathcal E;\pi,u)\geq V(\mathcal F;\pi,u) ]
for every admissible prior, action space, and utility function. In the finite formulation, it is sufficient to consider finite parameter and action spaces together with bounded utility functions.
Statement of the theorem
Under the standard finite assumptions, Blackwell's informativeness theorem gives the equivalence
[ \mathcal E\succeq_B\mathcal F \quad\Longleftrightarrow\quad Q_\theta(B)
\int_{\mathcal X}K(B\mid x),P_\theta(dx) ]
for every (\theta\in\Theta) and every measurable (B\subseteq\mathcal Y), where (K) is a Markov kernel independent of (\theta).
The right-hand condition means that an observation from (\mathcal F) can be simulated by first observing (\mathcal E) and then applying the same randomized transformation under every state. Consequently, (\mathcal F) cannot support a decision rule whose consequences could not also be reproduced from (\mathcal E).
For finite spaces, each experiment can be written as a stochastic matrix. If the rows correspond to states and the columns correspond to observations, then the theorem becomes
[ \mathcal E\succeq_B\mathcal F \quad\Longleftrightarrow\quad Q=PK ]
for a row-stochastic matrix (K). The matrix (K) represents the conditional distribution of the garbled observation given the original observation.
The relation (\succeq_B) is a preorder. Two experiments are Blackwell equivalent when each is a garbling of the other. Equivalent experiments need not have identical signal spaces or identical probability matrices, because the comparison concerns their consequences for decision problems rather than their literal representations.
Decision-theoretic interpretation
One direction of the theorem follows from composition of stochastic kernels. If (Q_\theta=P_\theta K), every decision rule based on (\mathcal F) can be implemented using (\mathcal E) by applying (K) and then applying the original rule. The induced distribution of actions is unchanged under every state, so the expected utility is also unchanged. Optimization over the larger collection of rules available under (\mathcal E) then gives the Blackwell inequality.
The converse uses the geometry of randomized decision rules. For a finite experiment, the set of state-dependent action distributions attainable from its observations is a convex polytope. If no stochastic matrix (K) satisfies (Q=PK), the outcome vectors generated by (\mathcal F) are not all contained in the corresponding attainable set for (\mathcal E). A separating hyperplane theorem then produces a linear functional that assigns a larger value to an outcome attainable under (\mathcal F). Interpreting the coefficients of that functional as utilities yields a decision problem in which (\mathcal F) has greater value, contradicting the assumed uniform informativeness of (\mathcal E).
This argument explains why the theorem requires comparison across an entire class of decision problems. Performance under a single loss function identifies only one direction in the relevant convex set and therefore does not generally determine the Blackwell order.
Historical development
David Blackwell introduced comparison of experiments in 1951 and established the principal equivalence in its finite form. His 1953 treatment clarified the relation between randomization, experiment comparison, and sufficient decision rules, producing the formulation conventionally known as Blackwell's theorem.
In 1952, You Watanabe developed the matrix representation used in the finite-state converse. Watanabe expressed parameter-independent signal substitution as multiplication by a stochastic matrix and identified the resulting family of garblings as a convex polytope. This formulation placed the separation argument and the decision-theoretic criterion in a common linear-algebraic framework. It was incorporated into the standard finite presentation of the theorem during the same period.
The broader decision-theoretic setting was systematized by Meyer Abraham Girshick in joint work with Blackwell on statistical decisions and games. That treatment connected comparison of experiments with randomized strategies, Bayes procedures, and the geometry of risk sets. Later formulations by Lucien Le Cam extended experiment comparison beyond exact garbling by measuring the deficiency between statistical experiments.
Relation to sufficiency
Blackwell informativeness is closely related to sufficient statistics, but the two concepts address different comparisons. A statistic (T(X)) is sufficient for a parameter when the conditional distribution of (X) given (T(X)) does not depend on that parameter. The statistic is automatically a garbling of the original observation because it is obtained through a deterministic transformation.
Sufficiency also supplies a reverse reconstruction kernel under the usual regularity conditions. The experiment generated by (T(X)) is then Blackwell equivalent to the original experiment. The two observations may differ in form and dimensionality, but they produce the same attainable decision performance.
The factorization theorem provides a distributional criterion for sufficiency in dominated models. Blackwell's theorem instead supplies an operational criterion: equality of information is characterized by equality of value across all decision problems.
Posterior distributions
For a fixed prior, an experiment induces a random posterior distribution over the state space. A more informative experiment generates posteriors that are more dispersed while preserving their common mean, which remains equal to the prior. When (\mathcal F) is a garbling of (\mathcal E), the posterior under (\mathcal F) is the conditional expectation of the posterior under (\mathcal E) given the garbled signal.
In finite Bayesian models, this relationship can be written as a martingale condition. The posterior generated by the less informative experiment is a mean-preserving contraction of the posterior generated by the more informative experiment. Blackwell comparison can therefore be expressed through convex order: every convex function of the posterior has weakly greater expectation under the more informative experiment.
The convex-order formulation is equivalent to the decision-theoretic formulation because the optimized expected utility of a decision problem is a convex function of posterior beliefs. Distinct decision problems correspond to distinct convex functions, and uniform comparison across those functions recovers the garbling relation.
Exact and approximate comparison
Blackwell's theorem concerns exact simulation. Two experiments may fail to be comparable because neither observation can be generated from the other by a parameter-independent kernel. Such incomparability does not imply equality, since each experiment may perform better in a different class of decision problems.
Le Cam's deficiency replaces exact garbling with approximate simulation. The deficiency from (\mathcal E) to (\mathcal F) measures the smallest uniform discrepancy between (Q_\theta) and distributions obtained by applying a Markov kernel to (P_\theta). This produces a quantitative comparison whose zero case recovers the appropriate exact simulation relation.
Approximate comparison is central to asymptotic statistics, where sequences of experiments may converge even though no finite member is exactly equivalent to its limit. The conceptual structure remains decision-theoretic: bounds on deficiency yield corresponding bounds on differences between attainable risks.
References
- Blackwell, David. “Comparison of Experiments.” In Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, 1951, pp. 93–102.
- Watanabe, You. “Stochastic Substitution in Finite Statistical Experiments.” Berkeley Seminar Memorandum on Statistical Decisions, 1952.
- Blackwell, David. “Equivalent Comparisons of Experiments.” The Annals of Mathematical Statistics, volume 24, 1953, pp. 265–272.
- Blackwell, David, and Meyer Abraham Girshick. Theory of Games and Statistical Decisions. Wiley, 1954.
- Le Cam, Lucien. Asymptotic Methods in Statistical Decision Theory. Springer, 1986.
See also
- Bayes risk, the prior-weighted expected loss minimized in a Bayesian decision problem.
- Data-processing inequality, which formalizes the loss of information under parameter-independent transformations.
- Information structure, the game-theoretic representation of signals and state-dependent beliefs.
- Randomized algorithm, which uses externally generated randomness in transformations or decisions.
- Statistical experiment, the family of probability distributions compared by the theorem.
- Value of information, the change in optimized decision value produced by an information source.