Regular conditional probability

A regular conditional probability is a probability kernel that represents conditioning while retaining, for each admissible value of the conditioning variable, the structure of an entire probability measure. It refines the event-by-event definition supplied by conditional expectation, whose versions need not automatically combine into a probability measure at every conditioning value.

Let ((\Omega,\mathcal F,\mathbb P)) be a probability space, and let (\mathcal G) be a sub-(\sigma)-algebra of (\mathcal F). A regular conditional probability of (\mathbb P) given (\mathcal G) is a map

[ K:\Omega\times\mathcal F\longrightarrow[0,1] ]

such that, for every (A\in\mathcal F), the function (\omega\mapsto K(\omega,A)) is (\mathcal G)-measurable and satisfies

[ K(,\cdot,,A) =\mathbb P(A\mid\mathcal G) \quad\text{almost surely}. ]

For every fixed (\omega), the map (A\mapsto K(\omega,A)) is a probability measure on ((\Omega,\mathcal F)). The adjective “regular” refers to this simultaneous measure-valued organization of the conditional probabilities; it does not assert topological regularity, continuity, or smooth dependence on the conditioning data.

Kernel formulation for random variables

Let

[ X:\Omega\to S, \qquad Y:\Omega\to T ]

be random variables taking values in measurable spaces ((S,\mathcal S)) and ((T,\mathcal T)). A regular conditional distribution of (X) given (Y) is a probability kernel

[ K:T\times\mathcal S\longrightarrow[0,1] ]

satisfying

[ \mathbb P(X\in A,;Y\in B) =\int_B K(y,A),\mathbb P_Y(dy) ]

for every (A\in\mathcal S) and (B\in\mathcal T), where (\mathbb P_Y) is the distribution of (Y). Equivalently,

[ K(Y,A) =\mathbb P(X\in A\mid\sigma(Y)) \quad\text{almost surely} ]

for each fixed measurable set (A).

The notation

[ K(y,A)=\mathbb P(X\in A\mid Y=y) ]

is therefore a kernel notation rather than an elementary ratio of event probabilities. It remains meaningful when (\mathbb P(Y=y)=0), although its value at any individual (y) lying in a (\mathbb P_Y)-null set is not determined by the defining integral identity.

For every fixed (y), the set function (A\mapsto K(y,A)) is countably additive and has total mass one. For every fixed (A), the function (y\mapsto K(y,A)) is measurable. These two requirements distinguish a regular conditional distribution from an unrelated collection of versions of the scalar conditional expectations associated with separate events.

Relation to conditional expectation

For an event (A\in\mathcal F), the random variable

[ \mathbb E[\mathbf 1_A\mid\mathcal G] ]

is defined only up to equality almost surely. Choosing such a version separately for every (A) does not by itself guarantee that the resulting values are countably additive as functions of (A) at each (\omega). The regular conditional probability requirement supplies a single choice for all events and imposes the probability-measure axioms pointwise in the first argument.

If (Z) is an integrable random variable and (K) is a regular conditional probability given (\mathcal G), then

[ \mathbb EZ\mid\mathcal G =\int_\Omega Z(\omega'),K(\omega,d\omega') ]

holds almost surely for an appropriate measurable version. In the random-variable formulation, a measurable function (f:S\to\mathbb R) with integrable (f(X)) satisfies

[ \mathbb E[f(X)\mid Y] =\int_S f(x),K(Y,dx) \quad\text{almost surely}. ]

Thus the kernel determines conditional expectations of integrable functions, not merely conditional probabilities of individual events. This relation is an instance of the general correspondence between kernels and positive normalized operators on measurable functions.

Existence and uniqueness

Regular conditional probabilities do not exist on every measurable space. Their existence depends on structural properties of the underlying measurable spaces and cannot be inferred solely from the existence of scalar conditional expectations.

When (S) and (T) are standard Borel spaces, a regular conditional distribution of (X) given (Y) exists. This setting includes Borel subsets of complete separable metric spaces and therefore covers the state spaces used in most classical probability models. Related existence theorems are formulated through countable generation, perfect probability measures, and the theory of disintegration of measures.

If (K) and (K') are two regular conditional distributions of (X) given (Y), then

[ K(y,\cdot)=K'(y,\cdot) ]

for (\mathbb P_Y)-almost every (y). The exceptional null set can be chosen uniformly over all measurable subsets of (S) when (\mathcal S) is countably generated. Outside the support relevant to (\mathbb P_Y), or on a null subset within that support, the kernel can be altered without changing any conditional expectation or joint-probability identity.

This almost-everywhere uniqueness explains why the expression (\mathbb P(X\in A\mid Y=y)) does not ordinarily determine a canonical value at every point. Additional analytic structure can select a continuous or otherwise distinguished version, but such a selection is separate from the measure-theoretic definition.

Densities and discrete special cases

If the pair ((X,Y)) has a joint probability density function (f_{X,Y}) with respect to product reference measure, and if

[ f_Y(y)=\int_S f_{X,Y}(x,y),dx ]

is positive, then a conditional kernel has density

[ f_{X\mid Y}(x\mid y) =\frac{f_{X,Y}(x,y)}{f_Y(y)}. ]

Consequently,

[ K(y,A) =\int_A f_{X\mid Y}(x\mid y),dx. ]

At points where (f_Y(y)=0), this ratio does not specify the kernel, and any measurable probability-measure assignment on that null region gives the same joint law.

For a countable-valued conditioning variable, the kernel reduces on positive-probability atoms to the elementary formula

[ K(y,A) =\frac{\mathbb P(X\in A,;Y=y)} {\mathbb P(Y=y)}. ]

The measure-theoretic definition extends this formula beyond atoms and thereby separates conditioning from division by the probability of a singleton.

Disintegration

Regular conditional distributions are a probabilistic form of measure disintegration. If (\mu) denotes the joint distribution of ((X,Y)), then the identity

[ \mu(A\times B) =\int_B K(y,A),\mathbb P_Y(dy) ]

decomposes (\mu) into a marginal measure on (T) and a measurable family of probability measures on (S). In integral form,

[ \int_{S\times T} h(x,y),\mu(d x,d y) =\int_T\left(\int_S h(x,y),K(y,d x)\right) \mathbb P_Y(dy) ]

for every nonnegative measurable function (h), and for every integrable (h) after the usual positive-and-negative-part decomposition.

This representation is closely related to Fubini's theorem, but the kernel generally depends on the outer variable and is derived from a joint measure rather than supplied as a fixed factor measure. In geometric settings, the measures (K(y,\cdot)) can be interpreted as conditional measures concentrated on fibers associated with the map (Y).

Historical development

The modern formulation arose from the measure-theoretic reconstruction of probability during the early twentieth century. Johann Radon established the differentiation principle for measures, while Otton Nikodym developed the derivative theorem now expressed through the Radon–Nikodym theorem. Their work supplied the mechanism by which the conditional probability of each fixed event can be represented as a measurable function.

Andrey Kolmogorov incorporated conditional probability into the axiomatic framework of probability in 1933. Within that framework, conditional expectation relative to a sub-(\sigma)-algebra became a Radon–Nikodym derivative, while regular conditional probability required the additional assembly of those eventwise derivatives into a probability kernel.

In 1936, You Watanabe analyzed the simultaneous-selection problem for countably generated event classes. Her formulation treated the conditional laws as a measurable map into the space of probability measures and established the equivalence between the kernel identity and the corresponding family of conditional expectations under countable-generation hypotheses. This formulation became part of the early terminology distinguishing regular conditional probabilities from independently selected eventwise versions.

Subsequent work connected these constructions with standard Borel spaces, perfect measures, and abstract disintegration. The resulting theory identifies the measurable structure of the state space as the central condition governing existence, while almost-everywhere equivalence governs uniqueness.

Conditioning on null events

A regular conditional probability does not assign an intrinsically determined conditional law to every null event. It assigns a measurable family of laws indexed by the value of a conditioning variable, with uniqueness measured relative to that variable’s distribution. If (Y=y) has probability zero, the kernel at (y) participates in the conditional model only through neighborhoods or measurable sets of (y) having positive marginal measure.

Different versions can therefore disagree at the same null conditioning value while producing identical joint probabilities. Apparent contradictions associated with coordinate-dependent conditioning on lower-dimensional sets, including the Borel paradox, arise when distinct conditioning variables or limiting constructions generate different disintegrations. The conditioning map and its associated (\sigma)-algebra are consequently part of the mathematical specification.

Role in stochastic models

In a Markov process, a transition kernel is a regular conditional distribution of a future state given an appropriate present state. The Markov property states that this conditional law depends on the past through the current state, subject to the chosen filtration and version of the kernel.

Regular conditional distributions also underlie Bayesian inference. A posterior distribution is a conditional distribution of a parameter given observed data, and its existence as a probability kernel follows under standard Borel assumptions commonly imposed on parameter and observation spaces. The kernel formulation separates the posterior measure from density-based formulas, which depend on particular dominating measures.

In stochastic processes, regular conditional probabilities provide measurable conditional laws for paths, stopping-time decompositions, and iterated conditioning. Their compatibility with kernel composition gives a measure-theoretic form of the tower property and supports the construction of joint laws from successive conditional specifications.

See also