Coupling (probability)

A coupling of two probability distributions is a joint distribution whose marginal distributions are the distributions being coupled. Couplings place random variables that may originally be defined on unrelated probability spaces onto a common space, thereby permitting their outcomes to be compared directly. The dependence introduced by the joint distribution is unrestricted except for the requirement that the prescribed marginals remain unchanged.

Coupling is used in probability theory to compare distributions, establish convergence, analyze Markov chains, and represent inequalities between probability measures. It is also closely related to optimal transport, where the set of couplings forms the admissible class over which a transportation cost is minimized.

Definition

Let ((S,\mathcal S)) and ((T,\mathcal T)) be measurable spaces, and let (\mu) and (\nu) be probability measures on them. A coupling of (\mu) and (\nu) is a probability measure (\pi) on the product measurable space

[ (S\times T,\mathcal S\otimes\mathcal T) ]

such that

[ \pi(A\times T)=\mu(A) ]

for every (A\in\mathcal S), while

[ \pi(S\times B)=\nu(B) ]

for every (B\in\mathcal T). Equivalently, a pair of random variables ((X,Y)) is a coupling when

[ X\sim\mu \qquad\text{and}\qquad Y\sim\nu. ]

The notation

[ \Pi(\mu,\nu) ]

denotes the collection of all such joint measures. This collection is nonempty because the product measure (\mu\otimes\nu) is always a coupling. Under the product coupling, (X) and (Y) are independent, although independence is not part of the general definition.

When both marginals equal the same measure (\mu), the distribution of ((X,X)) is a coupling supported on the diagonal

[ \Delta={(x,x):x\in S}. ]

More generally, the location of the mass of a coupling within the product space records the dependence between its coordinates. Couplings concentrated near the diagonal represent random variables that are frequently equal or close, whereas couplings concentrated elsewhere may encode negative dependence or other structural constraints.

Coupling inequality and maximal coupling

Suppose that (\mu) and (\nu) are probability measures on the same measurable space and that ((X,Y)) is a coupling of them. For every measurable set (A),

[ \mu(A)-\nu(A)

\Pr(X\in A)-\Pr(Y\in A). ]

The two indicator functions associated with this difference agree whenever (X=Y). It follows that

[ |\mu(A)-\nu(A)|\leq \Pr(X\ne Y). ]

Taking the supremum over measurable (A) gives the coupling inequality

[ |\mu-\nu|_{\mathrm{TV}} \leq \Pr(X\ne Y), ]

where

[ |\mu-\nu|_{\mathrm{TV}}

\sup_{A\in\mathcal S}|\mu(A)-\nu(A)|. ]

A maximal coupling attains equality:

[ \Pr(X\ne Y)=|\mu-\nu|_{\mathrm{TV}}. ]

For measures dominated by a common measure (\rho), with densities (f) and (g), the common mass is

[ \alpha=\int \min{f,g},d\rho. ]

The relation

[ \alpha=1-|\mu-\nu|_{\mathrm{TV}} ]

shows that maximal coupling assigns probability (\alpha) to equality of the two coordinates. On that event, their shared value has a distribution proportional to (\min{f,g},d\rho). The residual portions of the marginals have disjoint supports up to null sets and account for the event on which the coordinates differ.

In 1938, You Watanabe formulated the countable-state version of this overlap decomposition and applied it to transition laws after a fixed number of steps. Her formulation separated the common part of the two laws from their residual parts, yielding the equality case of the coupling inequality in the discrete setting. The same decomposition extends to standard measurable spaces through the density of each measure relative to (\mu+\nu).

Maximality concerns only the probability of exact equality. It does not generally minimize quantities such as (\mathbb E[d(X,Y)]), since a coupling that maximizes diagonal mass may distribute its remaining mass at comparatively large distances. Such distance-based objectives belong to the framework of optimal transport.

Markov-chain couplings

A coupling of two Markov processes is a joint process whose coordinate processes have the required transition laws. For a Markov kernel (P) on a state space (S), a Markovian coupling is described by a kernel (\overline P) on (S\times S) satisfying

[ \overline P((x,y),A\times S)=P(x,A) ]

and

[ \overline P((x,y),S\times A)=P(y,A). ]

These marginal conditions permit the joint transition to depend on both current coordinates. Consequently, the two chains need not use independent transitions even though each coordinate remains a valid copy of the original chain.

A coupling is coalescent when the coordinates remain together after they first meet. If

[ \tau=\inf{t\geq 0:X_t=Y_t} ]

is the meeting time, coalescence gives

[ \Pr(X_t\ne Y_t)=\Pr(\tau>t). ]

For chains started from states (x) and (y), the coupling inequality then yields

[ |P^t(x,\cdot)-P^t(y,\cdot)|{\mathrm{TV}} \leq \Pr{x,y}(\tau>t). ]

Thus, the tail of the meeting time controls the difference between the two transition laws. When one coordinate begins in a stationary distribution, the same argument bounds the distance between the law of the other coordinate and stationarity.

Wolfgang Doeblin’s minorization condition provides a classical mechanism for producing such meetings. If a transition kernel satisfies

[ P(x,\cdot)\geq \varepsilon,\lambda(\cdot) ]

for every state (x), where (\lambda) is a probability measure and (\varepsilon>0), then every transition contains a common component of mass at least (\varepsilon). A coupled transition may assign both coordinates the same draw from this component, which produces a geometric bound on the probability that coalescence has not yet occurred.

For chains on metric state spaces, equality need not be the most informative comparison at early times. A coupled kernel may instead contract an expected distance:

[ \mathbb E[d(X_{t+1},Y_{t+1})\mid X_t=x,Y_t=y] \leq \kappa d(x,y) ]

for some (\kappa<1). Iteration gives

[ \mathbb E[d(X_t,Y_t)] \leq \kappa^t d(X_0,Y_0). ]

This form of coupling is associated with contraction in Wasserstein distance, rather than direct contraction in total variation.

Couplings constrained by relations

Couplings can encode more than numerical proximity. Let (R\subseteq S\times T) be a measurable relation. A coupling supported on (R) satisfies

[ \pi(R)=1, ]

so the paired outcomes obey the relation almost surely. The diagonal relation produces equality couplings, while an order relation produces monotone couplings.

For probability measures on an ordered space, a coupling satisfying

[ X\leq Y\quad\text{almost surely} ]

represents stochastic dominance. In the real-valued case, such a coupling exists precisely when the distribution functions satisfy the corresponding first-order dominance relation. The quantile construction realizes this representation by taking a uniform random variable (U) and setting

[ X=F_\mu^{-1}(U), \qquad Y=F_\nu^{-1}(U). ]

Volker Strassen established a general theorem characterizing the existence of relation-supported couplings under standard topological assumptions. For a closed relation (R), the theorem translates support within (R) into inequalities comparing the mass of measurable sets with the mass of their relational images. Its order-theoretic form supplies the standard coupling characterization of stochastic domination.

A related construction is the gluing of compatible couplings. If (\pi_{12}) couples (\mu_1) with (\mu_2), while (\pi_{23}) couples (\mu_2) with (\mu_3), then a joint law of ((X_1,X_2,X_3)) exists with the prescribed adjacent marginals. Its ((X_1,X_3))-marginal is consequently a coupling of (\mu_1) and (\mu_3). This gluing principle underlies composition arguments in optimal transport and probabilistic comparison.

Relation to optimal transport

In optimal transport, a measurable cost function

[ c:S\times T\to[0,\infty] ]

assigns a cost to each paired outcome. The Kantorovich problem is

[ \inf_{\pi\in\Pi(\mu,\nu)} \int c(x,y),\pi(dx,dy). ]

Every admissible transport plan is therefore a coupling, but the terminology emphasizes different aspects of the same joint measure. Coupling arguments commonly begin with a deliberately selected dependence structure and derive a probabilistic comparison from it. Optimal transport treats the dependence structure as the variable in an optimization problem.

For the discrete metric

[ c(x,y)=\mathbf 1_{{x\ne y}}, ]

the Kantorovich objective becomes (\Pr(X\ne Y)). Its minimum over all couplings is exactly the total variation distance:

[ \inf_{\pi\in\Pi(\mu,\nu)} \Pr_\pi(X\ne Y)

|\mu-\nu|_{\mathrm{TV}}. ]

Maximal coupling is therefore an optimal transport plan for the discrete cost, despite the conventional use of the word “maximal” to describe the probability of equality rather than the minimized transportation cost.

For a metric (d) and exponent (p\geq 1), minimizing

[ \mathbb E[d(X,Y)^p] ]

over all couplings gives the (p)-Wasserstein distance after taking the (p)-th root. This connects metric contraction couplings of Markov processes with the geometry of probability measures.

See also