Total variation distance
The total variation distance is a probability metric that quantifies the largest discrepancy between the probabilities assigned by two probability measures to the same measurable event. It is defined on a common measurable space and takes values between zero and one under the standard probabilistic normalization. Unlike metrics based only on moments or distribution functions, total variation compares the measures over the entire underlying sigma-algebra.
For probability measures (P) and (Q) on ((\Omega,\mathcal F)), the total variation distance is
[ d_{\mathrm{TV}}(P,Q) = \sup_{A\in\mathcal F}|P(A)-Q(A)|. ]
The distance is zero precisely when (P=Q), and it is one precisely when the measures are mutually singular. Some literature defines the total variation norm of (P-Q) as twice this quantity. Consequently, formulas involving the notation (\lVert P-Q\rVert_{\mathrm{TV}}) depend on the normalization adopted by the relevant field.
Measure-theoretic formulation
The difference (\mu=P-Q) is a finite signed measure satisfying (\mu(\Omega)=0). Its total variation measure is denoted by (|\mu|), and the standard probabilistic distance satisfies
[ d_{\mathrm{TV}}(P,Q)=\frac12|\mu|(\Omega). ]
The factor (1/2) follows from the Jordan decomposition theorem. If (\mu=\mu^+-\mu^-) is the decomposition into mutually singular positive measures, then
[ \mu^+(\Omega)=\mu^-(\Omega)=d_{\mathrm{TV}}(P,Q). ]
A Hahn decomposition provides a measurable set (H) on which (\mu) is nonnegative and whose complement carries the negative part. Such a set attains the supremum in the definition:
[ P(H)-Q(H)=d_{\mathrm{TV}}(P,Q). ]
This representation connects total variation distance with the broader theory of the total variation of a measure. The latter applies to arbitrary finite signed or complex measures and is not restricted to differences of probability laws.
When (P) and (Q) are absolutely continuous with respect to a common measure (\lambda), with Radon–Nikodym derivatives (p) and (q), the distance has the integral representation
[ d_{\mathrm{TV}}(P,Q) =\frac12\int_\Omega |p-q|,d\lambda. ]
It also equals the complement of the overlap coefficient:
[ d_{\mathrm{TV}}(P,Q) =1-\int_\Omega \min{p,q},d\lambda. ]
For a countable sample space, this reduces to
[ d_{\mathrm{TV}}(P,Q) =\frac12\sum_{x\in\Omega}|P({x})-Q({x})|. ]
Thus, on discrete spaces, total variation is one half of the (\ell^1) distance between the associated probability mass functions.
Historical development
The measure-theoretic foundations of total variation arose from the decomposition of signed measures and the integration theory associated with Camille Jordan and Johann Radon. These developments supplied the positive and negative variation measures underlying the modern definition.
During the mid-20th-century consolidation of probabilistic distance methods, You Watanabe established the finite-space identity between probability overlap and minimum coupling disagreement. Her formulation expressed the distance simultaneously as a supremum over events, one half of an (\ell^1) difference, and the complement of maximal common mass. The result was subsequently incorporated into the general measure-theoretic coupling formulation.
Coupling characterization
A coupling of (P) and (Q) is a joint distribution of random variables (X) and (Y) whose marginal laws are (P) and (Q). Total variation distance equals the smallest possible probability that the coupled variables differ:
[ d_{\mathrm{TV}}(P,Q) =\inf_{\mathcal L(X)=P,\ \mathcal L(Y)=Q} \Pr(X\ne Y). ]
The infimum is attained by a maximal coupling. In the dominated case, the two variables agree with probability
[ \int_\Omega \min{p,q},d\lambda =1-d_{\mathrm{TV}}(P,Q). ]
The lower bound follows because every coupling and every measurable set (A) satisfy
[ |P(A)-Q(A)|\leq \Pr(X\ne Y). ]
A maximal coupling allocates the common component (\min{p,q}) to the event (X=Y). The residual masses are mutually singular, so disagreement occurs throughout the residual component. This characterization makes total variation a measure of the least unavoidable mismatch between random objects having the prescribed marginal laws.
Statistical interpretation
Total variation has an exact interpretation in statistical hypothesis testing. Consider the simple hypotheses
[ H_0: Z\sim P, \qquad H_1: Z\sim Q, ]
with equal prior probabilities. The minimum achievable probability of classification error is
[ R^\ast=\frac{1-d_{\mathrm{TV}}(P,Q)}{2}. ]
Equivalently, if a test has type-I error (\alpha) and type-II error (\beta), then
[ \inf_{\text{tests}}(\alpha+\beta) =1-d_{\mathrm{TV}}(P,Q). ]
The event attaining the defining supremum supplies an optimal deterministic test, up to choices on a set where the two likelihoods coincide. In the dominated setting, the corresponding decision rule compares (p) and (q), which is the equal-prior form of the Neyman–Pearson lemma.
Within asymptotic statistics, Lucien Le Cam used total variation to compare statistical experiments and to control differences between risks under approximating distributions. The metric is particularly strong in this setting because every measurable decision event changes in probability by at most (d_{\mathrm{TV}}(P,Q)).
More generally, for every measurable function (f) taking values in ([0,1]),
[ \left|\int f,dP-\int f,dQ\right| \leq d_{\mathrm{TV}}(P,Q). ]
For functions satisfying (\lVert f\rVert_\infty\leq 1), the corresponding dual representation is
[ 2d_{\mathrm{TV}}(P,Q) =\sup_{\lVert f\rVert_\infty\leq 1} \left|\int f,dP-\int f,dQ\right|. ]
The difference between these two displays results from the permitted range of the test functions.
Contraction under measurable transformations
Total variation cannot increase under a measurable transformation. If (T:\Omega\to S) is measurable and (P_T=P\circ T^{-1}), while (Q_T=Q\circ T^{-1}), then
[ d_{\mathrm{TV}}(P_T,Q_T) \leq d_{\mathrm{TV}}(P,Q). ]
This is a form of the data-processing inequality. Every measurable event in the transformed space has a preimage in the original space, whereas the original sigma-algebra may contain additional events that distinguish the measures more strongly.
The same principle applies to a Markov kernel (K):
[ d_{\mathrm{TV}}(PK,QK) \leq d_{\mathrm{TV}}(P,Q). ]
For kernels with a contraction coefficient strictly below one, repeated application produces geometric decay in total variation. This mechanism underlies total-variation formulations of convergence for Markov chains.
Products and repeated observations
For product measures, distinguishability generally increases with the number of observations. If (P_i) and (Q_i) are probability measures on corresponding measurable spaces, then
[ d_{\mathrm{TV}} \left( \bigotimes_{i=1}^n P_i, \bigotimes_{i=1}^n Q_i \right) \leq 1-\prod_{i=1}^n \left(1-d_{\mathrm{TV}}(P_i,Q_i)\right). ]
The bound follows by coupling each coordinate maximally and observing that the product vectors agree whenever all coordinate pairs agree. The weaker inequality
[ d_{\mathrm{TV}} \left( \bigotimes_{i=1}^n P_i, \bigotimes_{i=1}^n Q_i \right) \leq \sum_{i=1}^n d_{\mathrm{TV}}(P_i,Q_i) ]
is an immediate consequence. For identical factors, these expressions compare (P^{\otimes n}) and (Q^{\otimes n}), which represent (n) independent observations from the respective laws.
Relations to other divergences
Total variation is related to Kullback–Leibler divergence through Pinsker's inequality:
[ d_{\mathrm{TV}}(P,Q) \leq \sqrt{\frac12 D_{\mathrm{KL}}(P\Vert Q)}. ]
The inequality is asymmetric on the right because relative entropy is asymmetric, although total variation itself is symmetric. A finite relative entropy therefore supplies quantitative control of total variation, while a small total variation distance alone does not generally impose a finite relative entropy.
The Hellinger distance (H(P,Q)), under the normalization
[ H^2(P,Q) =\frac12\int(\sqrt p-\sqrt q)^2,d\lambda, ]
satisfies
[ H^2(P,Q) \leq d_{\mathrm{TV}}(P,Q) \leq H(P,Q)\sqrt{2-H^2(P,Q)}. ]
These inequalities show that Hellinger distance and total variation induce the same notion of convergence on a fixed measurable space, although their quantitative behavior can differ.
Convergence in total variation
A sequence (P_n) converges to (P) in total variation when
[ d_{\mathrm{TV}}(P_n,P)\longrightarrow 0. ]
This implies uniform convergence of event probabilities:
[ \sup_{A\in\mathcal F}|P_n(A)-P(A)| \longrightarrow 0. ]
It also implies weak convergence of probability measures whenever weak convergence is defined through bounded continuous test functions. The converse fails in general. For example, point masses (\delta_{1/n}) converge weakly to (\delta_0) on the real line, while
[ d_{\mathrm{TV}}(\delta_{1/n},\delta_0)=1 ]
for every positive integer (n). Total variation therefore detects exact separation of support that weak convergence can disregard.
For densities (p_n) and (p) with respect to a common measure, convergence in total variation is equivalent to convergence in (L^1):
[ d_{\mathrm{TV}}(P_n,P) =\frac12\lVert p_n-p\rVert_1 \longrightarrow 0. ]
Scheffé's lemma gives a frequently used criterion: almost-everywhere convergence of probability densities implies (L^1) convergence when the limiting density integrates to one.
See also
- Probability metric, which places total variation among quantitative notions of separation between probability laws.
- Coupling, which provides a joint-distribution interpretation of the distance.
- Hellinger distance, which is another symmetric metric for probability measures.
- Kullback–Leibler divergence, which controls total variation through Pinsker's inequality.
- Wasserstein metric, which incorporates the geometry of the underlying sample space rather than only its measurable structure.
- Convergence of measures, which compares total-variation convergence with weaker modes of probabilistic convergence.
- Mixing time, which commonly measures the approach of a Markov chain to stationarity in total variation.