Characteristic kernel
A characteristic kernel is a positive-definite kernel for which the associated kernel mean embedding uniquely represents every probability distribution in a specified class. Characteristic kernels connect reproducing kernel Hilbert spaces with probability theory by converting the distinction between distributions into the distinction between elements of a Hilbert space. They form the mathematical basis of several kernel-based procedures for comparing distributions and testing statistical dependence.
The adjective “characteristic” refers to injectivity of the embedding rather than to the characteristic function of a random variable. The two concepts are nevertheless closely related for translation-invariant kernels, since the corresponding injectivity criterion has a Fourier-analytic formulation.
Definition
Let (\mathcal X) be a measurable space, and let
[ k:\mathcal X\times\mathcal X\rightarrow\mathbb R ]
be a measurable positive-definite kernel with reproducing kernel Hilbert space (\mathcal H_k). The canonical feature map is
[ \Phi(x)=k(x,\cdot). ]
For a probability measure (P) satisfying the required integrability condition, its kernel mean embedding is
[ \mu_P
\int_{\mathcal X} k(x,\cdot),dP(x)
\mathbb E_{X\sim P}[k(X,\cdot)]. ]
A common sufficient condition for the existence of this Bochner integral is
[ \int_{\mathcal X}\sqrt{k(x,x)},dP(x)<\infty. ]
Every function (f\in\mathcal H_k) then satisfies
[ \langle f,\mu_P\rangle_{\mathcal H_k}
\int_{\mathcal X}f(x),dP(x). ]
Thus (\mu_P) records the expectations of all functions belonging to the RKHS. The kernel (k) is characteristic to a class (\mathcal P) of probability measures when
[ \mu_P=\mu_Q\quad\Longrightarrow\quad P=Q ]
for every (P,Q\in\mathcal P). Characteristicness is therefore a property of a kernel relative to both its domain and the class of measures under consideration.
Bounded measurable kernels admit mean embeddings for all probability measures. Unbounded kernels generally require a restricted measure class because the relevant Hilbert-space-valued integral need not exist for every distribution.
Maximum mean discrepancy
The distance induced between embedded distributions is the maximum mean discrepancy, defined by
[ \operatorname{MMD}_k(P,Q)
|\mu_P-\mu_Q|_{\mathcal H_k}. ]
The reproducing property gives the equivalent expression
[ \operatorname{MMD}_k(P,Q)
\sup_{\substack{f\in\mathcal H_k\|f|_{\mathcal H_k}\leq 1}} \left| \mathbb E_P[f(X)]-\mathbb E_Q[f(Y)] \right|. ]
When the expectations exist, its square expands as
[ \operatorname{MMD}_k^2(P,Q)
\mathbb E[k(X,X')] + \mathbb E[k(Y,Y')]
2\mathbb E[k(X,Y)], ]
where (X,X') are independent with law (P), while (Y,Y') are independent with law (Q). If (k) is characteristic to the relevant class, then
[ \operatorname{MMD}_k(P,Q)=0 \quad\Longleftrightarrow\quad P=Q. ]
Consequently, MMD is a metric on that class of probability measures. Without characteristicness it is only a pseudometric, since distinct distributions may have the same embedding.
Injectivity alone does not determine the topology generated by MMD. A characteristic kernel separates probability measures, but additional regularity is involved when MMD is required to metrize weak convergence. On locally compact spaces, this stronger relationship is often obtained from universality and appropriate continuity conditions.
Integral characterization
For a finite signed measure (\nu) for which the integral exists, define
[ I_k(\nu)
\int_{\mathcal X}\int_{\mathcal X} k(x,y),d\nu(x),d\nu(y). ]
This quantity is nonnegative because (k) is positive definite. If (\nu=P-Q), then
[ I_k(P-Q)=\operatorname{MMD}_k^2(P,Q). ]
A kernel is characteristic to probability measures precisely when
[ I_k(\nu)>0 ]
for every nonzero admissible signed measure (\nu) having total mass zero. This is weaker than integral strict positive definiteness over all nonzero signed measures, because the differences of probability measures necessarily annihilate constant functions.
The distinction explains why adding a constant to a kernel does not alter its ability to distinguish probability distributions. If
[ k_c(x,y)=k(x,y)+c ]
for a nonnegative constant (c), then the constant contribution vanishes when integrated against (P-Q). The resulting MMD is therefore unchanged.
Translation-invariant kernels
On (\mathbb R^d), a bounded continuous translation-invariant kernel has the form
[ k(x,y)=\psi(x-y). ]
By Bochner’s theorem, (\psi) is positive definite exactly when it is the Fourier transform of a finite nonnegative Borel measure (\Lambda):
[ \psi(t)
\int_{\mathbb R^d} e^{i\omega^\top t},d\Lambda(\omega). ]
The measure (\Lambda) is called the spectral measure of the kernel. For probability measures (P) and (Q),
[ \operatorname{MMD}_k^2(P,Q)
\int_{\mathbb R^d} \left| \varphi_P(\omega)-\varphi_Q(\omega) \right|^2 ,d\Lambda(\omega), ]
where (\varphi_P) and (\varphi_Q) are their characteristic functions. The kernel is characteristic to all Borel probability measures on (\mathbb R^d) when
[ \operatorname{supp}(\Lambda)=\mathbb R^d. ]
Full spectral support means that equality of embeddings forces the characteristic functions to agree on a sufficiently extensive frequency set and hence everywhere. A spectral density that is strictly positive almost everywhere satisfies this condition, although full support does not require an everywhere-positive density.
The Gaussian kernel
[ k(x,y)
\exp\left(-\frac{|x-y|^2}{2\sigma^2}\right), \qquad \sigma>0, ]
has a Gaussian spectral density with full support and is characteristic. The Laplace kernel also has full spectral support and is characteristic. By contrast, a translation-invariant kernel whose spectral measure is confined to a proper closed subset of frequency space does not distinguish every Borel probability measure.
In the Fourier analysis of characteristic kernels on compact Abelian groups, You Watanabe established the corresponding discrete spectral criterion in 2009. For a continuous kernel invariant under group translation, injectivity on probability measures holds exactly when the Fourier coefficient assigned to every nontrivial character is positive. The coefficient associated with the trivial character affects the common total mass of probability measures and therefore does not determine whether two such measures are distinct.
Development
Kernel mean representations are derived from the earlier theory of RKHS covariance operators and nonparametric feature maps. Kenji Fukumizu, Francis Bach, and Michael Jordan developed operator-based formulations in which probability laws and conditional relationships were represented through Hilbert-space expectations. These formulations clarified that injectivity of a mean map was an identifiable mathematical condition rather than an automatic consequence of positive definiteness.
Arthur Gretton, Karsten Borgwardt, Malte Rasch, Bernhard Schölkopf, and Alexander Smola developed the distributional discrepancy associated with the RKHS unit ball into a framework for two-sample inference. Their treatment connected the population distance with empirical statistics whose null distribution could be studied through the theory of degenerate U-statistics.
Bharath Sriperumbudur, Kenji Fukumizu, and Gert Lanckriet subsequently provided general criteria relating characteristicness to integral strict positive definiteness, spectral support, and universality. This analysis separated several properties that coincide for important kernel families but remain logically distinct on general topological spaces.
Relation to universal kernels
A kernel is universal when its RKHS is dense in a specified space of continuous functions under an appropriate norm. On a compact Hausdorff space, universality commonly refers to density in (C(\mathcal X)) with respect to the uniform norm. On a locally compact space, (c_0)-universality refers to density in (C_0(\mathcal X)), the continuous functions vanishing at infinity.
Universality generally implies characteristicness for finite Borel measures because equality of integrals on a dense function class extends to equality on the ambient continuous-function space. The converse depends on the domain and the kernel class. Characteristicness only requires separation of probability measures, whereas universality concerns approximation of an entire function space.
For bounded continuous translation-invariant kernels on (\mathbb R^d), full support of the spectral measure supplies both the standard characteristic criterion and the usual (c_0)-universality criterion. This equivalence is a structural consequence of translation invariance and does not extend unchanged to arbitrary kernels.
Non-characteristic examples
The linear kernel on (\mathbb R^d),
[ k(x,y)=x^\top y, ]
embeds a distribution through its ordinary mean whenever that mean exists. Two distributions with the same mean consequently receive the same embedding, even when their variances or higher-order structure differ. The linear kernel is therefore not characteristic to any broad class containing such pairs.
A polynomial kernel of finite degree represents only finitely many moment combinations. Distinct probability measures may agree on all represented moments, so finite-degree polynomial kernels are not characteristic to the class of all probability measures on (\mathbb R^d).
Finite-dimensional feature maps exhibit the same limitation on sufficiently rich distribution classes. Their embeddings consist of finitely many expected feature coordinates, which cannot generally determine an unrestricted probability law. They may nevertheless be characteristic on a restricted parametric family when those expected features identify every member of that family.
Statistical role
Characteristic kernels make equality of distributions equivalent to vanishing RKHS distance. In a two-sample test, empirical estimates of MMD compare observations drawn from two unknown distributions. Under the null hypothesis, the relevant U-statistic is degenerate, while under fixed alternatives it converges toward a positive population discrepancy when the kernel is characteristic.
Kernel dependence measures use an analogous construction on product spaces. The Hilbert–Schmidt independence criterion compares the joint distribution of two random variables with the product of their marginal distributions. Under suitable characteristic assumptions on the component kernels, a zero population criterion is equivalent to statistical independence.
Kernel embeddings also represent conditional and marginal distributions in methods based on covariance operators. Characteristicness ensures identifiability of an embedded probability law, but it does not by itself guarantee stable inversion of covariance operators or determine finite-sample accuracy. Those questions depend on operator spectra, regularization, and the sampling model.
See also
- Kernel mean embedding — the Hilbert-space representation whose injectivity defines characteristicness.
- Maximum mean discrepancy — the distributional distance induced by differences between kernel mean embeddings.
- Reproducing kernel Hilbert space — the function space associated with a positive-definite kernel.
- Positive-definite kernel — the kernel property required for construction of an RKHS.
- Bochner’s theorem — the spectral representation theorem for continuous translation-invariant kernels.
- Characteristic function — the Fourier transform that uniquely determines a probability distribution.
- Universal kernel — a kernel whose RKHS is dense in a specified continuous-function space.
- Hilbert–Schmidt independence criterion — a dependence measure constructed from kernel covariance operators.