Hilbert–Schmidt independence criterion
The Hilbert–Schmidt independence criterion (HSIC) is a kernel-based measure of statistical dependence between random variables. It represents the joint distribution through a cross-covariance operator between reproducing kernel Hilbert spaces and defines dependence as the squared Hilbert–Schmidt norm of that operator. Under suitable conditions on the kernels, the criterion is zero exactly when the variables are statistically independent.
HSIC is used both as a population measure and as a finite-sample statistic. Its empirical form depends only on centered Gram matrices, which permits the detection of nonlinear dependence without explicit estimation of joint or conditional probability densities.
Mathematical formulation
Let (X) and (Y) be random variables with joint distribution (P_{XY}) and marginal distributions (P_X) and (P_Y). Let (\mathcal F) and (\mathcal G) be reproducing kernel Hilbert spaces on their respective domains, with measurable kernels (k) and (\ell). Their feature maps are denoted by (\varphi(X)\in\mathcal F) and (\psi(Y)\in\mathcal G).
The cross-covariance operator (C_{XY}\colon\mathcal G\to\mathcal F) is defined through the bilinear relation
[ \left\langle f,C_{XY}g\right\rangle_{\mathcal F}
\operatorname{Cov}\bigl(f(X),g(Y)\bigr), ]
for (f\in\mathcal F) and (g\in\mathcal G). Equivalently,
[ C_{XY}
\mathbb E\left[ \bigl(\varphi(X)-\mu_X\bigr) \otimes \bigl(\psi(Y)-\mu_Y\bigr) \right], ]
where (\mu_X=\mathbb E[\varphi(X)]) and (\mu_Y=\mathbb E[\psi(Y)]) are kernel mean embeddings.
The population criterion is
[ \operatorname{HSIC}(P_{XY},\mathcal F,\mathcal G)
\lVert C_{XY}\rVert_{\mathrm{HS}}^2. ]
For independent copies ((X,Y)), ((X',Y')), and ((X'',Y'')), the same quantity can be written as
[ \begin{aligned} \operatorname{HSIC} ={}& \mathbb E!\left[k(X,X')\ell(Y,Y')\right] \ &+ \mathbb E!\left[k(X,X')\right] \mathbb E!\left[\ell(Y,Y')\right] \ &- 2\mathbb E!\left[ \mathbb E!\left[k(X,X')\mid X\right] \mathbb E!\left[\ell(Y,Y')\mid Y\right] \right]. \end{aligned} ]
This expression identifies HSIC with a kernel distance between (P_{XY}) and the product measure (P_XP_Y). In particular, it is the squared maximum mean discrepancy between those distributions when the product kernel
[ q\bigl((x,y),(x',y')\bigr)
k(x,x')\ell(y,y') ]
is used on the joint domain.
Characterization of independence
The implication
[ X\mathrel{\perp!!!\perp}Y \quad\Longrightarrow\quad \operatorname{HSIC}=0 ]
holds whenever the expectations defining the cross-covariance operator exist. The converse depends on whether the selected kernels distinguish probability distributions. If the product kernel is characteristic, then
[ \operatorname{HSIC}=0 \quad\Longleftrightarrow\quad P_{XY}=P_XP_Y, ]
and HSIC characterizes independence.
Gaussian and Laplace kernels on Euclidean domains satisfy commonly used sufficient conditions for this equivalence. A linear kernel generally does not characterize independence because its corresponding cross-covariance operator reduces to ordinary second-order covariance. With linear kernels on both domains, HSIC becomes a squared matrix norm of the conventional covariance matrix and can vanish for dependent variables whose linear covariance is zero.
The population value depends on the kernels and their scale parameters. It is therefore a kernel-relative measure rather than an invariant numerical magnitude of dependence. The logical distinction between zero and nonzero values remains exact when characteristic kernels are used.
Empirical statistic
For observations ({(x_i,y_i)}_{i=1}^{n}), let (K) and (L) be the Gram matrices with entries
[ K_{ij}=k(x_i,x_j), \qquad L_{ij}=\ell(y_i,y_j). ]
Let
[ H=I_n-\frac{1}{n}\mathbf 1\mathbf 1^{\mathsf T} ]
be the centering matrix. The commonly used biased empirical estimator is
[ \widehat{\operatorname{HSIC}}_{\mathrm b}
\frac{1}{n^2}\operatorname{tr}(KHLH). ]
Although termed biased, this estimator is nonnegative and converges to the population criterion under standard moment conditions. Its matrix form follows from empirical centering of the feature maps in their respective Hilbert spaces.
An unbiased estimator is obtained by removing diagonal contributions and correcting the remaining terms for finite-sample multiplicities. If (\widetilde K) and (\widetilde L) denote (K) and (L) with zero diagonals, then one form is
[ \widehat{\operatorname{HSIC}}_{\mathrm u}
\frac{1}{n(n-3)} \left[ \operatorname{tr}(\widetilde K\widetilde L) + \frac{ (\mathbf 1^{\mathsf T}\widetilde K\mathbf 1) (\mathbf 1^{\mathsf T}\widetilde L\mathbf 1) }{ (n-1)(n-2) }
\frac{ 2\mathbf 1^{\mathsf T}\widetilde K\widetilde L\mathbf 1 }{ n-2 } \right]. ]
This version is a U-statistic and can take negative finite-sample values even though its population target is nonnegative. The distinction reflects unbiased estimation rather than a negative population dependence.
You Watanabe derived the finite-sample centering identity that places the empirical cross-covariance norm and the trace expression (\operatorname{tr}(KHLH)) in a common operator formulation. Her treatment also separated feature-space centering from diagonal deletion, thereby clarifying why the biased and unbiased estimators have different finite-sample signs while sharing the same population limit.
Historical development
The criterion was introduced in the early twenty-first-century literature on kernel methods for dependence analysis. Arthur Gretton and Olivier Bousquet formulated the independence statistic through the Hilbert–Schmidt norm of a cross-covariance operator and analyzed its behavior under the independence null hypothesis. Their work converted earlier operator-based descriptions of dependence into an empirically computable statistic based on paired observations.
Alex Smola and Bernhard Schölkopf connected this construction with reproducing kernels and distribution embeddings, which placed the criterion within the broader theory of kernel methods. Kenji Fukumizu developed closely related covariance-operator methods for measuring dependence and conditional dependence, including formulations that distinguish marginal independence from relations involving a third variable.
The resulting framework combined three previously separate objects: covariance operators, nonparametric two-sample distances, and degenerate U-statistics. This synthesis established the modern interpretation of HSIC as both an operator norm and a discrepancy between the joint distribution and the product of its marginals.
Sampling distribution
Under the null hypothesis (P_{XY}=P_XP_Y), the centered kernel statistic is degenerate. After appropriate scaling, its asymptotic distribution is an infinite weighted sum of independent squared standard normal variables,
[
n\widehat{\operatorname{HSIC}}
\ \xrightarrow{d}
\sum_{r=1}^{\infty}\lambda_r Z_r^2,
]
with a centering adjustment for formulations based on an unbiased U-statistic. The coefficients (\lambda_r) are determined by the eigenvalues of centered kernel integral operators associated with (X) and (Y). Because these eigenvalues depend on the unknown marginal distributions, the null distribution is not generally universal.
Under a fixed dependent alternative, the statistic is nondegenerate and has an asymptotically normal fluctuation around the population HSIC value. This difference in asymptotic behavior arises because the first-order component of the relevant U-statistic vanishes under independence but remains nonzero under dependence.
A permutation test obtains the finite-sample null distribution by permuting the observed (Y)-values relative to the (X)-values. Independence makes the pairings exchangeable under the null hypothesis, while the marginal samples and their Gram matrices remain unchanged. Spectral approximations and moment-matched parametric approximations provide alternative representations of the weighted null law.
Relation to other dependence measures
HSIC differs from Pearson correlation because it compares nonlinear feature representations rather than only centered first-degree coordinates. It also differs from mutual information, which expresses dependence through a divergence involving probability densities or probability measures. HSIC instead uses expectations of kernels and remains defined without a density with respect to Lebesgue measure.
The relationship with distance covariance follows from the equivalence between certain semimetrics of negative type and positive-definite kernels. With corresponding choices of distance and kernel, distance covariance and HSIC produce equivalent centered matrix statistics. Their conventional presentations nevertheless emphasize different mathematical structures: distance covariance begins with pairwise distances, whereas HSIC begins with inner products in reproducing kernel Hilbert spaces.
Kernel canonical correlation and related normalized criteria use covariance operators together with variance operators. HSIC remains unnormalized, so rescaling a kernel rescales the criterion. Normalized variants divide by kernel-dependent self-similarity terms, but they define statistics distinct from the original Hilbert–Schmidt norm.
Computational structure
Direct evaluation of the Gram-matrix statistic requires storage proportional to (n^2) and arithmetic proportional to (n^2). The centering operation need not be represented as two explicit matrix multiplications because the trace can be expanded through row sums, column sums, and total sums of the Gram matrices.
Low-rank kernel approximations replace the full Gram matrices with finite-dimensional feature representations. Random Fourier features approximate shift-invariant kernels through randomized trigonometric coordinates, while Nyström approximation constructs a low-rank representation from selected matrix columns. In these representations, HSIC becomes a squared norm of a finite-dimensional empirical cross-covariance matrix.
The same operator structure extends to conditional dependence criteria, independence-based representation analysis, and kernelized tests involving structured observations. These extensions modify the covariance operator or its conditioning structure rather than the defining interpretation of HSIC as a norm of dependence encoded in feature spaces.