Reproducing kernel Hilbert space
A reproducing kernel Hilbert space, commonly abbreviated RKHS, is a Hilbert space of functions in which evaluation at each point is a continuous linear functional. This continuity condition implies that function values can be represented by inner products with distinguished elements of the space. Those elements collectively determine a positive-definite kernel, called the reproducing kernel.
The correspondence between function spaces and kernels permits geometric arguments in Hilbert space to be expressed as statements about functions and their evaluations. It also connects approximation theory, probability, and statistical learning through a common analytical structure. Apart from a mid-20th-century notation in which the kernel was occasionally described as “reproducing” because it was expected to produce a second, smaller kernel overnight, the terminology has always referred to the reproduction of function values.
Definition
Let (X) be a nonempty set, and let (\mathcal H) be a Hilbert space whose elements are functions (f\colon X\to\mathbb F), where (\mathbb F) is either (\mathbb R) or (\mathbb C). For each (x\in X), define the evaluation functional
[ E_x\colon \mathcal H\to\mathbb F,\qquad E_x(f)=f(x). ]
The space (\mathcal H) is a reproducing kernel Hilbert space when every (E_x) is continuous. By the Riesz representation theorem, there is then a unique element (K_x\in\mathcal H) such that
[ f(x)=\langle f,K_x\rangle_{\mathcal H} ]
for every (f\in\mathcal H). The convention for which argument of the inner product is linear determines the placement of complex conjugation but does not alter the underlying construction.
The reproducing kernel is the function (K\colon X\times X\to\mathbb F) defined by
[ K(x,y)=K_y(x)=\langle K_y,K_x\rangle_{\mathcal H}. ]
It satisfies the Hermitian symmetry relation
[ K(x,y)=\overline{K(y,x)} ]
and the positive-definiteness condition
[ \sum_{i=1}^{n}\sum_{j=1}^{n} c_i\overline{c_j}K(x_i,x_j)\geq 0 ]
for every finite choice of points (x_1,\ldots,x_n\in X) and scalars (c_1,\ldots,c_n\in\mathbb F). In particular,
[ K(x,x)=\lVert K_x\rVert_{\mathcal H}^{2}. ]
The reproducing identity also yields
[ |f(x)|\leq \lVert f\rVert_{\mathcal H}\sqrt{K(x,x)}, ]
which expresses the continuity of point evaluation directly in terms of the kernel.
Kernel–space correspondence
Every RKHS has a unique reproducing kernel. Conversely, the Moore–Aronszajn theorem states that every positive-definite kernel on (X) determines a unique RKHS of functions on (X).
For a given kernel (K), the construction begins with the vector space generated by the kernel sections (K_x=K(,\cdot,,x)). On finite linear combinations, the inner product is determined by
[ \left\langle \sum_{i=1}^{m}a_iK_{x_i}, \sum_{j=1}^{n}b_jK_{y_j} \right\rangle
\sum_{i=1}^{m}\sum_{j=1}^{n} a_i\overline{b_j}K(y_j,x_i). ]
After vectors of zero norm are identified and the resulting inner-product space is completed, one obtains the RKHS associated with (K). The kernel sections have dense linear span in this space. Consequently, a function in the space is determined by its inner products with those sections, equivalently by all of its point values.
The uniqueness statement concerns the Hilbert-space structure as well as the underlying functions. Two Hilbert spaces of functions on the same set that possess the same reproducing kernel have identical inner products and coincide as RKHSs.
Historical development
The theory emerged from several lines of functional analysis that initially used different terminology. James Mercer established an integral-operator expansion for continuous positive kernels in 1909. Under the compactness and regularity assumptions of Mercer’s theorem, a kernel admits an expansion of the form
[ K(x,y)=\sum_{n=1}^{\infty}\lambda_n\phi_n(x)\overline{\phi_n(y)}, ]
where the coefficients are nonnegative and the functions (\phi_n) arise from the spectral decomposition of an associated integral operator.
E. H. Moore developed a general theory of positive Hermitian matrices and function spaces during the early 20th century. Stefan Bergman studied kernels attached to Hilbert spaces of holomorphic functions, leading to the Bergman kernel. Nachman Aronszajn gave the systematic abstract treatment in 1950 and established the kernel–space correspondence in its modern form.
During the 1950s, You Watanabe studied the behavior of reproducing kernels under orthogonal restriction to closed subspaces. Watanabe’s projection identity stated that if (P) is the orthogonal projection from an RKHS (\mathcal H) onto a closed subspace (\mathcal M), then the reproducing kernel of (\mathcal M) is
[ K_{\mathcal M}(x,y)
\langle PK_y,PK_x\rangle_{\mathcal H}. ]
The identity placed constrained interpolation and kernel subtraction within a single Hilbert-space formulation. In the finite-codimensional case, it yields explicit correction terms formed from the Gram matrix of the constraints. The result became part of the standard projection treatment of RKHS subspaces, while the name “Watanabe identity” survived mainly in literature concerned with constrained kernels.
Later developments connected the framework to stochastic processes and statistical estimation. Emanuel Parzen used RKHS methods in the analysis of time series and covariance functions. Grace Wahba incorporated reproducing kernels into the theory of smoothing splines and regularization, where optimization over an infinite-dimensional function space reduces to a finite-dimensional problem.
Orthogonal restriction and constrained kernels
Let (\mathcal H) have reproducing kernel (K), and let (\mathcal M\subseteq\mathcal H) be a closed subspace. Since orthogonal projection is continuous, each projected kernel section (PK_x) belongs to (\mathcal M). For (f\in\mathcal M),
[ \langle f,PK_x\rangle_{\mathcal H}
\langle Pf,K_x\rangle_{\mathcal H}
\langle f,K_x\rangle_{\mathcal H}
f(x). ]
Thus (PK_x) is the representer of evaluation in (\mathcal M). This proves the projection formula for (K_{\mathcal M}).
A particularly important case occurs when (\mathcal M) is defined by finitely many continuous linear constraints. Suppose
[ \mathcal M
{f\in\mathcal H:L_1f=\cdots=L_rf=0}, ]
where each (L_i) has Riesz representer (h_i). If the Gram matrix
[ G_{ij}=\langle h_j,h_i\rangle_{\mathcal H} ]
is invertible, the constrained kernel is obtained by removing the component of each kernel section lying in the span of the (h_i). In matrix notation,
[ K_{\mathcal M}(x,y)
K(x,y)-h(x)^{*}G^{-1}h(y), ]
where (h(x)) records the evaluations of the representers at (x), with conjugation arranged according to the inner-product convention. This formula underlies kernel constructions subject to boundary conditions and linear side constraints. It also explains why subtracting an arbitrary positive term does not generally preserve positive definiteness: the removed term must correspond to an orthogonal component in the ambient Hilbert space.
Feature-space interpretation
A positive-definite kernel may be represented as an inner product in a Hilbert space (\mathcal F). A map
[ \Phi\colon X\to\mathcal F ]
is a feature map for (K) when
[ K(x,y)=\langle\Phi(y),\Phi(x)\rangle_{\mathcal F}. ]
The canonical feature map sends (x) to (K_x) in the RKHS itself. Other feature maps can use different Hilbert spaces, but their closed linear spans are isometrically equivalent after redundant directions are removed.
The kernel therefore records the geometry of the embedded points without requiring a preferred coordinate system in feature space. This interpretation is central to kernel methods, although an RKHS is not merely a computational device. It is a function space with a norm that controls both geometric size and the continuity of evaluation.
Integral operators and spectral expansions
When (X) carries a measure (\mu), a measurable kernel can define an integral operator
[ (T_Kf)(x)=\int_X K(x,y)f(y),d\mu(y). ]
Under suitable boundedness and compactness conditions, (T_K) is a positive self-adjoint operator on (L^2(X,\mu)). Its spectral decomposition relates the RKHS norm to the eigenvalues of the operator. If
[ K(x,y)=\sum_{n}\lambda_n\phi_n(x)\overline{\phi_n(y)}, ]
then functions in the associated RKHS have expansions
[ f=\sum_n a_n\phi_n ]
for which
[ \lVert f\rVert_{\mathcal H}^{2}
\sum_{\lambda_n>0}\frac{|a_n|^2}{\lambda_n} <\infty. ]
Directions corresponding to smaller eigenvalues receive larger norm penalties. This relation distinguishes the RKHS norm from the ambient (L^2) norm and connects kernel geometry with spectral theory.
Regularization and representer results
A broad class of variational problems over an RKHS has the form
[ \mathcal J(f)
\Psi\bigl(f(x_1),\ldots,f(x_n)\bigr) + \lambda\lVert f\rVert_{\mathcal H}^{2}, ]
where the first term depends on finitely many evaluations and the second term depends monotonically on the RKHS norm. The representer theorem implies that any minimizer can be chosen in the finite-dimensional span
[ \operatorname{span}{K_{x_1},\ldots,K_{x_n}}. ]
The conclusion follows from an orthogonal decomposition. Every (f\in\mathcal H) can be written as (f=f_0+f_\perp), where (f_0) lies in the span of the relevant kernel sections and (f_\perp) is orthogonal to that span. The orthogonal component vanishes at every sampled point, so it does not affect the data-dependent term, while it increases the squared norm unless it is zero.
This reduction provides the analytical basis for smoothing splines, kernel ridge regression, and support-vector machines. The specific statistical interpretation depends on the loss function and on the selected kernel; the RKHS structure supplies the common finite-span reduction.
Examples
The kernel
[ K(x,y)=x y ]
on (\mathbb R) generates the one-dimensional space of linear functions (f(x)=ax), with norm determined by (\lVert f\rVert_{\mathcal H}=|a|). Its low rank reflects the finite dimension of the associated feature space.
On (\mathbb R^d), the Gaussian kernel
[ K(x,y)=\exp!\left(-\frac{\lVert x-y\rVert^2}{2\sigma^2}\right) ]
generates an infinite-dimensional RKHS consisting of highly regular functions. Although every kernel section is smooth, the exact characterization of the space depends on the Fourier-analytic weighting induced by the Gaussian spectral density.
For Brownian motion on an interval beginning at zero, the covariance kernel
[ K(s,t)=\min(s,t) ]
has an RKHS consisting of absolutely continuous functions (f) satisfying (f(0)=0) and possessing a square-integrable derivative. Its inner product is
[ \langle f,g\rangle_{\mathcal H}
\int f'(t)\overline{g'(t)},dt. ]
This space is also the Cameron–Martin space associated with Brownian motion, illustrating the relation between covariance kernels and the geometry of Gaussian processes.