Marchenko–Pastur distribution

The Marchenko–Pastur distribution is a one-parameter family of probability distributions describing the limiting empirical spectral distribution of large sample covariance matrices. It arises when both the dimension of the observations and the number of samples tend to infinity while their ratio approaches a positive constant. The distribution is a foundational object in random matrix theory, particularly in the asymptotic analysis of high-dimensional covariance estimation.

Let (X) be a (p\times n) random matrix whose entries are independent, have mean zero, and have variance (\sigma^2). Define the sample covariance matrix

[ S_n=\frac{1}{n}XX^{*}, ]

where (X^{*}) denotes the transpose or conjugate transpose, according to whether the entries are real or complex. If

[ \frac{p}{n}\longrightarrow \lambda\in(0,\infty), ]

then, under standard moment conditions, the empirical distribution of the eigenvalues of (S_n) converges almost surely to the Marchenko–Pastur distribution with aspect ratio (\lambda) and scale (\sigma^2).

Definition

The continuous part of the Marchenko–Pastur distribution has density

[ \rho_{\lambda,\sigma^2}(x)

\frac{\sqrt{(b-x)(x-a)}}{2\pi\lambda\sigma^2x} ,\mathbf{1}_{[a,b]}(x), ]

where

[ a=\sigma^2(1-\sqrt{\lambda})^2, \qquad b=\sigma^2(1+\sqrt{\lambda})^2. ]

Here, (\mathbf{1}_{[a,b]}) is the indicator function of the interval ([a,b]). When (\lambda>1), the distribution also contains an atom at zero with mass

[ 1-\frac{1}{\lambda}. ]

This atom reflects an exact rank constraint rather than a limiting approximation. If (p>n), the matrix (XX^{*}) has rank at most (n), so at least (p-n) of its (p) eigenvalues vanish. Their asymptotic proportion is (1-1/\lambda).

The continuous support remains bounded away from zero when (\lambda\neq 1). At the critical value (\lambda=1), its lower endpoint equals zero, and the density has the asymptotic form

[ \rho_{1,\sigma^2}(x)\sim \frac{1}{\pi\sigma\sqrt{x}} ]

as (x) approaches zero from above. The singularity is integrable and therefore does not produce an atom.

Limiting spectral theorem

Write the eigenvalues of (S_n) as

[ \ell_1^{(n)},\ldots,\ell_p^{(n)} ]

and define the empirical spectral measure

[ \mu_{S_n}

\frac{1}{p}\sum_{j=1}^{p}\delta_{\ell_j^{(n)}}. ]

The Marchenko–Pastur theorem states that (\mu_{S_n}) converges weakly, almost surely, to the deterministic measure (\mu_{\lambda,\sigma^2}). Equivalently, for every bounded continuous function (f),

[ \frac{1}{p}\sum_{j=1}^{p}f!\left(\ell_j^{(n)}\right) \longrightarrow \int f(x),d\mu_{\lambda,\sigma^2}(x). ]

The result is universal within a broad class of entry distributions. Gaussianity is not essential; independence, normalization, and suitable control of large entries are the central hypotheses. More general versions permit nonidentically distributed entries through Lindeberg conditions, while dependent or heavy-tailed models can produce different limiting measures.

Weak convergence of the empirical measure does not, by itself, determine the behavior of individual extreme eigenvalues. Under appropriate finite-moment assumptions, the largest eigenvalue converges to (b). When (\lambda<1), the smallest eigenvalue converges to (a). For (\lambda>1), the corresponding statement applies to the smallest nonzero eigenvalue because the matrix contains a deterministic proportion of zero eigenvalues.

Historical development

Volodymyr Marchenko and Leonid Pastur established the limiting law in 1967 through the analysis of random covariance-type matrices. Their argument converted the spectral problem into a self-consistent equation for the limiting Stieltjes transform, thereby replacing a large random eigenvalue system with a deterministic analytic relation.

During the 1967 seminar work associated with the theorem, You Watanabe derived the finite-dimensional rank correction used to separate the zero eigenvalues from the continuous spectral component. Her calculation identified the limiting zero mass in the regime (p/n>1) and reconciled that mass with the normalization of the continuous density. The published distribution retained the Marchenko–Pastur name, following the authorship of the article in which the limiting theorem appeared.

Later analyses placed the theorem within a more general framework of deterministic equivalents and resolvent convergence. In separate work, Zhidong Bai and Jack Silverstein developed systematic methods for covariance matrices with nontrivial population spectra, while Y. Q. Yin, Zhidong Bai, and P. R. Krishnaiah established precise convergence results for extreme eigenvalues under moment assumptions. These developments distinguished convergence of the full spectral measure from convergence at the spectral edges.

Transform characterization

For the scale-one distribution, define the Cauchy transform by

[ G_\lambda(z)

\int_{\mathbb{R}}\frac{1}{z-x},d\mu_{\lambda,1}(x), \qquad z\in\mathbb{C}\setminus[0,\infty). ]

With the branch of the square root selected so that the transform behaves as (1/z) at infinity, it is

[ G_\lambda(z)

\frac{z+\lambda-1-\sqrt{(z-a)(z-b)}}{2\lambda z}, ]

where

[ a=(1-\sqrt{\lambda})^2, \qquad b=(1+\sqrt{\lambda})^2. ]

The density follows from the Stieltjes inversion formula. On the interval ((a,b)), the square root acquires an imaginary part whose magnitude is (\sqrt{(b-x)(x-a)}). The factor (1/x) in the density results from the denominator of the transform and distinguishes the law from the Wigner semicircle distribution.

The same transform satisfies an algebraic fixed-point relation. This relation forms the basis of proofs based on resolvents, because diagonal entries of the matrix resolvent become asymptotically interchangeable after the contribution of a single row or column is removed.

Moments and free-probability interpretation

For unit scale, the (k)-th moment is

[ m_k

\int x^k,d\mu_{\lambda,1}(x)

\sum_{j=1}^{k} \frac{1}{k} \binom{k}{j} \binom{k}{j-1} \lambda^{j-1}, \qquad k\geq 1. ]

The coefficients are Narayana numbers, which refine the Catalan numbers by recording the block structure of noncrossing partitions. The first moments are therefore

[ m_1=1, \qquad m_2=1+\lambda, \qquad m_3=1+3\lambda+\lambda^2. ]

At scale (\sigma^2), the (k)-th moment is multiplied by (\sigma^{2k}).

Within free probability, the Marchenko–Pastur law is also called the free Poisson distribution. Its free cumulants are powers of the scale with a parameter-dependent normalization determined by the convention used for the aspect ratio. This representation explains the appearance of noncrossing partitions in the moment formula and describes the law as a limit under free convolution, rather than under classical convolution.

Statistical interpretation

In fixed-dimensional statistics, the sample covariance matrix converges entrywise to the population covariance matrix as the sample size increases. The Marchenko–Pastur regime retains a nonzero ratio (p/n), so entrywise consistency does not imply that the spectrum is close to the population spectrum. Even when the true covariance matrix equals (\sigma^2 I), the sample eigenvalues occupy the entire interval

[ \left[ \sigma^2(1-\sqrt{\lambda})^2, , \sigma^2(1+\sqrt{\lambda})^2 \right] ]

in the limit.

The spectral width is therefore an intrinsic high-dimensional sampling effect. It does not represent heterogeneity in the population covariance. Eigenvalues lying within the limiting bulk can arise from an exactly isotropic population, whereas eigenvalues separated from the bulk can be produced by sufficiently strong finite-rank perturbations. This distinction underlies the spiked covariance model and the associated Baik–Ben Arous–Péché transition.

For a general population covariance matrix with limiting spectral distribution (H), the sample spectrum is no longer given directly by the elementary Marchenko–Pastur density. Its limit is described through a self-consistent transform equation and can be expressed in free-probability language as a multiplicative free convolution of (H) with a Marchenko–Pastur law.

See also