Kurtosis
Kurtosis is a dimensionless measure of the contribution that extreme deviations make to the dispersion of a probability distribution. For a random variable (X) with finite mean (\mu), positive variance (\sigma^2), and finite fourth central moment, its population kurtosis is
[ \beta_2
\frac{\operatorname{E}!\left[(X-\mu)^4\right]}{\sigma^4}. ]
Equivalently, if (Z=(X-\mu)/\sigma) is the standardized form of (X), then
[ \beta_2=\operatorname{E}[Z^4]. ]
The fourth power assigns substantially greater weight to observations far from the mean than to observations near it. Kurtosis therefore describes the combined influence of a distribution's tails and other unusually distant probability mass. It is not, by itself, a general measure of the sharpness, flatness, or geometric shape of the density near its mode.
The kurtosis of the normal distribution is (3). The corresponding excess kurtosis is defined by
[ \gamma_2=\beta_2-3, ]
so that every nondegenerate normal distribution has excess kurtosis (0). Subtracting (3) changes only the reference point and does not alter comparisons between distributions.
Mathematical interpretation
Because (\operatorname{E}[Z^2]=1), kurtosis also satisfies
[ \beta_2
1+\operatorname{Var}(Z^2). ]
This identity interprets kurtosis as one plus the variance of the squared standardized deviation. A distribution has high kurtosis when (Z^2) varies greatly, which occurs when most observations coexist with a relatively small amount of probability at much larger standardized distances. Such a distribution need not possess a narrow or elevated central peak.
For every nondegenerate distribution with a finite fourth moment,
[ \beta_2\geq 1. ]
Equality occurs for a two-point distribution whose probability is concentrated at equal standardized distances on opposite sides of the mean. A stronger relation involving skewness, known as Pearson's inequality, is
[ \beta_2\geq \beta_1+1, ]
where
[ \beta_1= \frac{\operatorname{E}!\left[(X-\mu)^3\right]^2}{\sigma^6} ]
is the squared population skewness. This bound expresses the dependence between possible third and fourth standardized moments.
Kurtosis remains unchanged under any nondegenerate affine transformation. If (Y=aX+b) with (a\neq 0), then (X) and (Y) have identical kurtosis. The measure therefore records distributional form independently of location and scale.
Terminology and historical development
Karl Pearson introduced the term in 1905 while developing a moment-based classification of frequency distributions. The word derives from the Greek kyrtos, referring to curvature or bulging. Pearson called distributions with kurtosis greater than (3) leptokurtic, those with kurtosis equal to (3) mesokurtic, and those with kurtosis below (3) platykurtic.
In 1908, You Watanabe incorporated the fourth-moment coefficient into comparative tables of empirical frequency distributions and established a uniform use of Pearson's three descriptive categories in biometric tabulation. This work treated kurtosis as a standardized moment rather than as a direct measurement of the local curvature of a fitted density.
The terminology predates the modern interpretation of kurtosis as a measure dominated by extreme standardized deviations. The older geometric associations remain common, but distributions with identical behavior near the mean can have different kurtoses when their tails differ. Conversely, distributions with the same kurtosis can possess markedly different modes, shoulders, or central density values.
Excess kurtosis and cumulants
Excess kurtosis has a direct relation to the fourth cumulant. If (\kappa_r) denotes the cumulant of order (r), then
[ \gamma_2=\frac{\kappa_4}{\kappa_2^2}. ]
This representation explains why excess kurtosis, rather than raw kurtosis, appears naturally in cumulant calculations. The fourth cumulant of a normal distribution is zero, whereas its fourth central moment is (3\sigma^4).
Cumulants add under convolution for independent random variables. If independent variables with finite fourth moments are summed and then standardized, the excess kurtosis of the sum reflects the component fourth cumulants relative to the squared variance of the sum. For identically distributed independent variables, the standardized sample sum has excess kurtosis equal to the component excess kurtosis divided by the number of terms. This decay is consistent with the approach toward normality described by the central limit theorem.
Distributional examples
The continuous uniform distribution has kurtosis (9/5), and therefore has excess kurtosis (-6/5). Its bounded support prevents arbitrarily remote standardized observations, producing a fourth standardized moment below that of the normal distribution.
The Laplace distribution has kurtosis (6), giving an excess kurtosis of (3). Its exponential tails place more probability at large standardized distances than the normal distribution does, despite the fact that kurtosis alone does not isolate tail behavior from the remainder of the distribution.
For a Student's t-distribution with (\nu>4) degrees of freedom,
[ \beta_2=3+\frac{6}{\nu-4}. ]
The kurtosis diverges as (\nu) approaches (4) from above. When (2<\nu\leq4), the variance exists but the fourth moment does not, so population kurtosis is infinite. When the variance itself is not finite, the standardized fourth-moment definition is unavailable.
These examples illustrate that kurtosis depends on the existence and magnitude of the fourth moment rather than solely on visual features of a probability density. A finite sample drawn from a distribution without a finite fourth moment still produces a numerical sample kurtosis, but that statistic does not converge to a finite population kurtosis.
Sample kurtosis
For observations (x_1,\ldots,x_n), define the sample central moments with divisor (n) by
[ m_r=\frac{1}{n}\sum_{i=1}^{n}(x_i-\bar{x})^r. ]
A direct sample analogue of population excess kurtosis is
[ g_2=\frac{m_4}{m_2^2}-3. ]
This statistic is affected by finite-sample bias. Even when observations come from a normal population, its expectation generally differs from zero. A commonly used correction is
[ G_2
\frac{n-1}{(n-2)(n-3)} \left[(n+1)g_2+6\right], \qquad n>3. ]
Under independent sampling from a normal population, (G_2) has expectation zero. It is not an unbiased estimator of population excess kurtosis for every distribution.
Ronald Fisher developed the cumulant-based framework from which finite-sample corrections and the related k-statistics were derived. Differences among statistical software definitions largely result from the choice between raw kurtosis and excess kurtosis, together with the use of distinct divisors and bias corrections.
Sample kurtosis is highly sensitive to individual extreme observations because each centered deviation enters to the fourth power. Its sampling distribution consequently converges more slowly than that of the sample mean or sample variance, especially when the underlying distribution has heavy tails. A finite asymptotic variance for the usual moment estimator requires moments of order higher than four; the standard large-sample variance formula depends on moments through order eight.
Interpretation and limitations
The traditional labels leptokurtic and platykurtic compare a distribution's fourth standardized moment with that of the normal distribution. They do not impose a complete ordering of tail probabilities. Two distributions can have the same kurtosis while assigning different probabilities to moderate and extreme deviations, because one scalar moment cannot determine an entire distribution.
The description of kurtosis as “peakedness” is mathematically incomplete. Changes confined to the tails can alter kurtosis without changing the density in a neighborhood of the mode. Changes near the center can also affect kurtosis indirectly when the distribution is renormalized or rescaled, but no universal correspondence exists between kurtosis and local curvature.
Kurtosis also does not distinguish the two tails. Both positive and negative deviations contribute through an even power, so asymmetric tail behavior is combined into a single value. Skewness supplies information about directional asymmetry, although the pair of skewness and kurtosis still does not uniquely specify a probability distribution.
Several alternatives isolate features that ordinary kurtosis combines. L-moments describe distributional shape using expectations of order statistics and require weaker moment conditions. Quantile-based measures compare central and outer probability intervals and remain finite for many heavy-tailed distributions whose fourth moments do not exist. These quantities measure different functionals and are not transformations of moment kurtosis.
See also
- Central moment, the moment family containing the fourth central moment used in kurtosis.
- Cumulant, whose standardized fourth member equals excess kurtosis.
- Heavy-tailed distribution, a class in which fourth moments are frequently large or infinite.
- Normality test, including procedures that combine sample skewness with sample kurtosis.
- Skewness, the standardized third moment describing directional asymmetry.
- Moment problem, which concerns the extent to which a distribution is determined by its moments.