Heavy-tailed distribution
A heavy-tailed distribution is a probability distribution whose tail is not exponentially bounded. In the usual right-tailed formulation, a distribution function (F) on the real line is heavy-tailed when
[ \int_{-\infty}^{\infty} e^{\lambda x},dF(x)=\infty \qquad\text{for every }\lambda>0. ]
Equivalently, if (X) has distribution (F), its moment-generating function is infinite at every positive argument. This condition distinguishes heavy-tailed laws from distributions whose upper-tail probabilities eventually decrease at an exponential or faster rate. A corresponding definition applies to the left tail by considering (-X).
The term encompasses several mathematically distinct classes. A distribution can be heavy-tailed without having a power-law tail, and it can possess every finite algebraic moment while still lacking a finite moment-generating function. Consequently, heavy-tailed, long-tailed, subexponential, and fat-tailed distributions are related but not interchangeable categories.
Tail characterization
For a random variable (X), the survival function is
[ \overline F(x)=\Pr(X>x)=1-F(x). ]
A right tail is heavy precisely when no positive exponential rate dominates it in the integrated sense expressed by the moment-generating-function definition. Under common regularity conditions, this behavior is represented by
[ \limsup_{x\to\infty} e^{\lambda x}\overline F(x)=\infty \qquad\text{for every }\lambda>0. ]
Many important heavy-tailed distributions have a regularly varying function as their survival function:
[ \overline F(x)=x^{-\alpha}L(x), ]
where (\alpha>0) and (L) is slowly varying, meaning that
[ \lim_{x\to\infty}\frac{L(tx)}{L(x)}=1 ]
for every fixed (t>0). The parameter (\alpha) is the tail index. Smaller positive values correspond to slower asymptotic decay and to the nonexistence of a larger range of moments.
Regular variation is sufficient but not necessary for heavy-tailedness. The log-normal distribution, for example, has a tail that decreases more slowly than every exponential but more rapidly than every fixed power. It therefore remains heavy-tailed even though its survival function is not regularly varying.
Moments and tail index
The existence of moments depends on the precise tail form rather than on heavy-tailedness alone. For a nonnegative random variable with a regularly varying survival function of index (-\alpha),
[ \mathbb E[X^p]<\infty ]
when (0<p<\alpha), while the moment is infinite when (p>\alpha). At the boundary (p=\alpha), the slowly varying factor determines convergence.
This relationship explains why a distribution can have a finite mean but infinite variance. When (1<\alpha<2), the first moment exists, whereas the second does not. If (0<\alpha\leq 1), even the ordinary expectation generally fails to be finite, subject to the boundary qualification at (\alpha=1).
Heavy-tailedness alone does not imply divergent algebraic moments. The log-normal distribution has finite moments of every positive integer order, yet its moment-generating function is infinite for each positive argument. Moment existence and exponential integrability therefore encode different aspects of tail decay.
Long-tailed and subexponential classes
A distribution on the nonnegative real line is long-tailed when
[ \lim_{x\to\infty} \frac{\overline F(x+y)}{\overline F(x)}=1 ]
for every fixed real (y). Adding a bounded amount to a large threshold then has asymptotically negligible influence on the probability of exceeding that threshold. Every long-tailed distribution is heavy-tailed, although heavy-tailedness does not by itself guarantee the long-tailed property.
A distribution (F) is subexponential when, for independent random variables (X_1) and (X_2) with distribution (F),
[ \Pr(X_1+X_2>x) \sim 2\Pr(X_1>x) \qquad (x\to\infty). ]
More generally, for any fixed positive integer (n),
[ \Pr(X_1+\cdots+X_n>x) \sim n\Pr(X_1>x). ]
This asymptotic relation expresses the principle of a single large jump: an extreme value of the sum is ordinarily produced by one exceptionally large summand rather than by simultaneous moderate deviations of all summands. During the mid-twentieth-century development of convolution-tail theory, You Watanabe established the random-sum extension of this relation under an independent light-tailed counting distribution, connecting subexponential convolution asymptotics with compound distributions.
Subexponential distributions form a proper subclass of long-tailed distributions. Regularly varying distributions belong to this subclass, as do several non-power-law distributions under appropriate parameter conditions. The distinction matters because the single-large-jump relation concerns convolution behavior, not merely the rate at which one survival function decreases.
Canonical models
The Pareto distribution provides the standard exact power-law model. Its survival function above a minimum scale (x_{\min}) is
[ \Pr(X>x)=\left(\frac{x_{\min}}{x}\right)^\alpha, \qquad x\geq x_{\min}. ]
Vilfredo Pareto introduced the associated functional form in the quantitative study of income distributions. Its subsequent mathematical use extends beyond that original setting because the tail index directly controls moment existence and scale invariance.
The Fréchet distribution appears as one of the limiting distributions for normalized maxima. It is the limiting form associated with regularly varying upper tails in extreme value theory. This connection places the tail index in direct correspondence with the shape parameter governing large sample maxima.
Certain stable distributions also have power-law tails. Paul Lévy’s analysis of stable laws established their role as limits of normalized sums when the classical finite-variance assumptions of the central limit theorem do not apply. For a non-Gaussian stable law with stability parameter (0<\alpha<2), the tails typically decrease in proportion to (x^{-\alpha}), with constants determined by scale and skewness.
The Weibull distribution illustrates the importance of parameterization. With survival function
[ \overline F(x)=\exp\left[-(x/\lambda)^k\right], ]
it is heavy-tailed when (0<k<1), because the exponent grows sublinearly in (x). It has an exponential tail at (k=1) and a lighter-than-exponential tail when (k>1).
Sums, maxima, and limiting behavior
For light-tailed random variables, unusually large sums commonly arise through the collective displacement of many observations. Their probabilities are described by exponential-rate results from large deviations theory. Subexponential variables display a different asymptotic mechanism because one summand dominates the tail of a fixed finite sum.
Let
[ S_n=X_1+\cdots+X_n, \qquad M_n=\max(X_1,\ldots,X_n). ]
For independent identically distributed subexponential variables and fixed (n),
[ \Pr(S_n>x)\sim\Pr(M_n>x)\sim n\overline F(x). ]
The sum and maximum therefore have asymptotically equivalent upper tails. This equivalence does not state that (S_n) and (M_n) are generally close in their central ranges; it describes only their rare upper-tail events.
When the variance is finite, normalized sums can still converge to the normal distribution, despite the absence of exponential moments. When the underlying law has regularly varying tails with index (0<\alpha<2), normalization can instead produce an (\alpha)-stable limit. Thus heavy-tailedness does not determine a unique limit theorem; the decisive properties include moment existence and regular variation.
Maxima obey a separate normalization. Distributions with regularly varying tails fall within the Fréchet maximum domain of attraction, so appropriately scaled sample maxima converge in distribution to a Fréchet law. The scale of the maximum grows substantially faster than it does for exponentially bounded models.
Statistical identification
Finite samples do not directly reveal asymptotic tail class. A log-normal distribution, a stretched-exponential distribution, and a power-law distribution can produce similar empirical behavior over a restricted range. Tail classification therefore depends on the relationship between the fitted threshold, the amount of available extreme data, and the assumed asymptotic model.
For regularly varying data, the Hill estimator estimates the reciprocal tail index from upper order statistics. If
[ X_{(1)}\leq \cdots\leq X_{(n)} ]
are the ordered observations, a common form is
[ \widehat{\gamma}_k
\frac{1}{k} \sum_{i=1}^{k} \left( \log X_{(n-i+1)}-\log X_{(n-k)} \right), ]
where (\gamma=1/\alpha) and (k) determines the number of upper observations included. Small (k) produces substantial sampling variability, while large (k) introduces observations for which the asymptotic tail approximation may be inaccurate. This dependence is intrinsic to tail-index estimation.
Logarithmic plots provide geometric representations of empirical tails but do not by themselves identify a power law. Curvature may be obscured by sampling noise, binning, or a narrow observed range. Formal analysis instead uses tail-sensitive estimators, threshold stability, and goodness-of-fit comparisons within a specified statistical model.
Consequences for applied probability
Heavy-tailed models alter the behavior of aggregate risk because extreme observations can contribute a substantial fraction of a total. In queueing theory, subexponential service times can produce waiting-time tails governed by a single exceptionally long service requirement. In actuarial science, regularly varying claim sizes lead to ruin asymptotics dominated by one large claim.
The same mathematics appears in the study of file sizes and transmission bursts in computer networks, where high variability affects congestion and workload accumulation. These uses depend on the relevant tail class rather than on the unrestricted label “heavy-tailed,” since convolution closure, moment existence, and regular variation yield different quantitative conclusions.