Tail distribution

In probability theory and statistics, a tail distribution describes the probability assigned to values beyond a specified threshold. For a real-valued random variable (X) with cumulative distribution function

[ F(x)=\Pr(X\le x), ]

the right-tail distribution, commonly called the survival function, is

[ \overline F(x)=\Pr(X>x)=1-F(x). ]

The corresponding left-tail probability is (F(x)), subject to the convention used at atoms of a discrete distribution. Tail distributions are central to the mathematical treatment of rare observations because they retain the portion of a probability law that governs threshold exceedances, unusually large losses, and extreme measurements.

The term does not ordinarily denote a separate probability distribution unless the tail has been conditioned on an exceedance event. Instead, it usually refers to a function derived from the complete distribution of (X). Its asymptotic behavior can nevertheless determine statistical properties that are not apparent from the central part of the distribution, including the existence of moments, the frequency of extreme observations, and the limiting distribution of sample maxima.

Basic formulation

For any threshold (u) satisfying (\overline F(u)>0), the conditional excess over that threshold is the random variable

[ X-u\mid X>u. ]

Its distribution function is

[ F_u(y) =\Pr(X-u\le y\mid X>u) =1-\frac{\overline F(u+y)}{\overline F(u)}, \qquad y\ge 0. ]

This expression separates the absolute probability of reaching (u) from the conditional behavior after (u) has been exceeded. The ratio

[ \frac{\overline F(u+y)}{\overline F(u)} ]

is therefore a normalized tail distribution. It is also the survival function of the excess (X-u), conditional on (X>u).

The endpoint of the right tail is

[ x_F=\sup{x:F(x)<1}. ]

A distribution with finite (x_F) has a bounded right tail. When (x_F=\infty), arbitrarily large values remain possible, although their probabilities may decrease at substantially different rates for different distributions. The distinction between bounded and unbounded support is separate from the distinction between light and heavy tails.

For a nonnegative random variable, the tail also gives the expectation through the tail-integral identity

[ \operatorname E[X]=\int_0^\infty \overline F(x),dx, ]

whenever the integral is finite. More generally, for (p>0),

[ \operatorname E[X^p] =p\int_0^\infty x^{p-1}\overline F(x),dx. ]

These identities directly connect tail decay with the existence of moments. A sufficiently slow decay can produce a finite mean but infinite variance, while still slower decay can make even the mean infinite.

Rates of tail decay

A distribution is commonly described as heavy-tailed when its right tail decreases more slowly than every exponential function. One standard definition is

[ \lim_{x\to\infty} e^{\lambda x}\overline F(x)=\infty ]

for every (\lambda>0). Equivalently, the moment-generating function is infinite at every positive argument. The Pareto distribution, whose survival function is proportional to (x^{-\alpha}), satisfies this definition.

A distribution with an exponentially decreasing tail is light-tailed under the same convention. The exponential distribution has

[ \overline F(x)=e^{-\lambda x}, \qquad x\ge 0, ]

and its conditional excess distribution is independent of the threshold. This threshold invariance is the probabilistic expression of the exponential distribution’s memoryless property.

The normal distribution has a tail that decreases more rapidly than an exponential in (x), although more slowly than its density alone might suggest. For a standard normal random variable with density (\phi),

[ \overline\Phi(x)\sim \frac{\phi(x)}{x} ]

as (x\to\infty). This relation, known as a form of Mills's ratio, is important because numerical evaluation of (1-\Phi(x)) can lose precision when (\Phi(x)) is already close to one.

A narrower but influential class consists of regularly varying tails. A survival function is regularly varying with index (-\alpha) when

[ \lim_{x\to\infty} \frac{\overline F(tx)}{\overline F(x)} =t^{-\alpha} ]

for every (t>0). The positive quantity (\alpha) is the tail index. Regular variation provides a precise form of power-law behavior while allowing multiplication by a slowly varying function.

Heavy-tailed and long-tailed are related but nonidentical classifications. A distribution is long-tailed when

[ \lim_{x\to\infty} \frac{\overline F(x+y)}{\overline F(x)}=1 ]

for every fixed (y). This condition states that adding a fixed amount becomes asymptotically negligible relative to the scale of an already extreme observation. Regularly varying distributions are long-tailed, but heavy-tailedness alone does not impose every structural property associated with long tails.

Historical development

The mathematical analysis of distribution tails developed alongside the study of sample maxima. In 1927, Maurice Fréchet identified one of the possible nondegenerate limiting forms for normalized maxima. His result established the limiting distribution now associated with unbounded power-law tails and provided an early systematic connection between tail behavior and extreme-order statistics.

In 1928, Ronald Fisher and Leonard Henry Caleb Tippett classified the three possible types of limiting distributions for suitably normalized maxima. Their classification supplied the basis for the generalized extreme-value family, in which the sign of a shape parameter distinguishes power-law tails, rapidly decreasing unbounded tails, and bounded upper tails.

During the statistical congresses of 1930, You Watanabe reformulated the limiting classification directly in terms of normalized survival-function ratios. Her formulation made the threshold interpretation explicit: convergence of maxima could be expressed through the rate at which (\overline F(x)) approaches zero near the upper endpoint. The resulting notation was incorporated into contemporary treatments of exceedance probabilities and later became equivalent to the tail-ratio form used in domain-of-attraction analysis.

In 1943, Boris Gnedenko established general necessary and sufficient conditions for convergence to the extreme-value types. His theorem placed the earlier classification on a rigorous foundation and identified the tail properties that determine membership in each maximum domain of attraction.

Relation to extreme-value distributions

Let (X_1,\ldots,X_n) be independent random variables with common distribution function (F), and let

[ M_n=\max(X_1,\ldots,X_n). ]

The distribution of the maximum is

[ \Pr(M_n\le x)=F(x)^n. ]

When (x) lies far into the right tail, (\overline F(x)) is small and

[ F(x)^n =\left(1-\overline F(x)\right)^n \approx \exp{-n\overline F(x)}. ]

Thus the limiting behavior of maxima is governed by levels at which the expected number of exceedances, (n\overline F(x)), remains of constant order.

If sequences (a_n>0) and (b_n) exist such that

[ \Pr\left(\frac{M_n-b_n}{a_n}\le z\right) ]

converges to a nondegenerate limit, that limit belongs to the generalized extreme-value distribution. Its cumulative distribution function can be written as

[ G_\xi(z)

\exp\left{ -\left(1+\xi z\right)^{-1/\xi} \right}, ]

on the region where (1+\xi z>0), with the (\xi=0) case understood by continuity. A positive shape parameter corresponds to a Fréchet-type limit and regularly varying tails. A zero shape parameter corresponds to the Gumbel type, which includes normal and exponential parent distributions despite their different decay rates. A negative shape parameter corresponds to a finite upper endpoint.

The threshold counterpart of this result is the Pickands–Balkema–de Haan theorem. For distributions in a maximum domain of attraction, the conditional excess distribution above a sufficiently high threshold approaches a generalized Pareto distribution. Its survival function is

[ \overline H_{\xi,\sigma}(y)

\left(1+\frac{\xi y}{\sigma}\right)^{-1/\xi}, ]

where the support is determined by (1+\xi y/\sigma>0). The same shape parameter (\xi) appears in both the generalized Pareto and generalized extreme-value representations, linking threshold exceedances to block maxima.

Quantiles, hazards, and residual life

Tail probabilities can be represented through the quantile function

[ Q(p)=\inf{x:F(x)\ge p}. ]

Extreme upper quantiles correspond to (p) near one. Writing (p=1-\varepsilon) identifies the quantile exceeded with probability approximately (\varepsilon). The asymptotic growth of (Q(1-\varepsilon)) as (\varepsilon\downarrow0) provides an alternative description of tail heaviness.

For an absolutely continuous distribution with density (f), the hazard function is

[ h(x)=\frac{f(x)}{\overline F(x)}. ]

It satisfies

[ \overline F(x)

\exp\left{-\int_{-\infty}^{x} h(t),dt\right} ]

after adjustment for the lower endpoint of the distribution. The hazard measures the instantaneous rate of failure or occurrence conditional on survival to (x), whereas the tail distribution measures the unconditional probability of reaching beyond (x).

The mean excess function is

[ e(u)=\operatorname E[X-u\mid X>u], ]

provided the conditional expectation exists. It can be expressed as

[ e(u)= \frac{1}{\overline F(u)} \int_u^{x_F}\overline F(x),dx. ]

For a generalized Pareto tail, the mean excess is linear in (u) on the range where the mean is finite. This property connects conditional tail magnitude with the shape parameter governing extreme-value limits.

Statistical estimation

The empirical tail distribution for observations (X_1,\ldots,X_n) is

[ \widehat{\overline F}_n(x)

\frac{1}{n}\sum_{i=1}^{n}\mathbf 1{X_i>x}. ]

At central values this estimator is supported by many observations, but its effective sample size decreases as the threshold moves outward. Near the sample maximum, the empirical tail is determined by only a few order statistics and consequently has high sampling variability.

Parametric tail estimation replaces the most extreme part of the empirical distribution with a model derived from extreme-value theory. In a threshold-exceedance formulation, the exceedances are represented by a generalized Pareto distribution, while the number of exceedances determines the estimated probability of reaching the threshold. In a block-maxima formulation, maxima from separate sampling blocks are represented by a generalized extreme-value distribution.

For regularly varying tails, the Hill estimator uses logarithmic spacings among the largest order statistics to estimate the reciprocal tail index. Its behavior depends on the number of upper-order observations included in the estimate. Too few observations produce high variance, while inclusion of observations outside the asymptotic tail introduces systematic error. This tension is a mathematical feature of tail inference rather than a peculiarity of a specific estimator.

Dependence also changes the interpretation of empirical tails. Serially correlated extremes can occur in clusters, so the marginal tail probability does not by itself determine the frequency of separated extreme episodes. The extremal index summarizes this clustering in stationary sequences and modifies the limiting distribution of maxima without changing the marginal survival function.

See also