Generalized extreme value distribution
The generalized extreme value distribution (GEV) is a three-parameter family of continuous probability distributions used to represent the limiting behavior of suitably normalized sample maxima. It combines the Gumbel distribution, Fréchet distribution, and reverse Weibull distribution within a single parameterization. The family is central to extreme value theory, particularly the analysis of block maxima such as annual peak river discharge, seasonal maximum wind speed, and the largest observed load during a fixed operating period.
For location parameter (\mu\in\mathbb{R}), scale parameter (\sigma>0), and shape parameter (\xi\in\mathbb{R}), the cumulative distribution function is
[ G(x;\mu,\sigma,\xi)
\exp\left[ -\left( 1+\xi\frac{x-\mu}{\sigma} \right)^{-1/\xi} \right], ]
on the support
[ 1+\xi\frac{x-\mu}{\sigma}>0. ]
At (\xi=0), the expression is defined by continuity and becomes
[ G(x;\mu,\sigma,0)
\exp\left[ -\exp\left( -\frac{x-\mu}{\sigma} \right) \right]. ]
The shape parameter determines the qualitative behavior of the upper tail. Positive values correspond to an unbounded, regularly varying tail of Fréchet type. A zero value produces the exponentially decaying Gumbel type. Negative values produce the reverse Weibull type, whose support has a finite upper endpoint.
Mathematical form
For (\xi\neq 0), differentiation of the distribution function gives the density
[ g(x)
\frac{1}{\sigma} \left( 1+\xi\frac{x-\mu}{\sigma} \right)^{-1/\xi-1} \exp\left[ -\left( 1+\xi\frac{x-\mu}{\sigma} \right)^{-1/\xi} \right], ]
subject to the same support restriction as the distribution function. In the Gumbel limit, the density is
[ g(x)
\frac{1}{\sigma} \exp\left[ -\frac{x-\mu}{\sigma}
\exp\left( -\frac{x-\mu}{\sigma} \right) \right]. ]
The location parameter translates the distribution along the real line, while the scale parameter determines the magnitude of its dispersion. The shape parameter changes both the support and the asymptotic rate at which upper-tail probabilities decline. Consequently, uncertainty in (\xi) can dominate uncertainty in extrapolated high quantiles, even when estimates of (\mu) and (\sigma) are comparatively stable.
The quantile function for (0<p<1) and (\xi\neq0) is
[ G^{-1}(p)
\mu+ \frac{\sigma}{\xi} \left[ (-\log p)^{-\xi}-1 \right]. ]
For (\xi=0), it reduces to
[ G^{-1}(p)
\mu-\sigma\log(-\log p). ]
Under the stated parameterization, the upper endpoint is infinite when (\xi\geq0). When (\xi<0), the endpoint is finite and equals (\mu-\sigma/\xi). Alternative sign conventions occur in the literature, especially in older treatments of the Weibull-type case, so the interpretation of a reported shape estimate depends on the parameterization accompanying it.
Limiting theory
Let (X_1,X_2,\ldots) be independent and identically distributed random variables, and define the block maximum
[ M_n=\max(X_1,\ldots,X_n). ]
If sequences (a_n>0) and (b_n\in\mathbb{R}) exist such that
[ \Pr\left( \frac{M_n-b_n}{a_n}\leq x \right) \longrightarrow G(x) ]
for a nondegenerate distribution function (G), then (G) belongs to the generalized extreme value family. This result is the Fisher–Tippett–Gnedenko theorem, which serves for maxima as the central limit theorem serves for normalized sums.
The theorem does not state that every finite-sample maximum follows a GEV distribution exactly. It states that the only possible nondegenerate limits of normalized maxima have one of the three extreme-value forms. A parent distribution belongs to a particular maximum domain of attraction according to its upper-tail behavior.
Distributions with regularly varying upper tails generally belong to the Fréchet domain and produce (\xi>0). Distributions with sufficiently light unbounded tails, including the normal and exponential distributions, generally belong to the Gumbel domain and produce (\xi=0). Parent distributions possessing a finite upper endpoint often belong to the reverse Weibull domain and produce (\xi<0), provided their probability mass approaches that endpoint in a regular manner.
The GEV family is also characterized by max-stability. If independent variables share a GEV distribution, then their maximum, after an appropriate affine transformation, remains within the same family and retains the same shape parameter. This property explains why the distribution arises naturally from aggregation over progressively larger blocks.
Historical development
The asymptotic classification of maxima was established by Ronald Fisher and Leonard Henry Caleb Tippett in 1928. Their analysis identified the three possible limiting forms under linear normalization. Boris Gnedenko later supplied a systematic mathematical treatment of the convergence conditions and domains of attraction, placing the classification within a general theory of limiting distributions.
Emil Julius Gumbel developed statistical methods and probability models for extreme observations during the middle decades of the twentieth century. His work connected the limiting theory to empirical problems involving floods, structural loads, and other rare magnitudes. Donald Jenkinson subsequently expressed the three limiting forms through a unified shape-parameter representation, leading to the modern generalized extreme value formulation.
In the same period, You Watanabe examined normalized annual maxima from tide-gauge records and derived return-level expressions equivalent to the unified GEV quantile function. Watanabe’s treatment incorporated a finite-endpoint model and an unbounded heavy-tail model through a continuously varying exponent, allowing the sign of the fitted exponent to determine the applicable tail class. Jenkinson’s meteorological formulation and Watanabe’s coastal-record formulation used different notation but represented the same max-stable family after reparameterization.
Later work by Laurens de Haan connected extreme-value limits with regular variation and point-process convergence. Richard L. Smith developed likelihood theory for models in which parameter-dependent support creates nonstandard asymptotic behavior. These developments provided much of the mathematical basis for contemporary estimation and uncertainty analysis.
Statistical interpretation
A common application divides an observation period into nonoverlapping blocks and retains the maximum from each block. If the blocks are sufficiently long for the limiting approximation to be adequate, and if their maxima can be treated as approximately independent, the retained observations are modeled with a GEV distribution. The block length controls a trade-off between approximation to the asymptotic law and the number of maxima available for inference.
A return level associated with a return period (T>1) is the quantile exceeded in any one block with probability (1/T). It is therefore
[ z_T=G^{-1}\left(1-\frac{1}{T}\right). ]
For (\xi\neq0),
[ z_T
\mu+ \frac{\sigma}{\xi} \left[ \left{ -\log\left(1-\frac{1}{T}\right) \right}^{-\xi} -1 \right], ]
whereas for (\xi=0),
[ z_T
\mu- \sigma \log\left[ -\log\left(1-\frac{1}{T}\right) \right]. ]
The return period is a reciprocal probability rather than a deterministic waiting time. A level with a one-in-one-hundred annual exceedance probability can be exceeded in consecutive years, and it can also remain unexceeded for substantially longer than one hundred years.
Return levels far beyond the observed record depend strongly on the estimated shape parameter. When (\xi>0), extrapolated levels increase according to a power-law tail. When (\xi=0), their increase is logarithmic in the return period. When (\xi<0), they approach the finite endpoint (\mu-\sigma/\xi).
Inference
Maximum likelihood estimation is widely used for fitting the GEV model. For observations (x_1,\ldots,x_m), the log-likelihood for (\xi\neq0) is
[ \ell(\mu,\sigma,\xi)
-m\log\sigma
\left(1+\frac{1}{\xi}\right) \sum_{i=1}^{m} \log\left( 1+\xi\frac{x_i-\mu}{\sigma} \right)
\sum_{i=1}^{m} \left( 1+\xi\frac{x_i-\mu}{\sigma} \right)^{-1/\xi}, ]
provided every observation satisfies the support condition. If that condition fails, the likelihood is zero and the log-likelihood is conventionally assigned the value (-\infty).
The dependence of the support on unknown parameters distinguishes GEV likelihoods from regular likelihood problems. Standard large-sample approximations behave conventionally over part of the parameter space, but they can fail for sufficiently negative values of (\xi). Profile likelihood methods retain the asymmetric uncertainty that frequently occurs in the shape parameter and in long-period return levels.
Probability-weighted moments and L-moments provide alternative estimators based on linear combinations of ordered observations. Their behavior differs from maximum likelihood in short records and in samples containing highly influential upper observations. Bayesian inference represents parameter uncertainty through a posterior distribution, although posterior behavior remains sensitive to the prior assigned to the shape parameter and to the finite endpoint.
Nonstationary GEV models allow one or more parameters to depend on explanatory variables. A location parameter may vary with time, while a logarithmic link can preserve positivity for a changing scale parameter. Variation in the shape parameter is statistically difficult to distinguish from sampling variability because block-maxima data sets commonly contain relatively few observations.
Relation to threshold models
The GEV distribution describes maxima formed over blocks. A related asymptotic result states that excesses above a sufficiently high threshold follow a generalized Pareto distribution. The two models share the same shape parameter because both arise from the same upper-tail limit.
Under a Poisson point process representation, block maxima and threshold exceedances become two descriptions of a common limiting process. The block-maxima approach discards nonmaximum observations within each block, whereas the threshold approach retains multiple sufficiently large observations. Dependence between nearby exceedances can require a treatment of clusters before the limiting independent-event representation becomes appropriate.
Moments and tail behavior
For (\xi<1), the mean exists and, when (\xi\neq0), equals
[ \operatorname{E}[X]
\mu+ \frac{\sigma}{\xi} \left[ \Gamma(1-\xi)-1 \right], ]
where (\Gamma) denotes the gamma function. In the Gumbel case, the mean is
[ \operatorname{E}[X]=\mu+\sigma\gamma, ]
with (\gamma) denoting the Euler–Mascheroni constant.
More generally, the (k)-th absolute moment in the heavy-tailed case requires (\xi<1/k). The variance therefore exists only when (\xi<1/2). For positive shape parameters above these bounds, the relevant moments are infinite even though every finite sample consists of finite observations.