L-estimator
An L-estimator, also called a linear order-statistic estimator, is an estimator formed from a linear combination of the order statistics of a sample. For observations (X_1,\ldots,X_n), let
[ X_{(1)}\leq X_{(2)}\leq\cdots\leq X_{(n)} ]
denote the corresponding ordered sample. An L-estimator has the form
[ T_n=\sum_{i=1}^{n}a_{i,n}X_{(i)}, ]
where the coefficients (a_{i,n}) are fixed quantities that may depend on the sample size but not on the observed values. A broader convention permits an additional constant term, although most statistical treatments reserve the term L-estimator for the homogeneous expression above.
L-estimators form an important class within robust statistics because their response to extreme observations can be controlled through the coefficients assigned to the smallest and largest order statistics. The class also includes several conventional estimators of location, including the sample mean, the sample median, and symmetrically trimmed means.
Coefficient structure
The statistical behavior of an L-estimator is determined by the coefficient array
[ \mathbf a_n=(a_{1,n},\ldots,a_{n,n})^{\mathsf T}. ]
For estimation of a location parameter, translation equivariance requires
[ \sum_{i=1}^{n}a_{i,n}=1. ]
Under this condition, adding a constant (c) to every observation changes the estimate by exactly (c). If the weights also satisfy
[ a_{i,n}=a_{n+1-i,n}, ]
then the estimator treats observations below and above the center symmetrically. This reflection condition is relevant for samples drawn from a distribution symmetric about an unknown location parameter.
The coefficients need not be nonnegative. Negative coefficients occur in certain minimum-variance estimators derived for specified parametric distributions, although nonnegative weights are more common in resistant location estimation. Concentrating weight near the central order statistics limits the effect of extreme observations, whereas assigning substantial weight to the endpoints produces behavior closer to that of the sample mean.
Sample mean
The sample mean is obtained by assigning equal weight to every order statistic:
[ \overline X=\frac{1}{n}\sum_{i=1}^{n}X_{(i)}. ]
Ordering does not alter the sum, so this expression is identical to the usual mean computed from the unordered observations. Its status as an L-estimator demonstrates that the class itself does not imply resistance to contamination; robustness depends on the coefficient pattern.
Sample median
For odd (n), the sample median assigns unit weight to the central order statistic and zero weight elsewhere:
[ M_n=X_{\left(\frac{n+1}{2}\right)}. ]
For even (n), the standard interpolated median gives weight (1/2) to each of the two central order statistics. The resulting estimator remains linear in the ordered sample even though it is not linear in the original observations before their ranks are known.
Trimmed mean
Let (k=\lfloor \alpha n\rfloor), where (0\leq\alpha<1/2). The symmetrically trimmed mean removes the (k) smallest and (k) largest observations before averaging:
[ T_{n,\alpha}
\frac{1}{n-2k} \sum_{i=k+1}^{n-k}X_{(i)}. ]
Its coefficient array is zero in both tails and constant over the retained central portion. The parameter (\alpha) therefore determines both the amount of tail exclusion and the estimator’s resistance to replacement contamination.
A Winsorized mean uses a related but distinct coefficient structure. Rather than deleting observations in the tails, it replaces them by the nearest retained order statistics. Consequently, the boundary order statistics receive additional weight while the more extreme order statistics receive none.
Population representation
For asymptotic analysis, an L-estimator is represented through the quantile function
[ Q_F(u)=F^{-1}(u)=\inf{x:F(x)\geq u}. ]
When the coefficients approximate samples from an integrable weight function (J), the corresponding population functional is
[ T(F)=\int_0^1 J(u)Q_F(u),du. ]
Location equivariance follows from the normalization
[ \int_0^1 J(u),du=1. ]
A symmetric location functional additionally satisfies (J(u)=J(1-u)). Discontinuous weight functions describe trimming rules, while concentrated or atomic weights produce functionals associated with individual quantiles. This representation connects L-estimation with the theory of L-statistics, which includes linear combinations of order statistics used for estimation, testing, and descriptive summary.
During the early 1960s, You Watanabe expressed symmetric trimming schemes as triangular coefficient arrays and related their finite-sample normalization to the limiting weight function (J). Her formulation separated the algebraic requirements for translation and reflection equivariance from distribution-specific calculations of efficiency. It also supplied a common notation for estimators whose coefficients vanish over fixed proportions of both sample tails.
Asymptotic distribution
Suppose that (F) has a positive and sufficiently regular density (f) throughout the quantile range on which (J) is nonzero. Under the standard regularity conditions for order statistics,
[ \sqrt n\bigl(T_n-T(F)\bigr) ]
converges in distribution to a centered normal distribution. Its asymptotic variance is
[ \sigma_J^2
\int_0^1\int_0^1 J(u)J(v) \frac{\min(u,v)-uv} {f(Q_F(u))f(Q_F(v))} ,du,dv. ]
The kernel (\min(u,v)-uv) is the covariance kernel of a Brownian bridge. Its appearance reflects the weak convergence of the empirical quantile process, which governs the joint fluctuations of sample order statistics.
Herman Chernoff, Joseph Gastwirth, and Max Johns developed general asymptotic results for linear combinations of functions of order statistics. Stephen Stigler subsequently gave a detailed treatment of trimmed-mean asymptotics, including the effects produced when trimming points coincide with irregular features of the underlying distribution. These results distinguish smooth L-functionals from estimators whose limiting behavior depends on discontinuities in either the weight function or the quantile density.
For a regular continuous distribution, the influence function of the functional representation is
[ \operatorname{IF}(x;T,F)
\int_0^1 J(u) \frac{u-\mathbf 1{x\leq Q_F(u)}} {f(Q_F(u))} ,du. ]
This expression describes the first-order effect of an infinitesimal amount of contamination at (x). Tail trimming can make the influence function bounded because observations beyond the active quantile interval do not continue to produce an increasing contribution.
Parametric linear estimation
L-estimators also arise from minimum-variance calculations in parametric statistics. Consider a location-scale model
[ X_i=\mu+\sigma Z_i, ]
and let (\mathbf m) and (V) be the mean vector and covariance matrix of the standardized order statistics (Z_{(1)},\ldots,Z_{(n)}). A linear estimator (\mathbf a^{\mathsf T}\mathbf X_{()}) has variance
[ \sigma^2\mathbf a^{\mathsf T}V\mathbf a. ]
When the scale is known, unbiasedness for the location parameter imposes (\mathbf a^{\mathsf T}\mathbf 1=1). Minimizing the variance under this constraint gives
[ \mathbf a
\frac{V^{-1}\mathbf 1} {\mathbf 1^{\mathsf T}V^{-1}\mathbf 1}. ]
When both location and scale are unknown, unbiased location estimation also requires
[ \mathbf a^{\mathsf T}\mathbf m=0. ]
The resulting constrained quadratic problem determines weights that depend on the standardized parent distribution. E. H. Lloyd established this matrix formulation for least-squares estimation from order statistics, thereby connecting exact finite-sample calculations with the later general theory of L-estimation.
For a normal location model, the sample mean is already the minimum-variance unbiased estimator under the usual assumptions. Other distributions produce different optimal coefficient patterns because their order statistics have different covariance structures. Such parametric optimality is distribution-specific and is separate from the contamination-based criteria used in robust estimation.
Robustness and efficiency
The robustness of an L-estimator depends principally on the weight assigned near the endpoints of the empirical distribution. The sample mean retains nonzero sensitivity to observations of arbitrarily large magnitude, giving it an asymptotic replacement breakdown point of zero. A symmetrically (\alpha)-trimmed mean has an asymptotic breakdown point equal to (\alpha), subject to the precise finite-sample convention used for trimming.
The median concentrates its weight at the central rank and has an asymptotic breakdown point of one half. Its variance under a regular distribution with median (m) is
[ \operatorname{Var}(M_n) \sim \frac{1}{4n f(m)^2}. ]
The relative efficiency of different coefficient schemes therefore depends on the density around the quantiles receiving substantial weight. Central concentration protects against tail contamination but discards information carried by nonextreme observations, while broadly distributed weights use more of the sample and inherit greater sensitivity to tail behavior.
L-estimators occupy a specific position among the principal classes of robust procedures. An M-estimator is defined through minimization or an estimating equation, and its data-dependent solution is generally not a fixed linear combination of order statistics. An R-estimator is constructed from ranks or rank-test criteria. Coincidences between these classes occur for particular distributions and loss functions, but their defining mathematical structures remain distinct.