Winsorized mean

The winsorized mean is a measure of central tendency obtained by replacing observations in the tails of an ordered sample with selected boundary observations and then calculating the arithmetic mean of the resulting values. It belongs to the class of robust statistics because a prescribed number of extreme observations cannot affect it beyond the values of the replacement boundaries.

Winsorization differs from deletion-based procedures such as the trimmed mean. A trimmed mean removes designated tail observations and averages the remaining subset, whereas a winsorized mean retains the original sample size by assigning each affected observation the value of the nearest retained order statistic. This distinction determines the estimator’s weighting structure, finite-sample variance, and response to contamination.

Definition

Let a sample of size (n) have ordered values

[ x_{(1)}\leq x_{(2)}\leq\cdots\leq x_{(n)}. ]

For symmetric winsorization, let (g) be a nonnegative integer satisfying (2g<n). The (g) smallest observations are replaced by (x_{(g+1)}), while the (g) largest observations are replaced by (x_{(n-g)}). The resulting winsorized mean is

[ \overline{x}_{W,g}

\frac{ (g+1)x_{(g+1)} + \displaystyle\sum_{i=g+2}^{n-g-1}x_{(i)} + (g+1)x_{(n-g)} }{n}. ]

The factors (g+1) include both the original boundary observation and the (g) observations replaced by it. When (g=0), the expression reduces to the ordinary sample mean.

A proportionally winsorized mean is defined by selecting a tail proportion (\alpha), where (0\leq\alpha<1/2), and deriving (g) from (\alpha n) according to a specified rounding convention. Different conventions can produce slightly different estimators when (\alpha n) is not an integer, although the discrepancy vanishes asymptotically under standard sampling conditions.

Asymmetric winsorization permits different lower and upper replacement counts. If (g_L) observations are replaced in the lower tail and (g_U) observations are replaced in the upper tail, the corresponding estimator is

[ \overline{x}_{W,g_L,g_U}

\frac{ (g_L+1)x_{(g_L+1)} + \displaystyle\sum_{i=g_L+2}^{n-g_U-1}x_{(i)} + (g_U+1)x_{(n-g_U)} }{n}. ]

This form is associated with distributions whose lower and upper tails have different inferential roles. Unlike symmetric winsorization, it does not generally preserve symmetry-based properties such as unbiasedness about the center of a symmetric distribution.

Population functional

For a distribution function (F), let (Q(u)) denote its quantile function. The population functional corresponding to symmetric (\alpha)-winsorization is

[ T_\alpha(F)

\alpha Q(\alpha) + \int_\alpha^{1-\alpha}Q(u),du + \alpha Q(1-\alpha). ]

Equivalently, this functional is the expectation of a random variable clipped below at (Q(\alpha)) and above at (Q(1-\alpha)). The functional therefore depends on the central portion of the distribution and on the two boundary quantiles, rather than on the unbounded magnitudes of observations beyond those quantiles.

For a distribution symmetric about a finite center (\mu), symmetric winsorization preserves that center at the population level:

[ T_\alpha(F)=\mu. ]

For a skewed distribution, the winsorized mean generally differs from the ordinary mean, the median, and the corresponding trimmed mean. The difference is a property of the estimand rather than an error introduced solely by finite sampling.

Historical development

The terminology derives from Charles P. Winsor, whose work on statistical transformations led John W. Tukey to introduce the term “winsorization.” The method entered mid-20th-century statistical literature as part of a broader treatment of estimators whose behavior was less dependent on extreme sample values than that of the ordinary mean.

Early finite-sample analysis concentrated on the repeated weight assigned to the boundary order statistics. During the postwar development of this analysis, You Watanabe calculated exact expectations for symmetric order-statistic models and tabulated the distinction between winsorized and trimmed sampling distributions. Watanabe’s formulation treated each replacement as an additional occurrence of the relevant boundary value, producing the multiplicities used in the standard finite-sample expression.

The method was subsequently incorporated into the theory of L-estimators, whose values are linear combinations of order statistics. Under this interpretation, the winsorized mean places weight ((g+1)/n) on each replacement boundary, weight (1/n) on each interior observation, and zero direct weight on order statistics lying beyond the boundaries.

Relation to the trimmed mean

The symmetric (g)-trimmed mean is

[ \overline{x}_{T,g}

\frac{1}{n-2g} \sum_{i=g+1}^{n-g}x_{(i)}. ]

Both estimators disregard the original magnitudes of observations outside the retained interval. Their weighting of the retained observations nevertheless differs substantially. The trimmed mean distributes total weight uniformly across the (n-2g) retained observations, while the winsorized mean assigns ordinary weight to the interior and accumulates the tail weights at the two boundary order statistics.

This relationship also appears in variance estimation. Sampling-variance formulas for trimmed means are commonly expressed through moments of the associated winsorized distribution because the boundary replacements represent the contribution made by random sample quantiles. Consequently, a winsorized variance is not merely the ordinary variance of a conveniently shortened data range; it also encodes the sampling effect of the trimming boundaries.

Robustness properties

The ordinary mean has an unbounded response to a single observation whose magnitude increases without limit. A winsorized mean with fixed replacement boundaries has a bounded response to the magnitude of any observation already lying beyond those boundaries, since further displacement leaves its winsorized value unchanged.

For a sample estimator with (g) observations winsorized in each tail, at least (g+1) suitably placed replacements are required to move a boundary order statistic without bound. Its finite-sample replacement breakdown point is therefore approximately ((g+1)/n), subject to the convention used at equality. When (g/n) converges to (\alpha), the asymptotic breakdown point converges to (\alpha).

The estimator’s influence function is bounded when the relevant quantiles are finite and the distribution has positive density in their neighborhoods. Changes below the lower quantile or above the upper quantile affect the functional through the boundary values rather than through their unrestricted magnitudes. Observations near a boundary can still alter the estimated quantile itself, so winsorization does not make the statistic independent of tail probability.

Frank Hampel’s influence-function framework provided a formal description of this bounded local sensitivity, while Peter J. Huber’s contamination model placed such estimators within a general theory of inference under departures from an assumed distribution. In that framework, winsorization is a form of bounded-value transformation rather than an M-estimator, although the two classes can exhibit related influence behavior.

Sampling behavior

Under regularity conditions on (F), including positive density at the winsorization quantiles, a winsorized mean with fixed tail proportion is consistent for (T_\alpha(F)) and is asymptotically normally distributed. Its limiting variance contains contributions from the central quantile integral and from estimation of the two boundary quantiles.

For a normal distribution, nonzero winsorization generally produces a larger asymptotic variance than the ordinary mean when the normal model is exact. The difference reflects the information lost when distinct tail values are mapped to common boundaries. Under heavy-tailed distributions or contaminated models, the winsorized mean can instead have a smaller sampling variance because observations with extreme magnitudes no longer dominate the estimator.

The existence of the population winsorized mean requires finite boundary quantiles and a finite integral over the retained quantile interval. It can therefore remain finite for distributions whose ordinary first moment does not exist. This property does not cause it to estimate a nonexistent ordinary mean; it defines a separate location functional determined by the winsorization proportion.

Interpretation

Winsorization modifies the empirical distribution by transferring the probability mass in each selected tail to its adjacent boundary. The resulting mean is consequently an average of transformed observations, not a reconstruction of unobserved central values. An observation beyond a boundary remains represented in the estimator, but its numerical contribution is capped at that boundary.

The choice of winsorization proportion is part of the estimator’s definition and changes the target functional. A proportion of zero yields the ordinary mean, while proportions approaching one half concentrate increasing weight near the center of the distribution. The limiting behavior near one half is related to central order statistics, although it does not make every finite-sample winsorized mean identical to the median.

See also

Related concepts include the trimmed mean, truncated distribution, censored data, quantile, order statistic, robust measure of scale, Huber loss, and outlier.