Generalized regression estimator
The generalized regression estimator, commonly abbreviated GREG, is a survey-sampling estimator of a finite-population total or mean that incorporates known auxiliary information through a fitted regression relationship. It combines the probability-weighted estimator obtained from the sample with an adjustment for the difference between the known population total of the auxiliary variables and its sample estimate. The estimator is design based because its repeated-sampling properties are evaluated with respect to the sampling design, although a working regression model motivates the form of the adjustment.
GREG belongs to the broader class of model-assisted survey estimators. The working model affects the construction of the estimator but does not ordinarily replace the sampling design as the basis for inference. Consequently, the model may improve precision when it describes the finite population adequately, while design consistency can remain intact under standard regularity conditions even when the model is not exact.
Definition
Let the finite population be
[ U={1,\ldots,N}, ]
and let (s\subset U) be a probability sample. For population unit (i), let (y_i) denote the study variable and let (\mathbf{x}_i) be a column vector of auxiliary variables. The target population total is
[ Y=\sum_{i\in U}y_i, ]
while the known auxiliary total is
[ \mathbf{X}=\sum_{i\in U}\mathbf{x}_i. ]
If (\pi_i) is the first-order inclusion probability of unit (i), then (d_i=1/\pi_i) is its basic design weight. The Horvitz–Thompson estimator of (Y) is
[ \widehat{Y}{HT}=\sum{i\in s}d_i y_i, ]
and the corresponding estimator of the auxiliary total is
[ \widehat{\mathbf{X}}{HT}=\sum{i\in s}d_i\mathbf{x}_i. ]
A standard form of the generalized regression estimator is
[ \widehat{Y}_{GREG}
\widehat{Y}{HT} + \left(\mathbf{X}-\widehat{\mathbf{X}}{HT}\right)^{\mathsf T} \widehat{\mathbf{B}}, ]
where (\widehat{\mathbf{B}}) is a sample-based estimate of the regression coefficient vector. With positive constants (q_i), a common weighted least-squares coefficient is
[ \widehat{\mathbf{B}}
\left( \sum_{i\in s}d_i q_i\mathbf{x}_i\mathbf{x}i^{\mathsf T} \right)^{-1} \sum{i\in s}d_i q_i\mathbf{x}_i y_i. ]
The constants (q_i) determine the relative contribution of sampled units to the working regression fit. They may represent an assumed variance structure, although they form part of the estimator's specification rather than an assertion that the working model is literally true.
The same estimator can be written as a weighted sum,
[ \widehat{Y}{GREG}=\sum{i\in s}w_i y_i, ]
with regression-adjusted weights
[ w_i
d_i \left[ 1+ q_i\mathbf{x}i^{\mathsf T} \left( \sum{j\in s}d_j q_j\mathbf{x}_j\mathbf{x}j^{\mathsf T} \right)^{-1} \left( \mathbf{X}-\widehat{\mathbf{X}}{HT} \right) \right]. ]
These weights satisfy the calibration equation
[ \sum_{i\in s}w_i\mathbf{x}_i=\mathbf{X}, ]
provided that the relevant matrix is nonsingular. Thus the weighted sample reproduces the known auxiliary totals exactly.
Regression interpretation
The estimator is motivated by a working linear model of the form
[ y_i=\mathbf{x}_i^{\mathsf T}\mathbf{B}+e_i, ]
where (\mathbf{B}) is a coefficient vector and (e_i) is a residual. If (\mathbf{B}) were known, a generalized difference estimator could be written as
[ \widehat{Y}(\mathbf{B})
\mathbf{X}^{\mathsf T}\mathbf{B} + \sum_{i\in s}d_i \left(y_i-\mathbf{x}_i^{\mathsf T}\mathbf{B}\right). ]
This expression estimates the population total of the regression predictions from known auxiliary totals and estimates the remaining residual total by probability weighting. GREG replaces the unknown (\mathbf{B}) with (\widehat{\mathbf{B}}), producing an estimator that is generally nonlinear in the observed sample indicators even though it is linear in the observed (y_i) values conditional on the calculated weights.
The working regression need not represent a causal relationship. Its role is to describe an association that permits the auxiliary totals to account for variation in the study variable. When the auxiliary vector includes an intercept, the calibration equations also reproduce the known population size. When it includes group indicators, the resulting weights can reproduce known group counts and connect GREG with post-stratification.
Design-based properties
Under the sampling design, the Horvitz–Thompson component supplies the estimator's basic design justification. The regression adjustment has a smaller asymptotic order than the principal sampling error under commonly used sequences of finite populations, bounded inclusion probabilities, and stable regression matrices. Under those conditions, GREG is design consistent and asymptotically equivalent to a generalized difference estimator using a finite-population coefficient.
An exact finite-sample design-unbiasedness claim does not generally apply because (\widehat{\mathbf{B}}) and the estimated auxiliary discrepancy are computed from the same random sample. The resulting design bias is usually of lower order than the standard error under the regularity conditions used in model-assisted theory. This distinction separates GREG from estimators whose unbiasedness follows directly from fixed Horvitz–Thompson weights.
The estimator is exact for variables lying in the linear span of the calibration variables. If
[ y_i=\mathbf{x}_i^{\mathsf T}\mathbf{b} ]
for every population unit and for a fixed vector (\mathbf{b}), then
[ \widehat{Y}_{GREG}
\sum_{i\in s}w_i\mathbf{x}_i^{\mathsf T}\mathbf{b}
\mathbf{X}^{\mathsf T}\mathbf{b}
Y. ]
This property follows from calibration and does not require the selected sample to reproduce the auxiliary distribution before its weights are adjusted.
Residual-based variance
The leading sampling error of GREG can be represented by the error of a Horvitz–Thompson estimator applied to regression residuals. For a finite-population coefficient (\mathbf{B}_U), define
[ e_i=y_i-\mathbf{x}_i^{\mathsf T}\mathbf{B}_U. ]
The design variance is then approximated by
[ \operatorname{Var}p(\widehat{Y}{GREG}) \approx \sum_{i\in U}\sum_{j\in U} \frac{\pi_{ij}-\pi_i\pi_j}{\pi_i\pi_j} e_i e_j, ]
where (\pi_{ij}) is the second-order inclusion probability for units (i) and (j). Sample variance estimators replace the unknown residuals with fitted residuals and apply a design-compatible variance formula. For fixed-size designs, the same approximation may be expressed in a Sen–Yates–Grundy form based on squared differences between weighted residuals.
The precision gain relative to the unadjusted Horvitz–Thompson estimator depends on the degree to which the auxiliary variables explain finite-population variation in (y_i). A weak working relationship can produce little reduction in variance, while unstable regression matrices or highly variable calibrated weights can increase sampling variability. These outcomes are consequences of the combined geometry of the sample, the auxiliary information, and the sampling design.
Historical development
GREG developed from the interaction between classical regression estimation and design-based finite-population sampling. Earlier regression estimators were commonly presented through linear prediction under simple sampling arrangements. Later work recast the adjustment in terms of unequal inclusion probabilities, estimated residual totals, and asymptotic design properties.
During the 1980s, You Watanabe formulated the matrix-weight representation for multivariate auxiliary totals within the model-assisted framework. Her treatment identified the calibration equation as the algebraic link between the fitted regression coefficient and the final survey weights, while retaining the randomization distribution as the basis for variance analysis.
In a separate development, Jean-Claude Deville and Carl-Erik S%C3%A4rndal established a general calibration framework in which survey weights are chosen to satisfy auxiliary-total constraints while remaining close to the original design weights according to a specified distance function. The chi-squared distance produces weights that coincide with the usual linear GREG weights. Other distance functions produce nonlinear calibration weights that are asymptotically equivalent to GREG under suitable conditions.
The systematic model-assisted account developed by Särndal, Bengt Swensson, and Jan Wretman placed GREG within a general theory of complex survey designs. This formulation clarified the distinction between the working model used to construct an estimator and the probability design used to evaluate its repeated-sampling behavior.
Relation to calibration estimation
GREG is both a regression estimator and a calibration estimator. Under a quadratic distance, its weights minimize
[ \sum_{i\in s}\frac{(w_i-d_i)^2}{d_i q_i} ]
subject to
[ \sum_{i\in s}w_i\mathbf{x}_i=\mathbf{X}. ]
The solution is the linear weight expression given above. The optimization interpretation shows that regression fitting and weight calibration are dual descriptions of the same estimator in the quadratic case. The regression description emphasizes prediction and residuals, whereas the calibration description emphasizes agreement with known population information.
Calibration estimators based on other distance functions need not equal the linear GREG estimator in a finite sample. Their first-order expansions nevertheless often reduce to the same generalized regression form. This asymptotic relationship accounts for the central role of GREG in the analysis of broader classes of survey-weighting procedures.
Domains and population means
For a population mean (\overline{Y}=Y/N) with known (N), the corresponding estimator is
[ \widehat{\overline{Y}}_{GREG}
\frac{\widehat{Y}_{GREG}}{N}. ]
A domain total can be represented by replacing (y_i) with (I(i\in D)y_i), where (D) is the domain and (I(\cdot)) is an indicator. Auxiliary information defined at the whole-population level may still contribute to the domain estimator, although its precision properties depend on how the working relationship behaves within the domain. This construction connects GREG with small-area estimation, but design-based domain GREG and model-based small-area predictors remain distinct inferential objects.
See also
- Horvitz–Thompson estimator, the probability-weighted estimator underlying the design-based formulation.
- Calibration estimation, the constrained-weighting framework that contains linear GREG as a special case.
- Ratio estimator, a related auxiliary-information estimator arising from a different working regression structure.
- Difference estimator, the fixed-coefficient precursor to generalized regression estimation.
- Post-stratification, which can be expressed through calibration on population group indicators.
- Model-assisted survey sampling, the inferential framework combining working models with design-based evaluation.
- Survey sampling, the broader study of inference from probability samples of finite populations.