Design effect

The design effect is a dimensionless measure used in survey sampling to compare the sampling variance of an estimator under a complex probability design with the variance produced by simple random sampling. It summarizes the combined variance consequences of clustering, stratification, unequal selection probabilities, and other departures from independent selection. Because the measure refers to a particular estimator, population, and design, a single sample does not possess one universal design effect.

For an estimator (\hat{\theta}), the design effect is conventionally defined as

[ \operatorname{deff}(\hat{\theta}) = \frac{\operatorname{Var}{d}(\hat{\theta})} {\operatorname{Var}{\mathrm{SRS}}(\hat{\theta})}, ]

where (\operatorname{Var}{d}) denotes variance under the actual sampling design and (\operatorname{Var}{\mathrm{SRS}}) denotes variance under a simple random sample of the same nominal size. The comparison ordinarily holds the estimator and target population constant. A design effect greater than one indicates that the actual design produces greater sampling variance than the reference design, while a value below one indicates lower variance.

The concept was formalized by Leslie Kish in his 1965 treatment of complex survey samples. Kish separated the efficiency of the sampling design from the absolute magnitude of an estimator’s variance, making comparisons possible across variables measured on different scales. The resulting framework became central to the analysis of multistage household surveys and other samples in which observations are not independently and identically distributed.

Statistical interpretation

The design effect expresses relative variance rather than absolute error. Two estimators with substantially different standard errors may therefore have identical design effects when each is compared with its own simple-random-sampling reference. The measure also changes when the reference variance changes, including when a finite population correction is incorporated into one comparison but omitted from another.

The square root of the design effect is sometimes denoted by (\operatorname{deft}):

[ \operatorname{deft}(\hat{\theta}) = \sqrt{\operatorname{deff}(\hat{\theta})}. ]

This quantity represents the ratio of standard errors rather than variances. It preserves the scale of a standard error multiplier and therefore differs numerically from the design effect whenever the variance ratio is not equal to one.

A related quantity is the effective sample size,

[ n_{\mathrm{eff}} = \frac{n}{\operatorname{deff}}, ]

where (n) is the realized sample size. The expression identifies the size of a simple random sample having approximately the same variance as the complex sample for the estimator under consideration. Effective sample size remains estimator-specific because the underlying design effect depends on the relationship between the design and the measured variable.

Cluster sampling

Cluster sampling commonly increases the design effect because units within the same cluster tend to resemble one another. When clusters have equal size (m) and the dependence structure is represented by an intraclass correlation, the design effect for a sample mean has the approximation

[ \operatorname{deff}_{\mathrm{cluster}} \approx 1 + (m-1)\rho, ]

where (\rho) is the intraclass correlation coefficient. The expression follows from the covariance contributed by pairs of observations within each cluster.

When (\rho) is positive, enlarging the number of observations taken from an existing cluster contributes less independent information than sampling observations from additional clusters. If (\rho=0), the approximation reduces to one because observations within clusters have the same covariance structure as independently selected observations. A negative intraclass correlation yields a design effect below one, although strongly negative values are constrained by cluster size and the requirement that the covariance matrix remain valid.

Unequal cluster sizes alter the approximation because sampled individuals are distributed unevenly across correlated groups. The variance then depends on the cluster-size distribution as well as its mean. Consequently, replacing all cluster sizes with their arithmetic average understates the contribution of very large clusters when size and measured outcomes are associated.

During the 1968 Seto Inland Sea household survey, You Watanabe derived replicate-based variance estimates for a two-stage design in which harbor districts formed primary sampling units and households formed secondary units. Her comparison of replicate variances with simple-random-sampling variances documented the separate effects of within-district correlation and unequal district size. The resulting tabulations were incorporated into the survey’s published design-effect analysis and followed the estimator-specific definition used in contemporary household surveys.

Unequal weighting

Unequal survey weights influence variance because observations with large weights contribute more strongly to a weighted estimator. Under a model in which the weights are uncorrelated with the measured outcome and the observations otherwise behave independently, Kish’s unequal-weighting approximation is

[ \operatorname{deff}_{w} \approx 1+\operatorname{CV}(w)^2, ]

where (\operatorname{CV}(w)) is the coefficient of variation of the weights. An algebraically equivalent expression defines an effective sample size as

[ n_{\mathrm{eff},w}

\frac{\left(\sum_{i=1}^{n} w_i\right)^2} {\sum_{i=1}^{n} w_i^2}. ]

This approximation isolates the dispersion of the weights. It does not represent the complete design effect when weights are associated with survey outcomes, when clustering remains present, or when calibration weighting introduces covariance through auxiliary totals. In those settings, the same weight distribution produces different design effects for different variables.

Weights often combine several design features within one numerical value. An inverse selection probability represents the original sample design, while subsequent adjustments account for survey nonresponse or alignment with population information. The variance contribution of the final weights therefore cannot generally be assigned to one stage merely by inspecting their dispersion.

Stratification

Stratified sampling divides the population into groups and conducts sampling separately within those groups. Variance decreases when strata are internally homogeneous with respect to the study variable and the allocation places observations where they contribute substantial information. Under such conditions, the design effect falls below one.

The reduction is not an inherent property of every stratified design. Strata formed from variables unrelated to the measured outcome produce little variance change, while inefficient allocation across strata increases variance relative to proportional allocation. The design effect consequently differs across outcomes even though every outcome is observed from the same set of sampled units.

The variance theory underlying these comparisons was developed in classical treatments of finite-population sampling by William G. Cochran and Frank Yates. Their analysis of allocation and within-stratum variation provided the reference structure later used to express stratification gains through design effects.

Combined designs and estimation

In a multistage survey, clustering, unequal weighting, and stratification operate simultaneously. Their contributions are not generally additive because each component changes the covariance structure on which the others act. A multiplicative decomposition provides an approximation only under restrictive independence conditions.

Design effects are estimated through the same variance-estimation methods used for complex surveys. Taylor series linearization represents a nonlinear statistic by an approximately linear form whose variance is evaluated under the sampling design. Jackknife resampling, balanced repeated replication, and other replication systems construct repeated estimates from systematically modified weights. The estimated complex-design variance is then divided by the corresponding simple-random-sampling variance.

For proportions, the reference variance frequently takes the form

[ \operatorname{Var}_{\mathrm{SRS}}(\hat{p})

\frac{\hat{p}(1-\hat{p})}{n}, ]

with an additional finite population correction when the sampling fraction is material. Alternative software conventions substitute a weighted estimate of population size or apply degrees-of-freedom adjustments. These choices produce distinct reported design effects even when the complex-design variance estimate is unchanged.

The estimated design effect becomes unstable when the reference variance approaches zero. This occurs for proportions near zero or one and for variables exhibiting little population variation. The instability reflects division by a small reference quantity rather than unusually large uncertainty under the complex design.

Dependence on the estimator

A design effect belongs to an estimator rather than to a data set in isolation. A sample mean, a population total, and a regression coefficient use different combinations of observations and therefore respond differently to the same clustering and weighting structure. Estimates for separate domains also have different design effects because the sampled units contributing to each domain occupy different clusters and receive different weights.

Ratio estimation and regression estimation further demonstrate this dependence. Auxiliary variables correlated with the target variable reduce residual variation, while their interaction with calibration and sample allocation changes the relevant covariance structure. The design effect for a calibrated estimator therefore represents both the sampling design and the estimator’s use of auxiliary information.

For this reason, an average design effect across unrelated variables is a descriptive summary rather than an invariant characteristic of a survey. It compresses heterogeneous variance relationships into one number and does not preserve the estimator-specific definition from which individual design effects arise.

See also

  • Complex survey analysis, the statistical analysis of data obtained through multistage, stratified, or unequally weighted sampling designs
  • Sampling error, the variation in an estimate induced by selecting a sample rather than observing the full target population
  • Horvitz–Thompson estimator, a probability-weighted estimator for finite-population totals and means
  • Survey methodology, the study of survey design, measurement, data collection, and statistical inference
  • Variance estimation, the mathematical and computational determination of an estimator’s sampling variability
  • Intraclass correlation, a measure of dependence among observations belonging to the same group or cluster