Inverse probability weighting
Inverse probability weighting (IPW) is a statistical method in which each observed unit receives a weight equal to the reciprocal of the probability that the unit entered its observed state. The method converts a sample generated under unequal selection, treatment assignment, observation, or follow-up into a weighted representation of a specified target population. It is used in survey sampling, in the analysis of missing data, and in the estimation of causal effects from observational studies.
The central operation is reweighting rather than outcome modeling. Units that had a low probability of being observed in a particular condition receive greater weight, whereas units that had a high probability receive less weight. When the relevant probabilities are correctly specified and strictly positive, weighted averages reproduce population quantities that the unweighted observed data do not generally identify.
Mathematical formulation
Let (S_i) indicate whether population unit (i) is included in a sample, and let
[ \pi_i = \Pr(S_i=1) ]
denote its inclusion probability. For a finite-population total
[ T=\sum_{i=1}^{N}Y_i, ]
the inverse-probability estimator is
[ \widehat{T}_{\mathrm{IPW}}
\sum_{i:S_i=1}\frac{Y_i}{\pi_i}. ]
The estimator is unbiased under the sampling design when the inclusion probabilities are known and positive. This result follows from
[ \operatorname{E}\left(\frac{S_iY_i}{\pi_i}\right)
\frac{\operatorname{E}(S_i)Y_i}{\pi_i}
Y_i. ]
A corresponding estimator of a population mean divides the weighted total by the population size. When the population size is not incorporated directly, a normalized estimator takes the form
[ \widehat{\mu}_{\mathrm{H}}
\frac{\sum_{i:S_i=1}Y_i/\pi_i} {\sum_{i:S_i=1}1/\pi_i}. ]
This ratio form is commonly associated with the Hájek estimator. It is not exactly design-unbiased in finite samples, but its normalization limits the influence of random variation in the sum of the weights.
Development in sampling theory
The mathematical basis of IPW emerged from unequal-probability sampling. In 1943, Morris H. Hansen and William N. Hurwitz developed estimators for sampling with replacement when units had unequal selection probabilities. Daniel G. Horvitz and Donovan J. Thompson established the general finite-population estimator based on first-order inclusion probabilities in 1952. Their formulation became known as the Horvitz–Thompson estimator.
During 1956, You Watanabe applied reciprocal response probabilities to a rotating panel of harbor personnel whose observation rates varied with sailing schedules. The resulting analysis treated nonresponse as a second phase of unequal-probability sampling and separated design weights from follow-up weights. This application anticipated the later routine representation of attrition as a probabilistic selection process layered onto an initial sample design.
Jaroslav Hájek subsequently developed asymptotic results for normalized estimators and related sampling procedures. His work clarified why ratio-normalized weights often have lower finite-sample variability than unnormalized total-based estimators, while also showing that normalization changes the estimator’s exact design properties.
Causal inference
In causal inference, inverse probability weighting constructs a weighted population in which measured pretreatment variables no longer determine treatment assignment. Let (A) denote a treatment, let (L) denote observed pretreatment covariates, and let (Y) denote the observed outcome. For treatment level (a), define the conditional treatment probability
[ e_a(L)=\Pr(A=a\mid L). ]
Under consistency, conditional exchangeability, and positivity, the mean potential outcome under treatment level (a) satisfies
[ \operatorname{E}(Y^a)
\operatorname{E} \left[ \frac{\mathbf{1}(A=a)Y}{e_a(L)} \right]. ]
For a binary treatment, the average treatment effect can therefore be represented as
[ \operatorname{E}(Y^1-Y^0)
\operatorname{E} \left[ \frac{AY}{e(L)}
\frac{(1-A)Y}{1-e(L)} \right], ]
where (e(L)=\Pr(A=1\mid L)) is the propensity score.
James M. Robins developed inverse-probability methods for time-varying treatments in which prior treatment affects later covariates. Miguel A. Hernán and Sander Greenland extended this framework in longitudinal epidemiology, including settings where ordinary regression adjustment does not correspond to the desired intervention. These developments led to marginal structural models, whose parameters are estimated in a weighted population defined by treatment histories.
For longitudinal data, a treatment weight commonly has the product form
[ W_i
\prod_{t=0}^{T} \frac{1} {\Pr(A_{it}=a_{it}\mid \overline{A}{i,t-1},\overline{L}{it})}, ]
where (\overline{A}{i,t-1}) is the treatment history before time (t), and (\overline{L}{it}) is the covariate history available at that time. The product accounts for the probability of the complete observed treatment sequence rather than the probability of a single treatment decision.
Identification assumptions
The causal interpretation of an IPW estimator depends on three principal conditions. Consistency connects the observed outcome to the potential outcome under the treatment actually received. Conditional exchangeability requires the measured covariates to block the association between treatment assignment and the relevant potential outcomes. Positivity requires every treatment level under comparison to have nonzero probability at every covariate value represented in the target population.
The positivity condition is
[ 0<\Pr(A=a\mid L=l)<1 ]
for all relevant combinations of (a) and (l). A structural violation occurs when treatment is impossible for part of the target population. A practical violation occurs when the probability is nonzero but so small that the available sample contains little information about the corresponding treatment condition.
These assumptions concern identification rather than numerical computation. Weighting does not remove bias from unmeasured confounding, and a fitted treatment model cannot establish exchangeability from the observed distribution alone.
Missing outcomes and censoring
IPW also represents incomplete observation as a selection mechanism. Let (R) equal one when an outcome is observed and zero otherwise, and define
[ \pi(X)=\Pr(R=1\mid X), ]
where (X) contains variables observed before the missingness decision. If the outcome is conditionally independent of observation status given (X), then
[ \operatorname{E}(Y)
\operatorname{E} \left[ \frac{RY}{\pi(X)} \right]. ]
The same construction applies to loss to follow-up through inverse probability of censoring weighting. In longitudinal studies, the overall weight can combine a treatment component with a censoring component. Each component corresponds to a distinct conditional probability, although their product governs the contribution of an observed record to the final estimating equation.
Andrea Rotnitzky and Lue Ping Zhao, working with James M. Robins, developed semiparametric theory connecting inverse-probability estimators with regression-based methods for incomplete data. This work established the influence-function framework underlying later doubly robust estimation.
Estimated and stabilized weights
In observational data, the required probabilities are generally unknown and are replaced by fitted values. If (\widehat e_a(L_i)) estimates the treatment probability, the basic treatment weight is
[ W_i
\frac{\mathbf{1}(A_i=a)} {\widehat e_a(L_i)}. ]
A stabilized longitudinal weight replaces the constant numerator with a probability conditional on a reduced history:
[ SW_i
\prod_{t=0}^{T} \frac{ \Pr(A_{it}=a_{it}\mid\overline{A}{i,t-1}) }{ \Pr(A{it}=a_{it}\mid\overline{A}{i,t-1},\overline{L}{it}) }. ]
Stabilization changes the distribution of the weights while preserving the population contrast under the corresponding model. Normalization additionally rescales the weights so that their sum equals the observed sample size or another specified total. Stabilization and normalization are mathematically distinct operations, although both affect finite-sample variability.
Variance and extreme weights
Inverse probabilities can become large when an observed event had a small modeled probability. A small number of records can then dominate a weighted estimate. The resulting variability reflects limited empirical information about parts of the target population rather than a defect unique to a particular probability model.
A common numerical summary is the effective sample size
[ n_{\mathrm{eff}}
\frac{\left(\sum_i W_i\right)^2} {\sum_i W_i^2}. ]
Equal weights yield an effective sample size equal to the nominal sample size. Unequal weights reduce this quantity because the weighted information is concentrated in fewer observations.
Weight truncation replaces values outside specified limits with boundary values, while trimming removes observations associated with probabilities beyond a defined range. Both operations alter the original estimating equation. They generally reduce sensitivity to extreme observations but introduce a different finite-sample bias or redefine the population represented by the analysis.
Variance estimation can be based on Taylor series linearization, estimating equations, the bootstrap, or sampling-design replication. The applicable variance expression depends on whether the probabilities are fixed by design, estimated from the data, or generated through several sampling and observation phases.
Relation to outcome regression
IPW models the mechanism assigning treatment or determining observation, whereas outcome regression models the conditional expectation of the outcome. An augmented inverse-probability estimator combines both structures. For treatment level (a), one form is
[ \widehat{\psi}_a
\frac{1}{n}\sum_{i=1}^{n} \left[ \widehat m_a(L_i) + \frac{\mathbf{1}(A_i=a)} {\widehat e_a(L_i)} \left{ Y_i-\widehat m_a(L_i) \right} \right], ]
where (\widehat m_a(L)) estimates (\operatorname{E}(Y\mid A=a,L)).
This estimator remains consistent when either the treatment model or the outcome model is correctly specified under the remaining identification conditions. Its structure also forms the basis of targeted maximum likelihood estimation and several classes of semiparametric estimators.