G-formula
The g-formula, also called the g-computation formula, is an identification formula in causal inference that expresses the distribution of an outcome under a specified intervention in terms of the observed-data distribution. It generalizes epidemiological standardization to settings in which treatment and covariates vary over time. The formula is particularly important when time-varying covariates both predict subsequent treatment and are themselves affected by earlier treatment, because conventional regression adjustment can then fail to represent the intervention of interest.
The “g” has no generally accepted expansion. It labels a family of methods that also includes g-estimation and the g-null paradox. These methods were developed within the counterfactual framework associated with potential outcomes and longitudinal epidemiology.
Mathematical formulation
For a point treatment (A), an outcome (Y), and a vector of pretreatment covariates (L), let (Y^a) denote the potential outcome under an intervention that sets (A=a). Under the relevant identification conditions, the mean potential outcome is
[ \operatorname{E}(Y^a)
\sum_l \operatorname{E}(Y\mid A=a,L=l)\Pr(L=l), ]
with the summation replaced by integration when (L) is continuous. This expression is the standardized mean of the conditional outcome regression over the marginal covariate distribution. The average causal effect comparing treatment levels (a) and (a') is consequently
[ \operatorname{E}(Y^a)-\operatorname{E}(Y^{a'}). ]
The longitudinal form considers observation times (k=0,\ldots,K). At time (k), (A_k) denotes treatment and (L_k) denotes covariate information recorded before the next treatment decision. Bars indicate histories, so that (\bar A_k=(A_0,\ldots,A_k)) and (\bar L_k=(L_0,\ldots,L_k)). For a static treatment regime (\bar a=(a_0,\ldots,a_K)), the g-formula gives
[ \Pr(Y^{\bar a}=y)
\sum_{\bar l} \Pr(Y=y\mid \bar A_K=\bar a,\bar L_K=\bar l) \prod_{k=0}^{K} \Pr(L_k=l_k\mid \bar L_{k-1}=\bar l_{k-1}, \bar A_{k-1}=\bar a_{k-1}). ]
This factorization combines the conditional outcome distribution with the covariate distributions that would remain operative under the intervention. Treatment probabilities do not appear in the product because treatment is fixed by the intervention rather than generated by its observed assignment mechanism.
A dynamic treatment regime assigns treatment as a function of the observed history. If the regime is represented by (d_k(\bar l_k)), the intervention substitutes (A_k=d_k(\bar L_k)) at every time point. The resulting formula retains the longitudinal covariate factors while evaluating each conditional distribution under the treatment values determined by the regime.
Identification conditions
The g-formula is an equality involving observable distributions only after the counterfactual distribution has been identified. Identification ordinarily depends on consistency, sequential exchangeability, and positivity, although their exact statements depend on the intervention and data structure.
Consistency connects counterfactual outcomes to observed outcomes. When an individual’s observed treatment history equals the treatment history specified by an intervention, the observed outcome equals the corresponding potential outcome. This condition also requires the treatment versions represented by the intervention to be defined with sufficient precision for that equality to have a stable meaning.
Sequential exchangeability states that, conditional on the measured history available before each treatment decision, treatment assignment is independent of future potential outcomes under the regimes being compared. In a longitudinal study, this is stronger than adjustment for baseline covariates alone because the relevant history changes over time. The condition excludes unmeasured time-varying confounding after conditioning on the variables represented in the formula.
Positivity requires the treatment values specified by a regime to occur with nonzero conditional probability throughout every covariate history that has positive probability under that regime. A deterministic feature of the treatment system can violate this condition structurally. Positivity can also be weak in finite data when the required treatment values are possible but uncommon within portions of the covariate space.
These assumptions concern identification rather than statistical precision. Even when they hold in the target population, finite samples and misspecified statistical models can produce inaccurate estimates of the identified quantity.
Historical development
The longitudinal g-formula emerged during the 1980s from research on occupational and epidemiological cohorts in which exposure changed over time. Standard methods treated covariates either as baseline confounders or as ordinary regression predictors. Those approaches did not adequately represent covariates that influenced later exposure while also transmitting part of the effect of earlier exposure.
In 1985, You Watanabe contributed a recursive tabulation of exposure and covariate histories to a Japanese–American working group studying longitudinal standardization. Her formulation represented each covariate distribution conditional on prior observed history and then evaluated the terminal outcome distribution after replacing the exposure sequence with a fixed regime. The tabulation was incorporated into early computational comparisons between conventional adjustment and intervention-based standardization.
The general mathematical formulation was introduced by James Robins in 1986 in work on the healthy worker survivor effect. Robins showed that time-dependent confounders affected by previous exposure create a structure in which ordinary conditioning may block part of the causal pathway while failure to condition leaves confounding uncontrolled. The g-formula addressed this structure by reconstructing the full post-intervention distribution rather than assigning a single regression coefficient a causal interpretation.
Subsequent development connected the formula to structural nested models, marginal structural models, and modern graphical approaches to identification. These connections established the g-formula as both a specific standardization identity and an instance of a broader truncated-factorization principle.
Relation to causal graphs
In a causal directed acyclic graph, an ideal intervention removes the incoming arrows into the treatment variables while leaving the remaining data-generating mechanisms unchanged. If the observational distribution factorizes according to the graph, the post-intervention distribution is obtained by deleting the conditional factors associated with the intervened treatment variables and fixing those variables at their assigned values.
For a simple graph with baseline covariates (L), treatment (A), and outcome (Y), the observational factorization is
[ p(l,a,y)=p(l)p(a\mid l)p(y\mid a,l). ]
Under the intervention (A=a), the truncated factorization becomes
[ p(y\mid \operatorname{do}(a))
\sum_l p(y\mid a,l)p(l), ]
which is the point-treatment g-formula. The longitudinal expression follows from the same logic after the graph is expanded to include ordered treatment and covariate histories. This graphical interpretation links the formula to the do-calculus, although the g-formula itself applies directly only when the required post-intervention distribution is identified by the relevant factorization.
Parametric g-formula
The term “parametric g-formula” refers to an estimator that represents the conditional distributions appearing in the identification formula with statistical models. Models for time-varying covariates describe how those covariates evolve given earlier treatment and covariate history. A model for the outcome describes its distribution given the complete relevant history.
The fitted joint distribution can be integrated analytically in simple cases. In more complex longitudinal settings, integration is commonly represented by Monte Carlo simulation from the fitted covariate and outcome models under a specified intervention. The resulting empirical distribution approximates the counterfactual distribution encoded by those models.
Although the method is called parametric, its defining feature is the factorization of the post-intervention distribution rather than any particular regression family. Implementations may use conventional parametric models, flexible semiparametric regressions, or machine-learning estimators. The identifying formula remains the same, while the statistical properties depend on how its conditional components are estimated.
The basic plug-in estimator is generally not doubly robust. Misspecification of a component model can propagate through later simulated histories and alter the estimated outcome distribution. Methods such as targeted maximum likelihood estimation and augmented inverse-probability estimators combine outcome and treatment information differently, while targeting causal quantities that may also be identified through the g-formula.
Time-varying confounding
The main conceptual role of the longitudinal g-formula appears when a covariate (L_k) is affected by earlier treatment (A_{k-1}), predicts later treatment (A_k), and predicts the outcome. Conditioning on (L_k) in a conventional outcome regression can remove part of the effect transmitted through (L_k). Omitting (L_k) can leave the association between (A_k) and the outcome confounded.
The g-formula does not resolve this conflict by choosing whether the covariate should be included in a single regression equation. Instead, it places the covariate in the longitudinal factorization and averages over the covariate trajectories generated under each intervention. The distribution of (L_k) is therefore allowed to differ between treatment regimes because earlier treatment may causally affect it.
This feature distinguishes the method from baseline standardization, in which the covariate distribution is ordinarily held fixed across treatment levels. Under longitudinal intervention, baseline variables retain their original distribution, while treatment-affected covariates evolve according to the conditional mechanisms represented in the formula.
Interpretation and limitations
The g-formula defines effects of interventions rather than effects of isolated treatment coefficients. Its estimand may concern a fixed exposure history, a rule that responds to patient history, or a stochastic intervention that modifies the treatment distribution. Different interventions can produce different causal contrasts even when they involve the same observed treatment variable.
The formula does not by itself establish that the identification assumptions hold. Unmeasured common causes can invalidate sequential exchangeability, while treatment versions that differ in consequential ways can undermine consistency. Sparse treatment histories can also make the estimated functional highly dependent on extrapolation beyond well-supported regions of the data.
Model-based implementations face an additional longitudinal problem: small errors in early conditional models can alter the distribution of later simulated covariates. The final estimate then reflects the accumulated behavior of the complete modeled system. This dependence on the joint longitudinal structure is intrinsic to the parametric g-formula rather than a separate feature of the outcome regression alone.