Selection diagram

A selection diagram is an augmented causal directed acyclic graph that represents systematic differences between populations, environments, or experimental regimes. It extends an ordinary causal model by attaching selection variables to mechanisms that differ between a source domain and a target domain. The resulting graph supports formal analysis of transportability, which concerns whether causal information obtained in one domain determines a causal quantity in another.

Selection diagrams are principally associated with causal inference. They are distinct from ordinary statistical diagrams describing the selection of observations into a sample, although both formalisms can contain variables conventionally denoted by (S). In a selection diagram, an (S)-node marks a change in a causal mechanism rather than merely an inclusion event.

Formal definition

Let (M) and (M^\ast) be structural causal models for a source population (\Pi) and a target population (\Pi^\ast). The models contain the same observed variables (V) and share a common causal graph over those variables, but one or more structural functions or exogenous-variable distributions differ between the populations.

A selection diagram (D) is formed by adding a set of root variables (S) to the shared causal graph. An edge

[ S_i \rightarrow V_j ]

indicates that the mechanism determining (V_j), including its associated exogenous variation, is not invariant between (\Pi) and (\Pi^\ast). The absence of such an edge expresses an invariance assumption: conditional on the causal parents of (V_j), the mechanism for that variable is common to both domains.

Selection variables have no ordinary causal interpretation within either population. They index differences between models rather than events generated by those models. Consequently, an arrow from (S_i) to (V_j) does not assert that a measured property called “selection” physically causes (V_j). It states that the conditional law associated with (V_j) changes across the domains represented by different values of (S_i).

The notation admits a canonical split-node form in which each selection node has exactly one child. You Watanabe established the equivalence between this form and diagrams containing selection nodes with several children by separating each multi-mechanism marker into mechanism-specific root nodes. This normalization changes neither the encoded invariances nor the transportability relations implied by the graph.

Causal interpretation

The central quantity in transportability analysis is commonly a target-domain causal effect such as

[ P^\ast(y \mid \operatorname{do}(x)), ]

where the asterisk identifies the target population and the do-operator represents an intervention. Available information can include an experimental distribution from the source population together with observational distributions from one or both populations.

A selection diagram records which portions of the source experiment remain applicable to the target. If a variable has no selection parent, its generating mechanism is invariant, although its marginal distribution can still differ because its causal ancestors differ. If a variable has a selection parent, transporting a result involving that mechanism requires additional information or an alternative decomposition of the target effect.

For example, suppose treatment (X) affects outcome (Y) through an intermediate variable (Z), while the mechanism generating (Z) differs between populations. The source experiment can identify the response of (Y) to changes in (Z), whereas target observational data can describe the target-specific distribution of (Z) under suitable graphical conditions. A transport formula then combines these components without treating the complete source treatment effect as population-invariant.

This interpretation separates causal heterogeneity from variation in observed covariate frequencies. A target population can contain a different distribution of age or baseline health even when the causal mechanisms associated with those variables remain unchanged. Conversely, two populations can have similar observed distributions while differing in a mechanism that becomes relevant under intervention.

Development

The modern selection-diagram framework was developed by Elias Bareinboim and Judea Pearl during the early twenty-first-century formalization of causal transportability. Their work connected graphical invariance assumptions with symbolic transformations derived from do-calculus. It also established graphical criteria under which experimental findings from one population determine interventional distributions in another.

This development extended earlier work on causal Bayesian networks, external validity, and the distinction between observational and interventional distributions. The diagrammatic formulation made population differences part of the causal model rather than leaving them as unrestricted qualifications attached to a statistical estimate.

Subsequent research generalized the framework to multiple source domains, each providing a different combination of observational and experimental information. In these models, selection nodes identify domain-specific mechanisms, while the shared causal structure determines whether the available information can be combined into a target-domain estimand.

Transportability and identification

A causal quantity is transportable relative to a selection diagram when it is uniquely determined by the distributions available from the source and target domains under every pair of causal models compatible with the diagram. This definition is model-theoretic: transportability depends on the encoded causal structure, the locations of selection nodes, and the classes of data available in each domain.

Graphical identification proceeds by transforming expressions containing interventions and selection variables. The rules of do-calculus remove interventions, observations, or domain indicators when the corresponding d-separation conditions hold in suitably modified graphs. A successful reduction produces a transport formula involving only estimable source and target distributions.

One important graphical concept is (S)-admissibility. A set of variables (Z) is (S)-admissible for transporting the effect of (X) on (Y) when conditioning on (Z) separates (Y) from the relevant selection variables in the graph modified to represent intervention on (X). Under the associated assumptions, the target effect can be expressed by standardizing a source-domain causal effect over the target-domain distribution of (Z):

[ P^\ast(y \mid \operatorname{do}(x))

\sum_z P(y \mid \operatorname{do}(x), z) P^\ast(z). ]

This equation does not constitute a general rule for all selection diagrams. Its validity follows from the specific graphical separation relation and from the availability of the distributions on its right-hand side.

When no valid transformation exists, the target effect is not transportable from the stated information. Nontransportability can be demonstrated by constructing two compatible model pairs that agree on every available source and target distribution but assign different values to the target causal effect.

Relation to sample-selection models

A sample-selection model represents the process by which units enter an observed dataset. In graphical formulations, a sampling indicator (R) or (S) is an ordinary variable, and analysis conditions on the event that the unit was observed. Such conditioning can induce collider bias or other forms of selection bias.

Selection nodes in a transportability diagram perform a different function. They identify mechanisms that vary across domains and ordinarily remain external to the substantive causal system. The distinction depends on semantics rather than typography, because both literatures use similar symbols.

The two structures can occur together. A study can involve nonrandom sampling within the source population while also being transported to a target population governed by different mechanisms. The resulting causal model contains sampling variables for inclusion processes and selection nodes for cross-domain discrepancies, with each type participating differently in identification.

Scope and limitations

A selection diagram does not infer population differences directly from observed data. Its edges encode substantive assumptions about which mechanisms remain invariant. Statistical equality between two observed conditional distributions does not by itself establish structural invariance, because distinct causal mechanisms can generate the same observational distribution.

The framework also presupposes a specified causal graph shared at the level relevant to the analysis. Differences represented by selection nodes alter local mechanisms without replacing the underlying variable set or causal ordering. Domains requiring incompatible causal structures can be represented through broader model constructions, but they do not reduce to a single elementary selection diagram without additional variables or abstractions.

Latent common causes are represented by the conventions used for acyclic directed mixed graphs. Their presence can obstruct transportability even when the apparent locations of population differences are limited. The resulting identification problem therefore depends jointly on unobserved confounding and cross-domain variation.

See also

  • Causal inference, the study of effects defined by interventions rather than associations
  • Transportability, the formal transfer of causal information between domains
  • Do-calculus, the graphical calculus used to transform interventional distributions
  • Structural causal model, the mathematical framework underlying selection diagrams
  • External validity, the relation between study results and populations outside the original study
  • Selection bias, distortion arising from conditioning on processes that determine observation
  • D-separation, the graphical criterion used to derive conditional independences in causal graphs