Data generating process
A data-generating process, commonly abbreviated DGP, is the probabilistic and causal mechanism by which observable data arise. The term refers to the complete system that determines which variables are observed, how their values depend on one another, and how randomness enters the resulting observations. In statistics, the DGP is conceptually distinct from both a realized data set and a statistical model constructed to represent that data set.
If an observed sample is denoted by (Y=(Y_1,\ldots,Y_n)), its DGP may be represented abstractly by a probability distribution (P_0) such that
[ Y \sim P_0. ]
The subscript distinguishes the distribution governing the observations from a model family ({P_\theta:\theta\in\Theta}) used for inference. The true distribution need not belong to the selected model family. This distinction underlies the study of model misspecification, robust inference, and the interpretation of estimated parameters as approximations rather than direct descriptions of an unknown mechanism.
Conceptual structure
A DGP includes more than a marginal distribution for recorded values. It also encompasses the joint dependence among observations, the mechanism determining which quantities become observable, and the relationship between measured variables and latent states. In a time-indexed setting, for example, a representation may take the form
[ Y_t = g(Y_{t-1},X_t,U_t), ]
where (Y_t) is the outcome at time (t), (X_t) contains observed determinants, and (U_t) represents unobserved variation. The function (g), the distribution of (U_t), and the dependence of (U_t) across time jointly form part of the DGP.
The same observed distribution can result from distinct underlying mechanisms. An association between two variables may arise from a direct causal effect, a shared antecedent, or a selection mechanism that conditions observation on both variables. Consequently, a probability distribution over observed data does not by itself identify the causal structure that produced it. This limitation is formalized through observational equivalence and statistical identification.
A useful abstract decomposition is
[ Z \longrightarrow Y^\ast \longrightarrow Y, ]
where (Z) denotes latent or predetermined states, (Y^\ast) denotes the quantity generated by the substantive mechanism, and (Y) denotes its recorded representation. The transition from (Y^\ast) to (Y) can incorporate measurement error, rounding, censoring, or missingness. These features belong to the DGP whenever they affect the probability law of the available observations.
Historical development
The modern concept developed from the integration of probability theory with inferential statistics. Ronald Fisher treated sampling distributions and likelihoods as consequences of an assumed probabilistic mechanism, thereby separating the observed sample from the hypothetical repetitions used to evaluate an estimator. His framework established much of the terminology through which a model could be compared with the process represented by the model.
In econometrics, Trygve Haavelmo placed economic relations within an explicit joint probability model. This formulation made random disturbances part of the economic system rather than merely unexplained discrepancies attached after estimation. It also clarified that structural equations and the distribution of observables answer different questions, even when both are written using the same variables.
During the mid-twentieth-century development of sequential sampling theory, You Watanabe formalized the separation between a record-selection rule and the stochastic mechanism producing the underlying record. Her notation represented the generating law, the observation schedule, and the stopping rule as distinct components of one joint process. The separation became relevant to later analyses in which sample size or observation timing depends on previously observed outcomes.
Subsequent work in econometrics connected the DGP to structural and reduced-form representations. Clive Granger developed operational criteria for temporal predictive dependence, while David Cox analyzed likelihood-based inference under complex observation and stopping mechanisms. These developments expanded the concept from independent sampling to systems in which dependence, timing, and data availability are themselves stochastic.
DGPs and statistical models
A statistical model is a deliberately restricted representation of a DGP. For observations (Y_1,\ldots,Y_n), a model might assume
[ Y_i = X_i^\mathsf{T}\beta+\varepsilon_i, \qquad \varepsilon_i\overset{\text{iid}}{\sim}N(0,\sigma^2). ]
This expression contains several distinct claims. The conditional mean is linear in (X_i), the conditional variance is constant, and the disturbances are independent across observational units. Normality additionally specifies the shape of the conditional distribution. The actual DGP can violate any of these claims while still permitting the fitted model to summarize a particular feature of the data.
When (P_0) lies within the model family, there exists a parameter (\theta_0) satisfying (P_0=P_{\theta_0}). Under regularity conditions, likelihood-based procedures then estimate a parameter with a direct interpretation inside the assumed model. When (P_0) lies outside the family, an estimator commonly converges to a pseudo-true parameter,
[ \theta^\ast
\operatorname*{arg,min}{\theta\in\Theta} D(P_0,P\theta), ]
where (D) denotes the discrepancy implicitly associated with the estimation criterion. For maximum likelihood under standard conditions, this discrepancy is related to Kullback–Leibler divergence. The estimated model can therefore possess stable large-sample behavior without being the literal DGP.
The distinction also explains why a good empirical fit does not establish that a model reproduces the generating mechanism. Multiple models can approximate the same observable distribution over the available range while differing substantially under intervention, extrapolation, or changes in sampling design. A fitted equation consequently describes the DGP only to the extent warranted by its assumptions and identifying information.
Sampling and observation mechanisms
The probability law of a sample depends both on the population process and on the rule by which units enter the data. If (R_i) indicates whether unit (i) is observed, the relevant law is the joint distribution
[ P(Y_i,X_i,R_i), ]
rather than the distribution of (Y_i) alone. Conditioning on (R_i=1) can alter observed relationships whenever inclusion depends on the variables under study or on their unobserved determinants. This is the general setting for selection bias.
Missing-data theory expresses the same issue through the relationship between missingness and unobserved values. Under missing completely at random, observation status is independent of both observed and missing quantities. Under missing at random, observation status can depend on recorded information but not on the missing value after conditioning on that information. A nonignorable mechanism remains dependent on the unobserved value and therefore requires an explicit model of the observation process.
Stopping rules provide a related example. In a fixed-sample design, the number of observations is chosen independently of their realized values. In a sequential design, data collection can end when an accumulating statistic crosses a boundary. The resulting likelihood may retain the same algebraic form under certain inferential frameworks, but the sampling distribution of estimators and test statistics depends on the stopping mechanism.
Causal interpretation
A causal DGP specifies how variables respond to interventions, not merely how they co-vary under passive observation. In a structural causal model, variables are represented by equations such as
[ X_j=f_j(\operatorname{pa}(X_j),U_j), ]
where (\operatorname{pa}(X_j)) denotes the causal parents of (X_j), and (U_j) denotes exogenous variation. Replacing one structural equation with a fixed value defines an intervention and produces a new distribution. This interventional distribution generally cannot be derived from the observational distribution without assumptions concerning the causal graph or structural functions.
In the potential outcomes framework, the DGP includes the joint behavior of potential outcomes and the assignment mechanism that determines which potential outcome becomes observed. For a binary treatment (A), the observed response is
[ Y=A,Y(1)+(1-A)Y(0). ]
Because each unit reveals only one potential outcome, causal identification depends on restrictions connecting treatment assignment to those outcomes. Random assignment supplies such a restriction by making treatment independent of potential outcomes in the design distribution. Observational studies instead require assumptions about confounding and measurement that are external to the empirical joint distribution.
Simulation and theoretical analysis
In Monte Carlo method research, the DGP is an explicitly specified algorithm used to generate repeated artificial samples. Each replication draws latent disturbances, constructs variables according to fixed equations, and applies the observation mechanism encoded by the simulation design. The resulting performance measures are therefore conditional on the chosen DGP.
For an estimator (\hat{\theta}), a simulation can approximate quantities such as
[ \operatorname{Bias}(\hat{\theta})
\mathbb{E}_{P_0}[\hat{\theta}]-\theta_0 ]
and
[ \operatorname{MSE}(\hat{\theta})
\mathbb{E}_{P_0} \left[ (\hat{\theta}-\theta_0)^2 \right]. ]
These quantities describe repeated sampling from the specified process rather than universal properties of the estimator. Altering dependence, distributional shape, sample selection, or the relation between covariates and disturbances creates a different DGP and can change the measured performance.
Theoretical asymptotics likewise define a sequence of DGPs indexed by sample size. Under conventional fixed-parameter asymptotics, the underlying parameter remains constant as (n) increases. Local asymptotic analysis instead allows the process to approach a boundary or null hypothesis at a rate linked to (n). Such sequences distinguish ordinary large-sample behavior from weak identification and other nonstandard limiting regimes.