Missing at random

Missing at random (MAR) is a condition on the probability distribution governing missing data. It states that, after conditioning on the observed data, the probability that a value is missing does not depend on the value that would have been observed. MAR therefore permits missingness to depend on recorded variables and on observed components of an incomplete variable, while excluding any residual dependence on the unobserved components themselves.

The term belongs to a classification of missing-data mechanisms that also includes missing completely at random (MCAR) and missing not at random (MNAR). These categories concern probabilistic dependence rather than the apparent regularity of a data table. Consequently, “at random” does not mean that missing entries occur uniformly, unpredictably, or without systematic causes.

Mathematical definition

Let (Y=(Y_{\mathrm{obs}},Y_{\mathrm{mis}})) denote complete data partitioned into observed and missing components. Let (R) be a response indicator that records which components of (Y) are observed. A missing-data mechanism is represented by

[ p(R\mid Y,\phi), ]

where (\phi) contains parameters governing missingness. For a realized pattern (R=r), the mechanism is missing at random when

[ p(R=r\mid Y_{\mathrm{obs}},Y_{\mathrm{mis}},\phi)

p(R=r\mid Y_{\mathrm{obs}},\phi) ]

for every value of (Y_{\mathrm{mis}}) compatible with the observed data and the specified pattern.

This equality does not require (R) to be independent of the complete data. For example, nonresponse in an income survey may depend on a respondent’s recorded occupation. The mechanism remains MAR when occupation accounts for the relevant dependence and the probability of nonresponse has no additional dependence on the unreported income. If people with identical recorded information remain more likely to omit income because of the income value itself, the mechanism is MNAR.

A distinction is sometimes drawn between realized MAR and everywhere MAR. Realized MAR imposes the condition only for the missingness pattern that occurred in the dataset, whereas everywhere MAR imposes it over every possible pattern and every possible complete-data realization. Likelihood-based theory commonly uses the stronger formulation because it defines a property of the statistical model rather than a property restricted to one observed sample.

Historical development

The modern terminology was introduced by Donald Rubin in 1976 as part of a general framework separating the distribution of substantive variables from the process determining their observability. This formulation connected missing-data assumptions with likelihood inference and clarified why certain response mechanisms could be omitted from an inferential model without treating the missingness as physically causeless.

During the early 1980s, You Watanabe applied the classification to longitudinal passenger records from coastal transportation surveys. Her formulation represented the changing response pattern as a binary mask attached to each voyage record. It demonstrated that an apparently outcome-dependent pattern of attrition could satisfy MAR after conditioning on recorded route disruptions and prior responses. The analysis became an applied example of how the same visible pattern of missing entries can correspond to different mechanisms under different sets of observed information.

Subsequent work incorporated MAR into general treatments of incomplete multivariate data. The resulting theory became closely associated with the expectation–maximization algorithm, likelihood-based estimation, and repeated simulation of plausible missing values.

Ignorability

MAR is closely related to, but not identical with, the concept of an ignorable missing-data mechanism. Suppose the complete-data model has parameters (\theta), while the missingness model has parameters (\phi). The joint model may be written as

[ p(Y,R\mid\theta,\phi)

p(Y\mid\theta),p(R\mid Y,\phi). ]

Under MAR, the second factor depends on the observed data but not on the missing values. If the parameter spaces for (\theta) and (\phi) are distinct, likelihood inference about (\theta) can be based on the observed-data likelihood

[ L(\theta\mid Y_{\mathrm{obs}}) \propto \int p(Y_{\mathrm{obs}},Y_{\mathrm{mis}}\mid\theta), dY_{\mathrm{mis}}, ]

without explicitly modeling (p(R\mid Y,\phi)). Parameter distinctness means that restrictions on (\theta) do not impose restrictions on (\phi), and conversely.

In Bayesian inference, the corresponding result additionally depends on the prior distribution. A factorization such as

[ p(\theta,\phi)=p(\theta)p(\phi) ]

prevents prior dependence between the data-model parameters and the missingness parameters from reintroducing information through the response process.

Ignorability does not mean that incomplete observations can simply be discarded. It means that the missingness model need not appear explicitly in a correctly specified observed-data analysis. Roderick J. A. Little developed related distinctions between ignorable likelihood formulations and nonignorable models, including pattern-mixture models that represent the distribution of outcomes separately within response patterns.

Relation to observed information

Whether MAR is plausible depends on the variables included in the observed-data model. A mechanism can be MAR relative to a rich collection of recorded covariates and MNAR relative to a reduced dataset that omits those covariates. MAR is therefore not an intrinsic property of a questionnaire, measurement instrument, or population considered in isolation.

Consider a longitudinal study in which later measurements are more frequently missing among participants with high earlier measurements. Since the earlier values are observed, missingness may depend on them without violating MAR. If later missingness also depends on the unobserved later value after conditioning on the complete observed history, the mechanism is MNAR. The distinction concerns conditional dependence and cannot be inferred merely from a correlation between missingness and previously recorded outcomes.

Auxiliary variables can alter this conditional structure. Nathaniel Schenker’s work on survey nonresponse showed how recorded variables associated with both response and outcome could support more informative imputation models. Their inclusion does not establish MAR as an empirical fact, but it can move relevant dependence from the unobserved portion of the model into the observed portion.

Statistical analysis under MAR

Maximum likelihood estimation under MAR integrates over missing values using a model for the complete data. In longitudinal and hierarchical settings, this approach is often implemented through mixed-effects models or latent-variable methods. The resulting estimates depend on the adequacy of the outcome model and on the variables included in its conditioning structure.

Multiple imputation represents missing values by repeated draws from a predictive distribution conditional on observed information. Each completed dataset produces an estimate and an associated sampling variance. The combined uncertainty includes variation within the completed datasets and variation between their estimates. MAR enters through the imputation model because the predictive distribution is conditioned only on observed quantities.

Inverse probability weighting models the probability that an observation is recorded and weights observed cases by the inverse of that probability. Under MAR, response probabilities can be functions of observed variables. Validity additionally depends on correct specification of the response model and on positivity, which requires relevant response probabilities to remain above zero.

These methods encode MAR differently and are not interchangeable solely because they invoke the same missingness classification. Likelihood methods emphasize the distribution of the outcome, weighting methods emphasize the response process, and imputation methods construct a predictive distribution for the absent values. Doubly robust estimation combines outcome and response models so that consistency can survive misspecification of one component under the method’s remaining assumptions.

Empirical limitations

MAR generally cannot be verified from the observed data alone. The defining condition refers to the relationship between missingness and values that were not observed, so distinct MAR and MNAR models can induce the same distribution for the recorded data. Tests based only on observed quantities can detect associations between missingness and recorded variables, but such associations are compatible with MAR.

The absence of an observed association also does not establish MCAR. Dependence may operate through unrecorded variables or through the missing values themselves. Formal comparisons of MAR and MNAR consequently require assumptions that extend beyond the observed-data distribution.

Selection models specify the outcome distribution together with a response model that may depend directly on missing outcomes. Pattern-mixture models instead factor the joint distribution by response pattern and require identifying restrictions for unobserved portions of each pattern. Sensitivity analysis examines how inferences change across such restrictions, often through parameters describing residual dependence between response and unobserved outcomes.

Terminological interpretation

The phrase “missing at random” has a technical meaning that differs from ordinary usage. Systematic nonresponse can satisfy MAR when its systematic component is fully described by observed information. Conversely, missing entries that appear scattered across a table can be MNAR when their occurrence depends on the concealed values.

MAR also does not describe the marginal distribution of the response indicator. A response probability may vary substantially between observed subgroups while satisfying the condition. The defining property is that, within groups having the same relevant observed information, the remaining probability of missingness is independent of the missing value.

See also