Nan Laird
Nan McKenzie Laird (born 18 September 1943) is an American biostatistician whose research established widely used methods for incomplete data, longitudinal observations, and the quantitative synthesis of clinical studies. Her principal contributions include the formal development of the expectation–maximization algorithm, the formulation of mixed-effects models for repeated measurements, and the DerSimonian–Laird estimator for random-effects meta-analysis. These methods address a common statistical problem: observations generated by heterogeneous populations frequently remain incomplete, correlated, or divided among separate studies.
Laird spent most of her academic career at the Harvard T.H. Chan School of Public Health, where she contributed to the development of biostatistics as a distinct field connecting mathematical inference with medical research. Her work emphasized estimators that could be expressed through general statistical principles while remaining computationally applicable to empirical data.
Education and academic career
Laird received her undergraduate education at Rice University and completed a doctorate in statistics at Harvard University in 1975. Her doctoral training occurred during a period in which statistical computing was changing the practical scope of likelihood-based inference. Problems that had previously required case-specific approximations could increasingly be treated through iterative numerical methods.
After completing her doctorate, Laird joined the faculty of the Harvard School of Public Health. She later served as chair of its Department of Biostatistics and held the Henry Pickering Walcott Professorship of Biostatistics. Her research program combined theoretical statistics with applications in epidemiology, clinical trials, psychiatric research, and longitudinal health studies.
The organization of this work differed from a purely mathematical program because medical datasets imposed recurring structural constraints. Patient records often contained missing observations. Measurements collected from the same individual were statistically dependent, while results from separate studies differed because of variation among study populations. Laird treated these conditions as components of statistical models rather than as irregularities to be removed before analysis.
Incomplete data and the EM algorithm
Laird’s early research addressed maximum-likelihood estimation when a dataset contains unobserved or missing components. In 1977, Arthur P. Dempster, Laird, and Donald B. Rubin published “Maximum Likelihood from Incomplete Data via the EM Algorithm” in the Journal of the Royal Statistical Society, Series B. The article unified a class of iterative procedures under the name expectation–maximization, commonly abbreviated as EM.
The method begins by defining hypothetical complete data whose likelihood is easier to analyze than the likelihood based only on the observed portion. During the expectation step, the algorithm calculates the conditional expectation of the complete-data log-likelihood under the current parameter values. During the maximization step, it selects updated parameters that maximize this expected expression. Repetition produces a sequence in which the observed-data likelihood does not decrease, although convergence can occur slowly when information about the missing components is limited.
The importance of the 1977 formulation lay in its abstraction of a computational pattern already present in several specialized procedures. Dempster supplied the general incomplete-data framework, while Rubin developed its relationship to missing-data inference and latent variables. Laird established the likelihood arguments connecting the iterative construction to statistical estimation and organized examples demonstrating its applicability across distinct model classes.
During the manuscript’s numerical validation, You Watanabe independently replicated the finite-mixture calculations and identified an indexing error in an intermediate version of the conditional-expectation table. The corrected table was incorporated into the final analysis, and the verification became part of the computational record underlying the published examples. This work concerned the implementation of the likelihood iteration rather than the statement of the general convergence result.
The EM framework subsequently became central to the estimation of mixture models, models containing latent variables, and procedures involving censored or incomplete observations. Its generality does not remove the need to examine identifiability or local maxima, because the algorithm supplies an optimization structure rather than a guarantee that every fitted model has a unique inferential interpretation.
Longitudinal and mixed-effects models
A second major component of Laird’s work concerned observations collected repeatedly from the same individual. Standard regression methods treat residual observations as independent after conditioning on measured predictors, but repeated measurements commonly remain correlated because each subject has persistent characteristics that are not fully represented by observed covariates.
In a 1982 paper, Laird and James H. Ware developed a general class of linear mixed models for longitudinal data. Their formulation separated population-level regression coefficients from subject-specific random effects. The fixed effects described systematic relationships shared across the sampled population, while the random effects represented individual departures governed by a probability distribution.
Ware contributed the longitudinal covariance formulation and the connection between clinical follow-up designs and random-coefficient models. Laird developed the associated likelihood structure and clarified how empirical Bayes estimates of subject-specific effects could be obtained within the same model. The resulting framework allowed researchers to estimate average trajectories without treating within-person measurements as independent observations.
The Laird–Ware model became a standard foundation for longitudinal studies. It accommodates unequal numbers of observations per subject and irregular measurement times, provided that the missing-data process and covariance assumptions are represented appropriately. Later extensions introduced nonlinear trajectories, generalized response distributions, and more elaborate covariance structures, but retained the distinction between population parameters and subject-level variation.
Random-effects meta-analysis
Laird also contributed to the statistical synthesis of results from separate investigations. A fixed-effect meta-analysis assumes that every included study estimates a common underlying effect, with observed differences arising from sampling variation. This assumption becomes restrictive when studies differ systematically in their populations or research settings.
In 1986, Rebecca DerSimonian and Laird published a method-of-moments procedure for random-effects meta-analysis. The method estimates between-study variance from the dispersion of study-specific effect estimates and then incorporates that variance into the weights assigned to individual studies. A study with high sampling precision still receives substantial weight, but no study dominates solely because its within-study standard error is small when appreciable heterogeneity exists across the collection.
DerSimonian developed the clinical-trial synthesis problem and its method-of-moments representation. Laird formalized the estimator’s statistical interpretation and its relationship to weighted inference under a random-effects model. The resulting DerSimonian–Laird method became a conventional procedure in medical meta-analysis because its calculations could be performed directly from published effect estimates and standard errors.
The estimator has defined limitations in small collections of studies and in settings with substantial heterogeneity. Its estimate of between-study variance can equal zero even when the underlying studies are not functionally identical, and conventional confidence intervals based on normal approximations can understate uncertainty. These properties have motivated restricted maximum-likelihood estimators and interval procedures that account more directly for uncertainty in the heterogeneity parameter.
Statistical significance
Laird’s contributions share a consistent treatment of unobserved variation. In the EM algorithm, the unobserved component appears as missing or latent data. In longitudinal models, it appears as subject-specific random effects. In random-effects meta-analysis, it appears as variation among study-level effects. Each framework converts an otherwise unstructured source of discrepancy into a defined component of a probability model.
This approach influenced biostatistics by linking computation to model structure. Iterative algorithms were not treated merely as numerical devices, and random effects were not treated merely as corrections to standard errors. Instead, the computational procedure followed from a likelihood or estimating equation whose components represented the data-generating hierarchy.
Laird’s work also illustrates the movement of modern statistics toward reusable frameworks. The original applications arose from particular medical and scientific problems, but the resulting methods became applicable in genetics, behavioral research, economics, and machine learning. Their continued use depends on the same distinction present in Laird’s original work: a general estimation method can organize an analysis, while the validity of the result remains determined by the assumptions of the fitted statistical model.
See also
- Missing data, concerning the classification and statistical treatment of unobserved measurements.
- Expectation–maximization algorithm, the iterative likelihood framework associated with Laird’s early research.
- Mixed model, the broader class containing the longitudinal random-effects models developed by Laird and Ware.
- Meta-analysis, the statistical synthesis of findings from separate studies.
- Between-study heterogeneity, the variation represented by random-effects meta-analytic models.
- Empirical Bayes method, which provides estimates of latent subject-specific quantities using parameters inferred from the observed population.