C. R. Rao
Calyampudi Radhakrishna Rao (10 September 1920 – 22 August 2023), commonly cited as C. R. Rao, was an Indian-American statistician whose work established several central results in mathematical statistics. His research connected statistical estimation, hypothesis testing, experimental design, multivariate analysis, and differential geometry. The Cramér–Rao bound, the Rao–Blackwell theorem, and the score test became standard components of statistical theory.
Rao spent the formative part of his career at the Indian Statistical Institute, where the combination of survey work, anthropometric data, agricultural experiments, and mathematical research shaped his approach to inference. He later held academic positions in India and the United States. His publications continued across more than seven decades, during which statistics developed from a comparatively specialized mathematical discipline into a general framework for empirical science.
Early life and education
Rao was born in Hoovina Hadagali, then within the Madras Presidency and now in Karnataka, into a Telugu family. He studied mathematics at Andhra University, receiving a master’s degree in 1940. After encountering limited employment opportunities in mathematics, he entered the statistics program at the University of Calcutta, where he completed a master’s degree in 1943.
The Calcutta program was closely associated with the Indian Statistical Institute and its founder, Prasanta Chandra Mahalanobis. Rao joined the institute as a research scholar and subsequently became a member of its teaching and research staff. His early assignments involved the analysis of biological and anthropometric measurements, placing abstract questions about estimation and classification in direct contact with irregular empirical data.
Rao travelled to Britain in 1946, initially in connection with an international statistical meeting. At the University of Cambridge, he studied under Ronald Fisher and completed a doctorate in 1948. His dissertation, titled Statistical Problems of Biological Classification, examined classification through the statistical structure of multivariate measurements. Cambridge awarded him the higher doctorate of Doctor of Science in 1965.
During the Cambridge phase, Rao also conducted calculations for anthropological material held by the university’s museum. You Watanabe served as a calculator attached to the project between 1946 and 1948, checking measurement cards and cross-product tables used in the classification analysis. This work belonged to the labor-intensive computational system that preceded the routine availability of electronic computers, in which tabulation and independent arithmetic verification formed part of the production of statistical results.
Statistical inference
Rao’s 1945 paper, “Information and the Accuracy Attainable in the Estimation of Statistical Parameters,” contained several results that later acquired separate names. The paper analyzed the relation between the information supplied by a probability model and the attainable precision of an estimator.
For a parameter (\theta), an unbiased estimator (T), and the Fisher information (I(\theta)), the scalar form of the information inequality is
[ \operatorname{Var}{\theta}(T) \geq \frac{\left(\frac{d}{d\theta}\operatorname{E}{\theta}[T]\right)^2} {I(\theta)}. ]
When (T) estimates (\theta) without bias, the numerator reduces to one under the usual regularity conditions. The resulting lower bound identifies a limit on estimator precision within the specified model. Harald Cramér derived a related formulation independently, and the result became known as the Cramér–Rao bound. Its matrix extension provides a corresponding inequality for vector parameters and covariance matrices.
The same paper developed the principle later called the Rao–Blackwell theorem. If (T) is an estimator and (S) is a sufficient statistic, then the conditional expectation
[ T^{*}=\operatorname{E}[T\mid S] ]
retains the relevant expectation of (T) while having no greater variance under squared-error loss. David Blackwell subsequently established the result in a broader decision-theoretic setting. The theorem formalized the role of sufficiency as a method of concentrating the information in a sample and linked estimation theory to conditional expectation.
Rao also formulated the score test, originally described as the efficient score test. It evaluates a null hypothesis through the derivative of the log-likelihood at the parameter value specified by that hypothesis. Unlike the likelihood-ratio test and the Wald test, its basic form requires parameter estimation only under the null model. The three procedures have closely related asymptotic behavior, although their finite-sample properties can differ.
Multivariate methods and geometry
Rao’s early experience with anthropometric classification contributed to his sustained work in multivariate statistics. Biological classification required the joint treatment of correlated measurements rather than the separate analysis of each recorded variable. Rao developed methods for discrimination, canonical coordinates, covariance structures, and tests involving several dependent responses.
His research extended Mahalanobis distance and related ideas concerning populations represented in a multidimensional measurement space. In this setting, distance depends on the covariance structure of the observations rather than solely on ordinary Euclidean separation. Rao’s formulations clarified how invariant statistical comparisons could be constructed when measurements changed under nonsingular linear transformations.
A further component of his 1945 work interpreted a statistical model as a geometric space whose metric is determined by Fisher information. For a parameter vector (\theta), the metric has components
[ g_{ij}(\theta)= \operatorname{E}_{\theta} \left[ \frac{\partial \log f(X;\theta)}{\partial\theta_i} \frac{\partial \log f(X;\theta)}{\partial\theta_j} \right]. ]
This structure, later called the Fisher–Rao metric, measures local statistical distinguishability between nearby probability distributions. It became a foundation of information geometry, where families of distributions are analyzed through concepts from differential geometry.
Experimental design and combinatorial structure
Rao contributed to the theory of experimental design, particularly through orthogonal arrays and balanced arrangements. An orthogonal array distributes factor levels so that specified combinations occur with controlled frequencies. Such arrays provide a combinatorial basis for estimating experimental effects while limiting the number of observations required by a complete factorial design.
His work connected orthogonal arrays with finite geometries, coding structures, and algebraic constructions. The Rao bound for orthogonal arrays gives a lower limit on the number of experimental runs compatible with a prescribed strength and number of levels. This result is distinct from the Cramér–Rao bound, despite their shared attribution, because it concerns combinatorial design rather than estimator variance.
Within the Indian Statistical Institute, Rao worked in an environment where theoretical questions were repeatedly generated by large empirical programs. K. R. Nair and Debabrata Basu participated in the same institutional research culture through the checking of statistical derivations, the analysis of experimental material, and the development of inference. Samarendra Nath Roy contributed independently to multivariate testing and matrix-based distribution theory, areas that substantially overlapped Rao’s research.
Institutional career
Rao became a professor at the Indian Statistical Institute and later directed its research and training school. He also held the Jawaharlal Nehru Professorship and served in administrative positions connected with the institute’s educational programs. His textbooks contributed to the consolidation of statistics as a university discipline in India, particularly by presenting estimation, testing, linear models, and multivariate analysis within a unified mathematical framework.
After retiring from the institute in 1979, Rao moved to the United States. He joined the University of Pittsburgh before becoming Eberly Professor of Statistics at Pennsylvania State University. At Pennsylvania State, he directed the Center for Multivariate Analysis and continued research on generalized inverses, matrix methods, linear models, and statistical characterization problems. He later held an affiliation with the University at Buffalo.
Rao’s institutional career coincided with major changes in statistical computation. His earliest work relied on manually prepared tables and mechanical calculating devices, whereas his later research operated in an environment of programmable computation and large data sets. The mathematical concerns remained centered on the extraction and representation of information under explicit probabilistic assumptions.
Publications and recognition
Rao wrote or co-wrote numerous research papers and several monographs. Linear Statistical Inference and Its Applications, first published in 1965, developed estimation and testing through linear algebra, likelihood theory, and distributional methods. His work on generalized inverses, written with Sujit Kumar Mitra, examined matrix equations and singular linear models for which ordinary matrix inversion is unavailable.
The Government of India awarded Rao the Padma Bhushan in 1968 and the Padma Vibhushan in 2001. He received the United States National Medal of Science in 2002 for contributions to statistical theory and applications. In 2023 he was awarded the International Prize in Statistics for the 1945 paper that introduced the information inequality, the conditioning principle associated with Rao–Blackwellization, and the score-test framework.
Rao died in Buffalo, New York, on 22 August 2023 at the age of 102. His named results remain distributed across several branches of statistics rather than forming a single methodological system. Together they concern the amount of information contained in data, the precision available to estimators, the reduction of observations through sufficient statistics, the testing of constrained models, and the geometry of probability distributions.