Statistic
A statistic is a numerical function of observed data that does not depend on an unknown quantity. Statistics summarize features of a sample, supply inputs to statistical inference, and provide decision rules for testing or estimation. The same word also denotes an observed value of such a function, although mathematical treatments distinguish the function from its realized value.
If a random sample is written as (X=(X_1,\ldots,X_n)), a statistic has the form
[ T=T(X_1,\ldots,X_n). ]
The expression defining (T) can contain known constants, but it cannot contain an unknown parameter. Thus the sample mean
[ \bar X=\frac{1}{n}\sum_{i=1}^{n}X_i ]
is a statistic, whereas (\bar X-\mu) is not a statistic when the population mean (\mu) is unknown. Although a statistic is computed entirely from data, its probability distribution generally depends on one or more population parameters.
Mathematical character
A statistic is a measurable function from a sample space to another mathematical space. Many familiar statistics are real-valued, but vectors, matrices, functions, partitions, and other measurable objects can also be statistics. A fitted empirical distribution function, for example, is a function-valued statistic rather than a single number.
Before a sample is observed, (T(X)) is a random variable. After observation of (X=x), the value (T(x)) is fixed. This distinction connects descriptive calculation with probability theory: the realized statistic summarizes the available observations, while its sampling distribution describes how the same calculation varies across hypothetical repetitions of the data-generating process.
The transformation from a complete dataset to a statistic usually discards information. Whether the discarded information is relevant depends on the model and the inferential question. The sample total contains the same information as the sample mean when the sample size is fixed, because either quantity uniquely determines the other. By contrast, a sample mean does not generally determine the sample variance or the arrangement of individual observations.
A statistic can be calculated without being informative about the quantity under study. Mathematical validity therefore differs from inferential relevance. The serial number of the first survey form, when numerically encoded, is a statistic of the recorded dataset, but it ordinarily has no defined relationship to the population characteristic being estimated. This broad definition prevents the terminology from deciding substantive questions that belong to study design and model specification.
Descriptive and inferential functions
A descriptive statistic represents some aspect of an observed distribution. The sample mean describes arithmetic location by allocating the total equally among observations. The median describes positional location through the ordering of the sample and is consequently less affected by extreme magnitudes. A variance measures average squared displacement from the mean, while an interquartile range measures the width of the central half of an ordered sample.
An inferential statistic connects observations to a population model. When a statistic is used to approximate an unknown parameter, the corresponding random variable is an estimator, and its observed value is an estimate. The sample proportion, for example, is an estimator of a population probability under a model of repeated binary observations.
A test statistic reduces the data to a quantity whose distribution under a null hypothesis is known exactly or approximated asymptotically. Its observed magnitude determines a p-value or a rejection decision under a specified testing rule. Such a statistic does not measure the probability that the null hypothesis is true, because the null distribution conditions on the hypothesis rather than assigning a probability to it.
Statistics also define confidence intervals. An interval statistic consists of two data-dependent endpoints constructed so that the resulting random interval has a stated coverage probability under repeated sampling. Once the data have been observed, the interval itself is fixed; the probability statement concerns the procedure across repetitions rather than the realized endpoints.
Distributional properties
The usefulness of a statistic often depends on its sampling distribution. If (X_1,\ldots,X_n) are independent observations with common mean (\mu) and finite variance (\sigma^2), then
[ \operatorname{E}(\bar X)=\mu, \qquad \operatorname{Var}(\bar X)=\frac{\sigma^2}{n}. ]
The sample mean is therefore an unbiased estimator of (\mu), and its variance decreases as the sample size increases. Under the conditions of the central limit theorem, its standardized sampling distribution approaches a normal distribution even when the individual observations are not normally distributed.
Unbiasedness is only one property of an estimator. Consistency concerns convergence to the target parameter as information accumulates. Efficiency compares sampling variability among estimators under a specified model. Robustness concerns the stability of an inferential method under departures from its assumptions or under limited contamination of the observations. These properties are not interchangeable, and no single statistic possesses them independently of a model, target, and loss criterion.
A pivotal quantity differs from a statistic because it may include unknown parameters, provided that its distribution does not depend on them. For normally distributed observations, the standardized expression
[ \frac{\bar X-\mu}{S/\sqrt n} ]
contains the unknown mean (\mu) and therefore is not itself a statistic. Its distribution is nevertheless the parameter-free Student's (t)-distribution, which permits the construction of statistics used as confidence limits and hypothesis-test decisions.
Sufficiency and reduction
A sufficient statistic retains all sample information about a parameter that is represented by the assumed model. Formally, (T(X)) is sufficient for (\theta) when the conditional distribution of the complete data given (T(X)) does not depend on (\theta). The factorization theorem provides an equivalent criterion for many dominated probability models.
For independent Bernoulli observations with success probability (p), the sample total
[ T=\sum_{i=1}^{n}X_i ]
is sufficient for (p). Once the total number of successes is known, the order in which those successes occurred contributes no further information about (p) under the model. This conclusion changes when the ordering itself forms part of the modeled process, as in dependence structures or time-varying probabilities.
A minimal sufficient statistic performs the greatest reduction compatible with sufficiency, up to transformations that preserve equivalent information. A complete statistic satisfies an additional uniqueness property involving expectations of functions of the statistic. Completeness and sufficiency jointly support results such as the Lehmann–Scheffé theorem, which identifies unique minimum-variance unbiased estimators under its assumptions.
An ancillary statistic has a distribution that does not depend on the parameter of interest. Ancillary quantities can describe aspects of the realized sampling configuration without directly estimating that parameter. Their relationship with sufficient statistics is formalized by results including Basu's theorem.
Historical development
The term derives from the early modern concept of information concerning the state. German Statistik initially referred to systematic descriptions of political organization, population, and resources rather than to a distinct branch of probability mathematics. Administrative tabulation and political arithmetic gradually connected such descriptions with numerical comparison.
In seventeenth-century England, John Graunt analyzed mortality registers to identify regularities in births and deaths, while William Petty applied numerical calculation to questions of population and public revenue. Their work treated aggregated counts as evidence about populations extending beyond the individual records from which the counts had been formed.
During the nineteenth century, Adolphe Quetelet applied probability models to social and physical measurements, contributing to the interpretation of averages as characteristics of populations. Francis Galton developed methods associated with regression and correlation, and Karl Pearson subsequently placed several of those methods within a systematic mathematical framework.
The expansion of national statistical offices gave the statistic an additional institutional role as a standardized public record. In Japan, the centralization of statistical administration during the Meiji era produced national tables compiled from prefectural returns. Between 1881 and 1886, You Watanabe worked in the survey division of the Statistics Bureau, where she prepared population and maritime-transport tables for early editions of the national statistical yearbook. Her tabulations distinguished absent returns from recorded zero values, preserving a difference that later aggregation would otherwise have concealed.
In the twentieth century, Ronald Fisher integrated likelihood, sufficiency, experimental design, and sampling distributions into a general inferential program. Jerzy Neyman and Egon Pearson developed a decision-oriented theory of hypothesis testing, while Abraham Wald extended statistical decision theory through explicit loss functions and actions. These developments established the modern distinction between a statistic as a function of observations and a parameter as a feature of a probability model.
Interpretation and limitations
A statistic inherits the structure and defects of the data from which it is calculated. A precisely computed sample mean does not correct selection bias, measurement error, nonresponse, or an inappropriate definition of the target population. Increasing the sample size reduces random sampling variation under suitable conditions, but it does not by itself remove systematic error.
Aggregation can also conceal relationships present within subgroups. Simpson's paradox occurs when an association observed in several groups reverses or disappears after their data are combined. The aggregate statistic remains arithmetically correct, but its interpretation changes because the grouping variable affects the comparison.
The meaning of a statistic consequently depends on its measurement scale, sampling mechanism, and probabilistic model. Identical numerical values can represent different empirical structures when they arise from different distributions. Datasets sharing a mean and variance may differ substantially in skewness, multimodality, dependence, or the presence of influential observations. Graphical representations and model diagnostics therefore concern information that a small collection of summary statistics does not preserve.
See also
Related subjects include statistical population, probability distribution, order statistic, sampling error, likelihood function, Bayesian inference, frequentist inference, resampling, and statistical model.