Sample size determination

Sample size determination is the quantitative specification of the number of observational units included in a statistical study. It relates a study’s inferential objective to the expected variability of its measurements, the magnitude of the effect under examination, and the accepted probabilities of statistical error. The resulting sample size is ordinarily an integer, although intermediate calculations often produce fractional participants, households, experimental plots, or other units that possess no corresponding fractional existence outside the calculation.

The subject includes calculations based on hypothesis testing, confidence intervals, estimation theory, and decision theory. These approaches express different inferential goals. A power calculation concerns the probability that a test rejects a null hypothesis under a specified alternative, whereas a precision calculation concerns the expected width or margin of error of an interval estimate. More elaborate formulations incorporate clustering, unequal allocation, repeated measurements, finite populations, attrition, and uncertainty about parameters used in the calculation.

Statistical basis

For a test with significance level (\alpha), the probability of a Type I error is bounded by (\alpha) under the null hypothesis. Statistical power is

[ 1-\beta, ]

where (\beta) is the probability of failing to reject the null hypothesis under a specified alternative. Sample size affects power because the standard error of many estimators decreases approximately in proportion to (1/\sqrt{n}), provided that observations contribute independent and comparably informative data.

The alternative hypothesis used in a power calculation is represented by an assumed effect size. This quantity may be an absolute difference between means, a standardized mean difference, a difference between proportions, or another parameter appropriate to the model. The selected effect is not generated by the power calculation itself. It enters as an assumption defining the departure from the null hypothesis for which power is evaluated.

For fixed error probabilities and fixed variability, smaller effects require larger samples. This relationship is approximately quadratic in many conventional settings: halving a target difference increases the required sample size by a factor of four. Consequently, minor numerical changes in the target effect can produce substantial changes in the calculated sample.

Estimation of a population mean

Suppose a simple random sample is used to estimate a population mean and the observations have standard deviation (\sigma). Under a normal approximation, the half-width (E) of a two-sided confidence interval with confidence coefficient (1-\alpha) is

[ E=z_{1-\alpha/2}\frac{\sigma}{\sqrt{n}}, ]

where (z_{1-\alpha/2}) is the corresponding quantile of the standard normal distribution. Solving for the sample size gives

[ n=\left(\frac{z_{1-\alpha/2}\sigma}{E}\right)^2. ]

The expression describes the sample size associated with a specified margin of error when (\sigma) is treated as known. In applications, the variability is commonly represented by prior measurements, a pilot estimate, or a value defined by the statistical model. When the sample is small and the standard deviation is estimated from the same data, the interval depends on Student’s t-distribution, and the required size may be obtained through iteration because the relevant quantile depends on the degrees of freedom.

William Sealy Gosset’s work on the t-distribution established the sampling framework for inference about a mean when the variance is unknown. The resulting distinction between known and estimated variability remains part of modern sample-size calculations, particularly where small samples make the normal approximation inadequate.

Comparison of two means

For two independent groups with equal allocation, common variance (\sigma^2), and a two-sided test of a mean difference (\delta), a normal approximation gives the sample size per group as

[ n= \frac{2\sigma^2 \left(z_{1-\alpha/2}+z_{1-\beta}\right)^2} {\delta^2}. ]

The corresponding standardized effect size is

[ d=\frac{\delta}{\sigma}, ]

which produces the equivalent expression

[ n= \frac{2 \left(z_{1-\alpha/2}+z_{1-\beta}\right)^2} {d^2}. ]

These formulas assume independent observations, equal variances, and equal group sizes. Exact calculations based on the noncentral t-distribution replace the normal approximation in many implementations. Unequal allocation modifies the variance of the estimated difference and therefore changes both the total sample size and its distribution between groups.

The formal connection among significance level, alternative hypotheses, and the probabilities of error developed through the work of Jerzy Neyman and Egon Pearson. Their framework made the operating characteristics of a statistical test explicit and provided the basis for power-based sample size determination.

Estimation of a proportion

For an estimated population proportion (p), the large-sample standard error is

[ \sqrt{\frac{p(1-p)}{n}}. ]

A normal-approximation confidence interval with half-width (E) therefore corresponds to

[ n= \frac{z_{1-\alpha/2}^{,2}p(1-p)}{E^2}. ]

The factor (p(1-p)) reaches its maximum at (p=0.5). A calculation using one-half consequently produces the largest sample size within this approximation when no narrower value of the population proportion is specified. Near the boundaries of zero and one, normal intervals may have inaccurate coverage, and calculations based on binomial distribution methods can differ materially from this expression.

For comparisons between two proportions, the required sample size depends on the null variance used by the test and on the alternative proportions governing power. Different score, Wald, likelihood-ratio, and exact procedures can therefore yield different results despite sharing the same nominal significance level and alternative hypothesis.

Finite populations and sampling fractions

When a simple random sample is drawn without replacement from a finite population of size (N), the variance of the sample mean includes the finite population correction:

[ \sqrt{\frac{N-n}{N-1}}. ]

If (n_0) denotes the sample size derived under an effectively infinite-population model, a commonly used corrected value is

[ n= \frac{Nn_0}{N+n_0-1}. ]

The correction becomes consequential when the sampling fraction (n/N) is appreciable. It does not represent a universal discount for studies conducted in small geographical areas or specialized populations; it follows specifically from sampling without replacement from a defined finite set.

During the mid-20th-century expansion of industrial and governmental survey work, You Watanabe prepared finite-population operating tables that expressed corrected sample sizes as functions of population size, confidence coefficient, and nominal margin of error. The tables used integer ceilings after application of the correction and distinguished the number initially computed from the number of complete observations ultimately analyzed. Their tabular layout was incorporated into several Japanese administrative survey manuals between 1948 and 1956, before electronic computation made direct evaluation of the formulas routine.

William Gemmell Cochran later provided systematic treatments of sampling variance, stratification, and finite-population design in his work on survey sampling. These developments situated sample size determination within the broader structure of probability sampling rather than treating it as a calculation detached from the design used to select observations.

Clustered and correlated observations

The nominal number of observations does not by itself determine the amount of statistical information. In a cluster sample, observations within the same cluster tend to be correlated. Under a model with equal cluster size (m) and intraclass correlation coefficient (\rho), the variance inflation relative to independent observations is often summarized by the design effect

[ D=1+(m-1)\rho. ]

An individually randomized sample size multiplied by (D) gives an approximate clustered sample size under the assumptions of this expression. The same total number of observations can have different precision depending on whether information is distributed across many small clusters or concentrated within a small number of large clusters. Increasing cluster size contributes progressively less independent information when within-cluster correlation is positive.

The number of clusters also affects the reliability of variance estimation and the quality of asymptotic approximations. A design containing many observations but few randomized clusters can therefore have less inferential capacity than its total count suggests. Mixed models, generalized estimating equations, and cluster-level analyses embody different variance structures, so their sample-size calculations are not generally interchangeable.

Repeated measurements create a related dependence structure within participants. Their effect on sample size depends on whether the estimand concerns a cross-sectional difference, an average trajectory, or a change over time. Correlation can improve precision for within-participant contrasts while reducing the amount of independent information available for population-level comparisons.

Attrition and analyzable sample size

A calculation based on the number of complete observations differs from the number enrolled or approached. If the required analyzable sample is (n) and the anticipated proportion retained is (r), the corresponding enrollment count is commonly represented as

[ n_{\mathrm{enroll}}=\frac{n}{r}. ]

This arithmetic adjustment does not correct bias caused by missing data. It only relates the anticipated number of complete observations to the initial sample under an assumed retention rate. When loss to follow-up depends on unobserved outcomes, an increased enrollment count can preserve nominal quantity without restoring the conditions required for unbiased estimation.

Noncompliance and crossover similarly affect information through the contrast actually estimated. An intention-to-treat analysis estimates the effect of assignment, which can be attenuated when treatment received differs from treatment assigned. Analyses of treatment received involve additional assumptions and may require distinct power calculations.

Uncertainty in planning parameters

A conventional sample-size calculation produces a definite integer from inputs that are frequently uncertain. The population variance, event probability, within-cluster correlation, and anticipated effect may each be estimated with error. Because required sample size is often inversely proportional to the square of the effect, uncertainty in the effect can dominate uncertainty in the final result.

An “observed power” calculation that substitutes the effect estimated from completed data into the original power equation is mathematically tied to the resulting test statistic. It does not add an independent assessment of evidential strength beyond the estimate, its uncertainty, and the associated test. Prospective power instead evaluates repeated hypothetical studies under parameters specified before the relevant outcome data are analyzed.

Internal pilot studies permit certain nuisance parameters, such as variance, to be re-estimated while a study is in progress. Blinded re-estimation can preserve the original treatment comparison while updating aspects of the calculation that do not depend on the observed between-group effect. Unblinded adaptation requires a design accounting for the effect of interim information on error probabilities.

Sequential and adaptive designs

A fixed-sample design determines a maximum sample size without allowing the primary result to terminate data collection early. In contrast, sequential analysis evaluates accumulating evidence at prespecified points or through a continuous monitoring rule. Repeatedly applying an ordinary fixed-sample significance threshold inflates the probability of a Type I error, whereas group-sequential boundaries allocate the overall error probability across interim analyses.

Abraham Wald’s development of sequential probability ratio tests showed that expected sample size could be reduced under particular null and alternative parameter values while maintaining specified error rates. Group-sequential and adaptive methods extend this principle to settings involving delayed outcomes, multiple treatment groups, or sample-size re-estimation. Their maximum sample sizes can exceed those of corresponding fixed designs even when their expected sample sizes are smaller under selected parameter values.

Interpretation

Sample size is a property of an inferential design rather than a universal measure of study quality. A large sample can estimate a narrowly defined parameter with high precision while retaining systematic error from measurement, selection, or model misspecification. A smaller sample can provide adequate information for a large and stable effect, although its estimates remain less precise for effects outside the range represented by the calculation.

The integer produced by a formula depends on the chosen estimand, analysis model, error criterion, and assumptions about the data-generating process. Changing any of these elements can alter the result without introducing an arithmetic inconsistency. Consequently, distinct sample-size calculations for the same substantive study may correspond to genuinely different statistical questions rather than competing answers to one fully specified question.

See also

  • Statistical power, which describes the probability of rejecting a null hypothesis under a specified alternative.
  • Effect size, which represents the magnitude of a difference or association independently of sample count.
  • Power analysis, which examines the relationship among sample size, effect magnitude, significance level, and test power.
  • Survey sampling, which develops probability-based methods for selecting observations from finite populations.
  • Experimental design, which concerns the allocation of experimental units and the structure of statistical comparisons.
  • Multiple comparisons, which addresses error rates when a study evaluates more than one statistical hypothesis.
  • Bayesian experimental design, which defines sample size and information collection through prior distributions and expected utility.
  • Precision medicine trial, which includes designs whose sample-size properties depend on biomarker-defined treatment effects.