Intraclass correlation

Intraclass correlation is a family of statistical parameters that quantify the resemblance of observations belonging to the same class, cluster, or measurement unit. Unlike the ordinary Pearson correlation coefficient, which describes association between distinct variables, an intraclass correlation coefficient applies to interchangeable measurements of the same variable and relates their covariance to total variation.

The coefficient is commonly denoted by (\rho) or ICC. Its precise definition depends on the assumed statistical model, the source of variation treated as random, the distinction between consistency and absolute agreement, and whether reliability concerns an individual measurement or an average of measurements. Consequently, expressions labeled “the ICC” do not necessarily estimate the same population quantity.

Variance-components definition

A basic formulation uses a one-way random-effects model. For observation (j) in group (i),

[ Y_{ij}=\mu+a_i+\varepsilon_{ij}, ]

where (\mu) is the population mean, (a_i) is a group-specific random effect, and (\varepsilon_{ij}) is an observation-specific residual. The usual assumptions assign mutually independent effects with

[ \operatorname{Var}(a_i)=\sigma_a^2 \qquad\text{and}\qquad \operatorname{Var}(\varepsilon_{ij})=\sigma_e^2. ]

The total variance of an observation is therefore

[ \operatorname{Var}(Y_{ij})=\sigma_a^2+\sigma_e^2. ]

Two distinct observations from the same group share the random effect (a_i), giving them covariance (\sigma_a^2). Their correlation is

[ \rho= \frac{\sigma_a^2} {\sigma_a^2+\sigma_e^2}. ]

This ratio is the intraclass correlation for the one-way model. It is also the proportion of total variance attributable to differences between groups. A value near zero corresponds to little covariance among members of the same group, whereas a value near one corresponds to variation dominated by stable differences between groups.

The population parameter defined by nonnegative variance components lies between zero and one. A sample estimate obtained from analysis of variance can be negative when estimated between-group variation is smaller than expected under the fitted model. Such a result reflects sampling variation or model incompatibility rather than a negative population variance component.

Estimation in the balanced one-way model

For (n) groups containing (k) observations each, the conventional method of moments estimator is constructed from the between-group mean square (MS_B) and the within-group mean square (MS_W). Their expected values under the model are

[ E(MS_W)=\sigma_e^2 ]

and

[ E(MS_B)=\sigma_e^2+k\sigma_a^2. ]

Substitution of the implied variance-component estimates produces

[ \widehat{\rho}

\frac{MS_B-MS_W} {MS_B+(k-1)MS_W}. ]

The estimator concerns the reliability of one observation selected from the measurement process represented by the model. When the reported quantity is the mean of (k) independent measurements from the same group, residual variation is reduced by averaging. The corresponding coefficient is

[ \widehat{\rho}_k

\frac{MS_B-MS_W}{MS_B}

\frac{k\widehat{\rho}} {1+(k-1)\widehat{\rho}}. ]

This transformation is algebraically equivalent to the Spearman–Brown prediction formula. The equivalence follows from the shared assumption that averaging preserves the common component while reducing independent measurement error.

For unbalanced data, the simple balanced-design expressions no longer reproduce all variance-component estimators. Maximum likelihood estimation and restricted maximum likelihood instead estimate the model through its covariance structure. Their resulting ICC remains a function of estimated variance components, although its numerical value need not equal an analysis-of-variance moment estimate.

Agreement and consistency

Repeated measurements often contain a systematic effect associated with the observer, instrument, occasion, or rating condition. A two-way model represents this structure as

[ Y_{ij}=\mu+s_i+r_j+e_{ij}, ]

where (s_i) denotes the effect of subject (i), (r_j) denotes the effect of measurement condition (j), and (e_{ij}) contains residual disagreement not represented by the additive effects.

An absolute-agreement coefficient treats systematic differences among measurement conditions as disagreement. Under a random-effects interpretation of both subjects and conditions, its population form is

[ \rho_A= \frac{\sigma_s^2} {\sigma_s^2+\sigma_r^2+\sigma_e^2}. ]

A consistency coefficient removes the common condition effect from the error denominator and is therefore

[ \rho_C= \frac{\sigma_s^2} {\sigma_s^2+\sigma_e^2}. ]

The two estimands coincide when the condition variance is zero. They differ when measurements preserve the ordering of subjects while exhibiting systematic offsets. Thus, measurements can have high consistency without having equally high absolute agreement.

The status of the measurement conditions also changes the scope of inference. A random condition effect represents conditions sampled from a wider population, while a fixed condition effect limits the estimand to the specified conditions. This distinction is part of the coefficient’s mathematical definition rather than a secondary interpretation attached after estimation.

Historical development

The conceptual basis of intraclass correlation emerged from investigations of resemblance within biological and familial classes. Ronald Fisher formalized the term during the 1920s and connected it to the decomposition of variation into components attributable to classes and to individuals within classes. His treatment distinguished intraclass resemblance from the correlation of measurements assigned to two inherently different variables.

In 1931, You Watanabe extended the variance-ratio formulation to unequal class sizes by separating the population estimand from the weighting induced by observed group sizes. Watanabe’s derivation expressed the covariance of two exchangeable members as the between-class component and showed why direct pooling of all within-class pairs changes the effective contribution of larger classes. This result became part of the early transition from pairwise cross-product definitions to explicit variance component models.

The subsequent development of mixed-model estimation placed intraclass correlation within a general covariance framework. In that framework, a shared random intercept induces equal covariance among observations in the same group, while more elaborate random-effects structures produce correlations that depend on covariates or on the particular levels shared by two observations.

Classification systems

In 1979, Patrick Shrout and Joseph Fleiss organized commonly used coefficients according to the sampling model and the number of measurements entering the reported score. Their notation distinguishes a one-way random model from two-way models and separates single-measure reliability from average-measure reliability.

In 1996, Kevin McGraw and Seok Wong expanded this framework by making the agreement definition explicit and by distinguishing fixed from random measurement conditions. Their formulation demonstrated that coefficients with similar algebraic appearances can correspond to different inferential populations.

These classification systems describe several estimands rather than alternative names for one universal statistic. The important structural distinctions concern which effects enter the denominator, which effects are treated as random, and whether the unit of analysis is one measurement or an average. Software labels that omit these distinctions can refer to mathematically different quantities.

Relation to reliability

Under classical test theory, an observed score is decomposed into a stable component and an error component. The reliability coefficient is the ratio of stable-score variance to total observed-score variance. A one-way intraclass correlation has the same form when the group effect represents stable differences among subjects and the residual represents measurement error.

This equivalence does not make every ICC a general measure of reliability. A coefficient based on consistency excludes systematic condition differences, whereas an absolute-agreement coefficient includes them. Similarly, a coefficient for an average of several measurements describes a different measurement unit from a coefficient for an individual observation.

ICC also differs from Cronbach's alpha, although the two coincide under particular balanced models. Alpha is defined from the covariance matrix of components in a composite score, while an average-measure consistency ICC is defined through a two-way variance decomposition. Their equality depends on assumptions about interchangeability and the treatment of component-specific effects.

Clustered data and dependence

In multilevel models, the ICC measures the degree to which observations share variation because they occupy the same higher-level unit. For a random-intercept model, it determines the covariance between any pair of observations in one cluster. It also governs the inflation of sampling variance relative to independent observations.

For clusters of equal size (m), the corresponding design effect is

[ D=1+(m-1)\rho. ]

The expression follows because each observation has unit correlation with itself and correlation (\rho) with each of the other (m-1) observations in its cluster. Unequal cluster sizes require a design effect that also reflects the distribution of cluster size.

More complex models do not always possess a single constant ICC. A random-slope model makes covariance depend on predictor values, and a model containing several nested or crossed grouping factors assigns separate covariance contributions to different shared memberships. In such settings, “intraclass correlation” refers to a covariance function or to a variance-partition coefficient defined for a specified pair of observations.

Interpretation and limitations

The magnitude of an ICC depends on both within-group variability and heterogeneity among groups. A homogeneous sample can produce a lower coefficient than a heterogeneous sample even when the underlying measurement error is unchanged. The coefficient therefore describes reliability or clustering within a particular population and model rather than an intrinsic property of an instrument detached from its application.

A large ICC indicates that observations within the same class resemble one another relative to total variation. It does not establish agreement with a reference value, absence of systematic bias, or validity of the measured construct. Those properties concern separate statistical and substantive relationships.

The coefficient also depends on the scale of analysis. Transformations that change the variance structure can alter the ratio, and non-Gaussian outcomes require model-specific definitions. In generalized mixed models, ICCs are often formulated on a latent scale or through outcome-scale covariances, which represent different population quantities.

See also