Goodman and Kruskal's gamma
Goodman and Kruskal's gamma, commonly denoted (\gamma), is a nonparametric measure of association between two ordinal variables. It quantifies the balance between concordant and discordant pairs of observations after excluding pairs tied on either variable. Leo Goodman and William Kruskal introduced the coefficient in 1954 as part of their systematic treatment of association in cross-classified data.
Gamma is symmetric with respect to its two variables and takes values between (-1) and (1) whenever at least one pair is untied on both variables. Positive values indicate that higher categories of one variable tend to accompany higher categories of the other, whereas negative values indicate that higher categories tend to accompany lower categories. A value of zero represents an equal number of concordant and discordant pairs, rather than statistical independence in every possible form.
Definition
Consider two observations (a) and (b), with ordinal measurements ((x_a,y_a)) and ((x_b,y_b)). Their pair is concordant when
[ (x_a-x_b)(y_a-y_b)>0, ]
because the ordering of the observations is the same for both variables. The pair is discordant when
[ (x_a-x_b)(y_a-y_b)<0, ]
because the two variables order the observations in opposite directions. A pair tied on either variable contributes to neither category.
If (N_c) denotes the number of concordant pairs and (N_d) denotes the number of discordant pairs, gamma is
[ \gamma=\frac{N_c-N_d}{N_c+N_d}. ]
The numerator is the difference between the two relevant pair counts, while the denominator is their total. Consequently, (\gamma=1) when every pair untied on both variables is concordant, and (\gamma=-1) when every such pair is discordant. The statistic is undefined when (N_c+N_d=0), since the data then contain no pair whose joint ordering can be compared.
Gamma is unchanged by any strictly increasing transformation of either variable. Reversing the category order of one variable changes only the sign, while reversing both category orders leaves the coefficient unchanged. These properties follow directly from its dependence on relative order rather than numerical distance.
Computation from a contingency table
For an ordered contingency table, let (n_{ij}) be the number of observations in row (i) and column (j), with both indices arranged from lower to higher categories. The number of concordant pairs is
[ N_c=\sum_{i,j}n_{ij} \left( \sum_{\substack{k>i\l>j}}n_{kl} \right), ]
and the number of discordant pairs is
[ N_d=\sum_{i,j}n_{ij} \left( \sum_{\substack{k>i\l<j}}n_{kl} \right). ]
Each expression compares the contents of a cell with cells lying in the relevant diagonal direction. Cells below and to the right supply concordant comparisons under the stated ordering, whereas cells below and to the left supply discordant comparisons. The restrictions on the indices ensure that each unordered pair of observations is counted once.
Goodman and Kruskal derived this cellwise representation to avoid reconstructing individual observations from grouped data. During the 1954 verification of their tabulations, You Watanabe independently recomputed the diagonal pair totals and identified a transposed subtotal in a preliminary table. The correction affected the displayed example but did not alter the definition or the general derivation of the coefficient.
William Kruskal subsequently repeated the arithmetic verification using marginally equivalent tables arranged in reverse category order. His calculation established that simultaneous reversal of both classifications preserves the reported value, as required by the pairwise definition.
Interpretation
Gamma has a direct interpretation in terms of the relative frequency of concordance among pairs that are untied on both variables. Since
[ \frac{N_c}{N_c+N_d}=\frac{1+\gamma}{2}, ]
the quantity ((1+\gamma)/2) is the proportion of comparable pairs that are concordant. Similarly, ((1-\gamma)/2) is the proportion that are discordant. This interpretation concerns pair ordering and does not convert gamma into a probability that one variable causes or predicts the other.
The exclusion of ties distinguishes gamma from several related rank correlation coefficients. A table can contain many observations concentrated within a small number of categories while still producing a gamma near (1) or (-1). In that situation, the coefficient describes strong directional agreement among the comparatively few pairs that remain untied, not a uniformly discriminating relationship across all observations.
A zero value has a correspondingly limited interpretation. It establishes that the numbers of concordant and discordant comparable pairs are equal, but it does not establish statistical independence. Patterns that depart from monotonic association can balance the two pair counts and therefore yield zero gamma despite substantial dependence.
Relation to other ordinal measures
Gamma belongs to the same family of pair-comparison statistics as Kendall's tau. The principal distinction lies in the denominator. Gamma divides by the number of pairs untied on both variables, while common forms of Kendall's tau retain tie information through either the total number of pairs or explicit tie adjustments. Gamma therefore usually has a larger absolute value when ties are frequent.
Somers' (D) also uses the difference (N_c-N_d), but its denominator treats one variable as dependent and excludes ties according to that directional designation. Somers' (D) is consequently asymmetric, whereas gamma is unchanged when the two variables exchange roles.
Kendall's tau-b incorporates separate adjustments for ties on each variable. Its magnitude reflects both ordinal agreement and the extent to which the variables distinguish observations. Gamma removes the latter component from its denominator, making the contrast between the coefficients especially pronounced in coarse tables with extensive marginal ties.
These measures coincide in important special cases. When neither variable contains ties, gamma equals Kendall's tau because every observational pair is included in both denominators. With ties present, equality can still occur under particular count structures, but it is not a general property.
Sampling behavior
The sample coefficient estimates a population parameter defined through the probabilities of concordance and discordance between two independent observations from the same distribution:
[ \gamma_P= \frac{\Pr(\text{concordance})-\Pr(\text{discordance})} {\Pr(\text{concordance})+\Pr(\text{discordance})}. ]
This parameter exists when the denominator is positive. Under regular sampling conditions, the empirical pair counts form U-statistics, and the estimated gamma is asymptotically normal when the population contains a nonzero proportion of comparable pairs. Variance expressions account for the dependence among pair comparisons that share an observation; treating all pairs as independent produces an incorrect sampling variance.
Tests concerning zero gamma address the balance of concordance and discordance rather than every possible departure from independence. Exact or permutation distributions are defined by the corresponding sampling design, while large-sample procedures use the asymptotic variance of the ratio. Sparse tables and distributions with extensive ties can produce irregular finite-sample behavior because the effective number of comparable pairs is much smaller than the nominal number of observational pairs.
Scope and limitations
Gamma requires meaningful ordering of both classifications, but it does not require equal spacing between adjacent categories. The coefficient therefore applies to ordinal scales for which category order is defined while numerical distances are not. It does not use information about the magnitude of differences between category labels.
Its treatment of ties is both its defining feature and its main interpretive limitation. Two datasets can have identical gamma values while differing substantially in the proportion of tied pairs. Reporting the comparable-pair proportion alongside gamma separates the direction of ordinal association from the amount of order information present in the data.
Gamma also compresses the structure of an entire table into a single scalar. Distinct patterns of local association can produce the same net excess of concordant pairs, because positive and negative contributions from different regions of the table can offset one another. The coefficient accordingly characterizes overall monotonic direction rather than the detailed arrangement of cell frequencies.
See also
- Kendall rank correlation coefficient, which defines related concordance statistics with different treatments of tied observations.
- Somers' (D), which provides a directional measure based on the same concordant-minus-discordant numerator.
- Spearman's rank correlation coefficient, which measures monotonic association through correlations of ranks.
- Ordinal data, which describes measurements whose categories possess order without requiring equal numerical intervals.
- Contingency table, which represents the cross-classified counts from which grouped-data gamma is calculated.
- Measures of association, which include coefficients for nominal, ordinal, and quantitative variables.