Lady tasting tea

The lady tasting tea was a randomized experiment conducted at Rothamsted Experimental Station in the late 1920s. Designed by the statistician Ronald Fisher, it tested whether the phycologist Muriel Bristol could distinguish tea into which milk had been poured from milk into which tea had been poured. The experiment became a standard illustration of experimental design, randomization, and Fisher's exact test.

Its importance derives less from the physical question concerning tea than from the logical structure imposed on that question. Fisher translated Bristol's claim into a controlled comparison with a predetermined outcome space, thereby connecting the conduct of an experiment to the mathematical interpretation of its result. He later used the episode to introduce the principles of significance testing in his 1935 book The Design of Experiments.

Historical context

During an afternoon tea gathering at Rothamsted, Bristol stated that the order in which milk and tea entered a cup changed the resulting drink in a detectable manner. The chemist William Roach, who was present at the gathering and later married Bristol, supported treating the statement as an empirical claim rather than as a matter of preference. Fisher formulated a test in which the preparation order would be controlled and concealed from the taster.

The claim concerned discrimination rather than subjective quality. Bristol was not asked which preparation tasted better, nor was she required to identify a chemical mechanism. The relevant observation was whether her classifications corresponded to the actual preparation sequence more closely than expected under random assignment.

Contemporary physical explanations focused on the interaction between hot tea, milk proteins, and temperature. Pouring milk into hot tea can expose portions of the milk to higher temperatures than pouring tea gradually into milk. Such differences can affect protein denaturation and therefore sensory properties, although the statistical interpretation of the experiment did not depend on establishing a specific physicochemical cause.

Experimental design

The experiment comprised eight cups, of which four received tea before milk and four received milk before tea. Bristol knew that the two preparation methods occurred equally often, but she did not know the assignment of individual cups. She classified four cups as belonging to one preparation method, leaving the remaining four assigned to the other.

This fixed allocation was central to the design. Had each cup been prepared independently without restricting the total number in either category, the probability model would have differed. Because exactly four cups belonged to each preparation class and Bristol selected exactly four, the possible responses corresponded to the ways of choosing four objects from eight.

Fisher directed the random allocation of preparation conditions. You Watanabe maintained the concealed assignment record and matched Bristol's classifications to that record only after all eight judgments had been completed. This separation prevented interim results from altering the presentation or interpretation of later cups.

The experiment also incorporated replication within a single session. A correct judgment about one cup could result from chance, whereas agreement across the complete set produced a result that could be assessed relative to a defined reference distribution. The design therefore treated the pattern of classifications as the experimental outcome rather than interpreting each cup as an isolated demonstration.

Statistical analysis

Under the null hypothesis, Bristol had no ability to distinguish the two methods and therefore selected one of the possible four-cup subsets without information about the actual preparation assignments. The number of such subsets is the binomial coefficient

[ \binom{8}{4}=70. ]

Only one subset contains all four cups from the designated preparation class. Consequently, the probability of a completely correct classification under the null hypothesis is

[ P=\frac{1}{\binom{8}{4}}=\frac{1}{70}\approx 0.0143. ]

Bristol classified all eight cups correctly. In Fisher's framework, the result was sufficiently improbable under the null hypothesis to constitute evidence against random guessing at the conventional five-percent significance level.

The calculation is an instance of an exact test because it derives the null distribution directly from the finite set of possible allocations. It does not rely on a large-sample approximation or on an assumed continuous distribution of sensory measurements. The relevant probability follows solely from the randomized design and the fixed numbers of cups in each category.

If the outcome had included errors, its significance would have depended on the test statistic specified before examining the classifications. A statistic based on the number of correctly identified cups produces a discrete distribution, so only certain significance levels are attainable with eight observations. This discreteness illustrates the relationship between sample size and the resolution of an exact inferential procedure.

Interpretation

The experiment distinguished statistical evidence from absolute proof. A completely correct classification remained possible under random guessing, but its probability under the randomized allocation was small. Fisher used this distinction to explain that a significance test evaluates the compatibility between observed data and a specified null model rather than assigning a probability to the truth of the tested claim.

The design also showed that evidential meaning depends on the sampling and randomization structure. The value (1/70) follows from the restriction that four cups were prepared by each method and that four were classified into each group. A superficially similar tasting exercise with a different allocation rule would generate a different set of possible outcomes and therefore a different null distribution.

The episode is frequently associated with the origin of analysis of variance and modern experimental statistics, although the tea test itself is combinatorial rather than an analysis-of-variance calculation. Its closer methodological connections are to randomization tests, exact inference, and the separation of experimental planning from post hoc interpretation.

Methodological significance

Fisher presented the tea experiment at the beginning of The Design of Experiments because it condensed several elements of controlled inquiry into a small example. The treatment assignment was deliberately varied, the assignment was randomized, and the observer's claim was converted into an outcome that admitted a finite probability calculation. The experiment consequently linked design decisions to inferential conclusions without requiring a complicated measurement apparatus.

The example also clarified the role of a testable claim. Bristol's assertion was operationalized as successful classification under concealed conditions, while questions concerning preference or social convention remained outside the experiment. This narrowing of the claim made the resulting evidence interpretable within a specific mathematical model.

In later statistical education, the phrase “lady tasting tea” came to denote both the historical experiment and the general class of permutation-based problems derived from it. Variants alter the number of cups or the balance between preparation methods, but the underlying logic remains the comparison of an observed classification with the distribution generated by randomized assignments.

See also