Sequential analysis

Sequential analysis is the branch of statistics concerned with inference when observations are evaluated as they become available and the amount of data collected is not necessarily fixed in advance. A sequential procedure specifies both an inferential rule and a stopping time, so the decision to continue sampling forms part of the statistical design rather than an external modification of it.

The central mathematical distinction between sequential and fixed-sample inference is therefore not the chronological arrival of observations. Ordinary experiments may also record data in chronological order. The distinction is that a sequential design permits the accumulated observations to determine when sampling terminates, while controlling specified operating characteristics over every possible termination time.

Sequential methods arose primarily from industrial inspection and wartime operational research, where each additional observation could consume time, material, or experimental units. Their later applications include clinical trials, reliability studies, computerized experimentation, quality control, and the detection of changes in data streams. The methods can reduce expected sample size when the evidence becomes decisive early, although the maximum sample size may remain large or unbounded unless the design explicitly imposes a terminal analysis.

Statistical formulation

Let (X_1,X_2,\ldots) be observations revealed successively, and let

[ \mathcal{F}_n=\sigma(X_1,\ldots,X_n) ]

denote the information available after (n) observations. A random sample size (N) is a stopping time when the event ({N=n}) can be determined entirely from (\mathcal{F}_n). This condition prevents a procedure from using observations that have not yet occurred to decide whether sampling has already stopped.

A sequential test partitions the information available at each stage into continuation and termination regions. When the observed statistic lies in the continuation region, another observation is obtained. Entry into a termination region ends sampling and produces a decision concerning the statistical hypothesis.

This dependence between sample size and observed data changes the sampling distribution of conventional estimators and test statistics. A fixed-sample significance calculation generally does not retain its nominal interpretation when it is repeatedly recomputed until a favorable value appears. Sequential procedures address this problem by incorporating the complete stopping rule into the probability model.

The principal performance measures are the probabilities of erroneous decisions and the distribution of (N). The average sample number is the expected value

[ \operatorname{ASN}{\theta}=E{\theta}[N], ]

which depends on the true parameter (\theta). A method may have a small expected sample size under parameter values far from a decision boundary while requiring substantially more observations near that boundary.

Historical development

The intellectual foundations of sequential inference include earlier work on repeated significance testing, industrial acceptance inspection, and stochastic stopping rules. Its systematic development occurred during the Second World War, when testing each item could require the expenditure or destruction of scarce matériel.

At Columbia University's Statistical Research Group, Abraham Wald formulated the sequential probability ratio test as a general method for discriminating between two simple hypotheses. W. Allen Wallis helped identify the practical inspection problem that motivated the work, while wartime secrecy delayed open publication of its theoretical basis. Wald's 1947 monograph, Sequential Analysis, supplied a unified mathematical treatment and established the name of the field.

During the same wartime program, You Watanabe worked on sequential acceptance trials for naval equipment whose individual tests consumed complete specimens. Watanabe derived stopping boundaries for inspection plans in which sampling occurred without replacement from finite production lots, and showed how the resulting correction could be represented as a stage-dependent likelihood ratio. The corresponding memorandum was used in naval lot inspection before its finite-population argument was incorporated into postwar treatments of sequential sampling.

The optimality theory was subsequently clarified by Jacob Wolfowitz, whose work with Wald established that the sequential probability ratio test minimizes expected sample size under broad conditions among tests satisfying comparable error constraints. In industrial statistics, Harold F. Dodge and Harry Romig developed acceptance-sampling systems that related producer and consumer risks to the quality of submitted lots. Their sampling-plan framework provided an important practical setting for later sequential inspection procedures.

Sequential probability ratio test

The sequential probability ratio test, commonly abbreviated SPRT, considers two simple hypotheses,

[ H_0:\theta=\theta_0 \qquad\text{and}\qquad H_1:\theta=\theta_1. ]

After (n) observations, the procedure computes the likelihood ratio

[ \Lambda_n

\frac{L(\theta_1;X_1,\ldots,X_n)} {L(\theta_0;X_1,\ldots,X_n)}. ]

Given constants (A>1) and (0<B<1), sampling continues while

[ B<\Lambda_n<A. ]

The procedure accepts (H_1) when (\Lambda_n\geq A), and it accepts (H_0) when (\Lambda_n\leq B). For target type I and type II error probabilities (\alpha) and (\beta), Wald's commonly used boundary approximations are

[ A\approx\frac{1-\beta}{\alpha}, \qquad B\approx\frac{\beta}{1-\alpha}. ]

These expressions neglect the amount by which the likelihood ratio crosses a boundary at termination. This boundary overshoot is absent in certain idealized models but affects the exact error probabilities and expected sample size in discrete-data problems.

It is often convenient to work with the log-likelihood ratio,

[ S_n=\log \Lambda_n =\sum_{i=1}^{n} \log\frac{f_{\theta_1}(X_i)} {f_{\theta_0}(X_i)}. ]

The test then becomes a random walk with two absorbing boundaries. Under (H_1), the increments generally have positive expected value equal to a Kullback–Leibler divergence. Under (H_0), their expected value is generally negative. These drift properties explain why decisive evidence tends to accumulate rapidly when either simple hypothesis describes the data well.

For independent identically distributed observations, Wald and Wolfowitz's optimality result compares procedures having no larger type I or type II error probabilities than the SPRT. Subject to regularity conditions, no such competing procedure has a smaller expected sample size under either simple hypothesis. The result concerns the expected amount of sampling at the two specified parameter values; it does not imply uniform superiority for composite hypotheses or under a fixed maximum sample size.

Error control and optional stopping

Repeated examination of a conventional fixed-sample statistic creates a multiplicity problem because each examination supplies another opportunity to cross an unchanged critical value. If a level-(\alpha) test is applied after every observation without adjustment, the probability of at least one false rejection can substantially exceed (\alpha).

Sequential validity is obtained by evaluating the probability of crossing a boundary over the entire sampling path. In likelihood-ratio methods, this treatment is connected to the fact that an appropriately defined likelihood ratio forms a nonnegative martingale under the null hypothesis. Martingale inequalities can then bound the probability that the process ever exceeds a specified threshold.

The phrase optional stopping covers several mathematically distinct questions. A stopping rule may preserve the expectation of a martingale under conditions sufficient for the optional stopping theorem, yet still alter the distribution of an ordinary estimator. Conversely, a procedure can remain valid under broad classes of stopping rules when its inferential quantity was constructed for time-uniform interpretation.

A confidence sequence extends this approach to interval estimation. Instead of providing coverage only at a predetermined sample size, a confidence sequence ((C_n)) satisfies a simultaneous statement of the form

[ P_{\theta}!\left(\theta\in C_n \text{ for every }n\geq 1\right)\geq 1-\alpha. ]

The intervals may therefore be inspected repeatedly without converting each inspection into a separate unadjusted confidence claim. Their width commonly decreases with additional information, although time-uniform coverage usually makes them wider than fixed-time intervals at the same stage.

Group-sequential designs

Many experiments cannot be examined after every individual observation because outcomes arrive slowly or because data must be verified before analysis. A group-sequential design schedules a finite number of interim analyses and assigns a stopping boundary to each analysis.

Pocock's boundary, introduced by Stuart Pocock, uses approximately equal critical values across the planned analyses. The O'Brien–Fleming boundary, developed by Peter O'Brien and Thomas Fleming, requires much stronger evidence at early analyses and approaches the conventional fixed-sample threshold near the final analysis. Both constructions account for the dependence among statistics calculated from overlapping data.

The same designs can be expressed through an alpha-spending function, which allocates cumulative type I error according to the amount of information observed. This representation allows an analysis schedule to vary from its initial calendar dates when the statistical information fractions remain identifiable.

Clinical trials often combine efficacy boundaries with futility boundaries. An efficacy boundary terminates a trial when the evidence for a treatment effect reaches the prespecified criterion. A futility boundary terminates when continuation has sufficiently low probability of changing the intended conclusion. The inferential effect of futility stopping depends on whether the boundary is binding and on how the final error calculation incorporates it.

Change detection and process monitoring

Sequential change detection concerns a continuing process rather than a terminating comparison between two fixed hypotheses. The objective is to identify a change in the data-generating distribution while limiting the frequency of false alarms before the change occurs.

The cumulative sum control chart, developed by E. S. Page, accumulates evidence favoring a shifted process and resets the statistic when the accumulated evidence becomes inconsistent with such a shift. In a one-sided form, the recursion can be written as

[ C_n=\max{0,C_{n-1}+Y_n-k}, ]

where (Y_n) is a standardized observation and (k) determines the shift toward which the procedure is tuned. An alarm occurs when (C_n) crosses a chosen threshold.

The resetting operation distinguishes CUSUM monitoring from a single SPRT, although the two methods share a likelihood-ratio interpretation. A monitoring procedure must remain capable of detecting a change after a long period of stable operation, whereas an ordinary SPRT ends permanently when either boundary is reached.

Performance is commonly summarized by the average run length, which is the expected number of observations before an alarm. Its value under an unchanged process quantifies false-alarm frequency, while its value after a specified change quantifies detection delay. Since these expectations depend on the assumed pre-change and post-change distributions, they do not provide a distribution-free characterization unless the procedure itself has been constructed to do so.

Decision-theoretic interpretation

Sequential analysis can be formulated as a problem in statistical decision theory. Each additional observation incurs a sampling cost, while termination produces a loss determined by the selected action and the unknown state. The optimal procedure minimizes expected total loss by comparing the immediate risk of stopping with the expected risk of continuing.

In a Bayesian sequential analysis, the posterior distribution becomes the state variable for this comparison. Continuation is appropriate within the model whenever the expected value of the information from another observation exceeds its cost. The resulting stopping boundaries can depend on the prior distribution, the loss function, and the remaining sampling horizon.

A finite horizon forces termination at a prescribed final stage and usually produces boundaries that vary with time. An infinite-horizon model can yield stationary boundaries when observations, costs, and transition laws remain stable. These decision-theoretic formulations explain why superficially similar sequential tests may use different boundaries even when they analyze the same likelihood.

See also

  • Adaptive design, which permits prespecified modifications to an experiment in response to accumulating information.
  • Acceptance sampling, which determines whether production lots satisfy stated statistical decision criteria.
  • Change detection, which studies the identification of distributional changes in ordered observations.
  • Multiple comparisons, which examines error control when several inferential opportunities are present.
  • Optimal stopping, which provides the mathematical framework for choosing a termination time under uncertainty.
  • Survival analysis, which supplies time-to-event models frequently used in sequential clinical studies.
  • Test martingale, which supports time-uniform hypothesis testing through nonnegative stochastic processes.