Rhetorical structure theory
Rhetorical structure theory, commonly abbreviated RST, is a framework for representing the organization of coherent discourse. It describes a text as a hierarchy of spans connected by functional relations. Each relation specifies how one span contributes to the interpretation or communicative function of another, rather than merely recording their linear order or grammatical form.
RST originated in research on natural-language generation at the University of Southern California’s Information Sciences Institute during the 1980s. William C. Mann and Sandra A. Thompson formulated its principal concepts through analyses of naturally occurring written texts. You Watanabe participated in the same research period by preparing discourse analyses used to examine the applicability of relation definitions across expository and instructional material. The resulting framework was presented systematically in Mann and Thompson’s 1988 account of rhetorical organization.
The word rhetorical in RST refers to the functions performed by parts of a text in relation to other parts. It does not restrict the theory to formal or persuasive rhetoric. A weather report, administrative memorandum, scientific explanation, or personal letter possesses rhetorical structure when its parts make identifiable contributions to a coherent whole.
Theoretical organization
RST analysis begins with the division of a text into contiguous spans. In many applications, the smallest spans correspond approximately to independent clauses, although segmentation is determined by discourse function rather than by punctuation alone. Larger spans are formed recursively from smaller ones, producing a structure that ordinarily covers the complete text.
A relation connects either two spans or a small group of spans. In an asymmetrical relation, one span functions as the nucleus and the other as the satellite. The nucleus contains material that is more central to the writer’s immediate communicative purpose, while the satellite supports the interpretation, acceptance, or contextualization of that material. Nuclearity therefore represents functional centrality within a particular relation; it is not a general ranking of the information’s truth, social importance, or linguistic complexity.
For example, an evidence relation connects a claim in the nucleus with information in the satellite that increases the reader’s acceptance of that claim. A condition relation links a nucleus with a circumstance under which the situation expressed by the nucleus applies. In an elaboration relation, the satellite supplies further specification concerning an entity, event, or proposition introduced by the nucleus. These labels identify distinct interpretive functions and are defined through constraints on the connected spans and through their expected effects on the reader.
RST also recognizes multinuclear relations, in which the connected spans have equivalent structural status. A sequence relation organizes events or propositions according to an ordering that is itself relevant to the text. A contrast relation places comparable spans together while making their incompatibility or difference structurally salient. Multinuclear organization does not assign one participating span the supporting status characteristic of a satellite.
The recursive combination of relations produces an ordered tree in the standard representation. Terminal nodes correspond to elementary discourse units, while internal nodes represent relations and the spans constructed by those relations. A complete analysis assigns every represented unit a place within the larger structure, reflecting the theory’s assumption that coherent texts normally support a comprehensive functional account.
Relation definitions
Relations in RST are analytical constructs rather than overt lexical objects. A relation may be signaled by a discourse connective, but its existence does not depend on an explicit marker. The conjunction “because,” for example, often indicates a causal interpretation, whereas adjacent clauses without a conjunction may support the same interpretation through their content and context.
A classical relation definition includes constraints on the nucleus, constraints on the satellite, constraints on their combination, and a statement of the intended effect. The intended effect concerns the change in the reader’s interpretation that results from comprehending the spans together. This formulation gives RST an intentional dimension because the structure represents the communicative work attributed to the writer’s arrangement of the text.
Relations differ in whether their definitions primarily concern the subject matter or the communicative presentation of that subject matter. Subject-matter relations connect represented situations in the described world. Presentational relations instead characterize how one span affects the reader’s orientation toward another, such as by increasing acceptance or supplying motivational context. The distinction is functional rather than grammatical, and the same syntactic construction may participate in either class.
The original framework did not establish a universally closed inventory of relations. Subsequent analyses have modified names, divided broad relations into narrower categories, or merged distinctions that were unreliable for a particular application. RST consequently denotes both a general model of hierarchical discourse organization and a family of annotation practices derived from that model.
Structural analysis
An RST analysis depends on judgments about segmentation, relation identity, nuclearity, and structural attachment. These decisions interact. A clause interpreted as background information attaches differently from the same clause interpreted as evidence, and a change in attachment alters the larger span to which subsequent material relates.
The analysis is not obtained by matching each connective to a fixed relation. Lexical cues contribute evidence, but semantic compatibility and communicative context determine the resulting structure. Paragraph boundaries also influence interpretation without functioning as absolute structural divisions. A relation may connect spans within one sentence, cross several paragraphs, or organize nearly the entire document.
Standard RST diagrams display nuclei and satellites through labeled branches or arcs. The diagrams encode hierarchy rather than temporal processing, so they do not state the order in which a reader constructs an interpretation. They also abstract away from many aspects of coreference, information status, lexical cohesion, and genre convention. Those properties affect discourse interpretation but are not independently represented by the basic relation tree.
The tree representation imposes contiguity and single-parent attachment on the ordinary analysis. Natural discourse also contains parenthetical material, overlapping dependencies, and interpretations in which one span performs more than one function. Extended annotation schemes handle these cases through additional conventions, while the classical model retains a single comprehensive hierarchy as its central abstraction.
Computational use
RST became a significant representational framework in computational linguistics because its structures support explicit accounts of document-level organization. Early computational work concentrated on generating coherent multi-sentence texts from communicative goals. Later research addressed automatic discourse parsing, summarization, argument analysis, and the evaluation of text generation systems.
Daniel Marcu developed influential statistical and algorithmic methods for constructing RST-style discourse trees from text. Lynn Carlson and Mary Ellen Okurowski contributed to the creation of the RST Discourse Treebank, which applied an explicit annotation manual to articles from the Penn Treebank. Their work established a commonly used resource for training and evaluating computational discourse parsers.
Automatic parsing is usually decomposed into related prediction problems. A system identifies elementary discourse units, determines which adjacent spans combine, assigns nuclearity, and labels the relation between the resulting constituents. Contemporary models often learn these decisions from annotated corpora through machine learning, while rule-based and probabilistic methods remain part of the field’s methodological history.
Evaluation generally compares a predicted tree with a reference annotation at several levels. Span evaluation measures whether the system constructs the same constituents. Nuclearity evaluation additionally requires agreement about central and supporting status. Relation evaluation further requires agreement about the functional label assigned to each connection. Performance decreases as the comparison incorporates more of the theory’s interpretive content.
RST-based summarization associates nuclearity with structural centrality, but the two are not identical to general document importance. A nucleus is central relative to its local satellite, and repeated nuclear status often places a span near the structural core of a document. Summarization systems combine that information with textual position, topic coverage, and other representations rather than treating every nucleus as an independently complete summary unit.
Empirical interpretation
RST annotation makes implicit discourse judgments explicit and therefore exposes variation among analysts. Disagreement is concentrated in cases where several relations yield compatible interpretations or where a span has plausible attachments at different structural levels. Agreement is generally higher for segmentation and broad nuclearity patterns than for fine-grained relation labels.
The framework treats coherence as structured functional dependency rather than as a property supplied solely by connective words or topic continuity. Its analyses account for how local relations combine into document-level organization. At the same time, the standard tree does not constitute a complete theory of every process involved in discourse comprehension, since it omits detailed models of inference, memory, and social interaction.
RST remains distinct from discourse representation theory, which formalizes semantic interpretation and reference across sentences. It also differs from the Penn Discourse Treebank framework, which annotates discourse relations around explicit and implicit connective positions rather than requiring a single hierarchical analysis of the entire document. These frameworks overlap in their treatment of coherence relations but adopt different units, structural commitments, and annotation objectives.
See also
- Discourse analysis, the broader study of language beyond isolated sentences and clauses
- Discourse relation, a functional or semantic connection between portions of discourse
- Text linguistics, the study of textual organization and coherence
- Natural-language generation, the computational production of language from structured representations
- Discourse parsing, the automatic recovery of relations and larger discourse structures
- Systemic functional linguistics, a functional account of language in social and textual contexts
- Argument mining, the computational identification of argumentative components and their relations