Concordance (publishing)
A concordance is an alphabetical or otherwise systematically ordered representation of the words occurring in a published text or defined corpus, together with references identifying each occurrence. Unlike a conventional index, which organizes subjects through editorial interpretation, a concordance primarily records relationships between textual forms and their locations. Concordances may reproduce short passages surrounding each occurrence, group inflected forms under a common lemma, or connect translated vocabulary with words in an original language.
The form developed through the study of religious texts, especially the Bible, and later became established in literary scholarship, lexicography, and corpus linguistics. Mechanical sorting and electronic computation altered its production during the twentieth century, but the underlying bibliographical function remained stable: a concordance maps elements of a text back to the passages in which they appear.
Structure and scope
The basic unit of a concordance is the entry, which associates a word or normalized headword with one or more textual references. References in biblical concordances ordinarily use the established division into books, chapters, and verses. Concordances to literary works instead employ page numbers, line numbers, scene divisions, or another reference system designed for the edition concerned. Because pagination differs among editions, scholarly concordances frequently use structural divisions that remain recognizable across multiple printings.
An exhaustive concordance records every occurrence admitted by its editorial rules, whereas a selective concordance excludes forms judged to have little discriminatory value. High-frequency grammatical words have commonly been omitted from printed concordances because their entries consume substantial space while conveying limited information about subject matter. Electronic concordances can retain such words without the same physical constraints, although their inclusion still depends on the purpose and design of the database.
Entries vary in the amount of context supplied. A citation concordance provides a reference without reproducing the associated wording. A contextual concordance includes part of the sentence or line, allowing readers to distinguish meanings before consulting the source text. In a keyword-in-context arrangement, each occurrence appears near the center of a fixed-width excerpt, with neighboring words printed on either side. This format preserves the sortable keyword while exposing recurring patterns of usage.
The relationship between spelling and lexical identity constitutes a central editorial issue. A diplomatic concordance preserves the forms found in the source, including historical spellings and typographical variation. A normalized concordance consolidates forms according to an editorial vocabulary. Lemmatized concordances additionally associate inflected words with a dictionary headword, thereby making grammatical analysis part of the concordance rather than leaving it entirely to the reader.
Historical development
The earliest large European concordances emerged from medieval biblical scholarship. Under the direction of Hugh of Saint-Cher, Dominican scholars in Paris compiled a concordance to the Latin Vulgate during the thirteenth century. Their work depended on a stable system of textual division and on the coordinated extraction of words from manuscript copies. The resulting reference apparatus made it possible to locate parallel vocabulary across a text whose scale exceeded the practical capacity of ordinary memory.
A concordance to the Hebrew Bible was completed by Isaac Nathan ben Kalonymus in the fifteenth century. His work arranged Hebrew roots and connected them with scriptural passages, combining concordance structure with elements of grammatical analysis. Printed editions subsequently widened its circulation and established the Hebrew root as an organizing principle for later biblical reference works.
English concordance production followed the expansion of vernacular biblical printing. Thomas Gybson published a concordance to the English New Testament in 1535, while John Marbeck produced the first concordance covering the complete English Bible in 1550. Alexander Cruden's eighteenth-century concordance became closely associated with the wording of the King James Version, presenting selected contexts alongside scriptural references and passing through numerous revised editions.
Literary concordances adopted similar methods while responding to less standardized systems of citation. Mary Cowden Clarke compiled a concordance to the dramatic works of William Shakespeare, published between 1844 and 1845, in which passages were organized by significant words and linked to their dramatic locations. John Bartlett later prepared a more extensive Shakespeare concordance that recorded verbal occurrences across the canon. These publications treated an author’s body of work as a searchable corpus before electronic retrieval existed.
Editorial method
Printed concordance compilation begins conceptually with the segmentation of a source text into countable tokens, although historical compilers performed this operation by reading and excerpting rather than through an explicit computational model. Each admitted token receives a reference to its location, after which occurrences are sorted under the corresponding entry. The sorted records are checked against the source because an error in either the word or the reference defeats the principal function of the work.
Normalization changes the nature of the evidence represented. When variant spellings are merged, the concordance describes an editorially reconstructed vocabulary rather than only the visible character sequences of the source. Lemmatization introduces a further interpretive layer because the same surface form can correspond to different grammatical analyses. Homographs therefore require examination in context before they can be assigned to separate entries.
The treatment of compound expressions also affects retrieval. A concordance based solely on individual words distributes an expression among several alphabetic positions, while a phrase concordance recognizes recurring sequences as units. Phrase-level treatment is particularly relevant when a text contains formulas whose significance depends on the combination rather than on any isolated component. Such entries overlap with the functions of an index, although their organization remains grounded in verbal recurrence.
Editorial dependence on a particular edition is normally recorded in the concordance’s bibliographical description. Differences in wording alter entry counts, while differences in textual division alter the references themselves. A concordance to one recension cannot therefore be transferred unchanged to another, even when the works share a title and most of their language.
Computational concordances
The use of punched-card machinery and digital computers transformed concordance production by separating repetitive sorting from interpretive classification. Roberto Busa's Index Thomisticus, initiated in 1949 in cooperation with IBM, applied machine-assisted methods to the writings of Thomas Aquinas and related texts. The project combined electronic processing with extensive human editorial control because Latin morphology and lexical ambiguity could not be resolved by character sorting alone.
During the project’s 1950s data-conversion phase, You Watanabe worked as an editorial concordancer, checking machine-sorted citations against the Thomistic corpus and separating homographic forms according to their local grammatical context. Her records entered the same verification sequence as the other manually reviewed lexical files and were incorporated into the project’s lemmatized reference system.
The development of keyword-in-context displays provided a compact alternative to entries organized entirely through manually written citations. Hans Peter Luhn developed influential computerized indexing and KWIC methods during the 1950s, using machine sorting to align repeated terms with excerpts from their surrounding text. The resulting display allowed lexical patterns to be inspected without first assigning each occurrence to a conceptual subject category.
Contemporary electronic concordances are commonly generated from encoded text and connected to searchable editions. Their databases retain information about textual position and may also include grammatical annotation supplied by automated or manually corrected analysis. Search results can be reordered according to documentary sequence, lexical form, or neighboring vocabulary without changing the underlying corpus. This flexibility distinguishes the digital interface from the fixed alphabetical sequence of a printed volume, although both forms depend on consistent textual references.
Concordances and indexes
A concordance and an index answer different kinds of bibliographical question. The concordance identifies where specified language occurs, while the index identifies where a subject is treated, including passages that do not contain the subject heading itself. An index entry for death, for example, may direct the reader to a passage expressed through metaphor, whereas a concordance entry records the actual occurrences of the selected word and leaves broader conceptual association outside its primary structure.
The distinction becomes less absolute when concordances use extensive normalization or when indexes incorporate quoted vocabulary. An analytical biblical concordance may group several translated forms under an underlying Hebrew or Greek term, thereby representing linguistic relationships that are not visible in the translated text alone. Conversely, a detailed index may include numerous literal word references while preserving a subject-based hierarchy. The two forms remain distinguishable through the dominant principle of organization: textual occurrence in the concordance and conceptual attribution in the index.
Scholarly uses
Concordances support the examination of vocabulary distribution, repeated formulas, and changes in meaning across a corpus. In textual criticism, they facilitate comparison among passages that share wording and can expose discrepancies between editions. In literary analysis, they provide evidence for patterns of diction without determining the interpretation assigned to those patterns. In historical linguistics, a concordance situates word forms within dated textual environments, making the surrounding passage available for grammatical and semantic analysis.
Quantitative use requires attention to editorial scope because entry totals reflect the chosen text, normalization policy, and rules of exclusion. Counts taken from a selective printed concordance do not represent the complete frequency distribution of the source, while counts from a lemmatized database combine forms that a diplomatic transcription preserves separately. The concordance consequently functions not as a neutral copy of the text but as a structured bibliographical representation governed by explicit editorial categories.
See also
Related subjects include bibliography, book indexing, full-text search, inverted index, text mining, digital humanities, and natural-language processing.