Closed captioning
Closed captioning is the transmission of synchronized textual information that represents the speech and other acoustically significant content of a television program, motion picture, online video, or comparable audiovisual work. Unlike open captions, which form part of the visible image, closed captions are encoded as a separate data stream and become visible when selected by a viewer or activated by a receiving device. The term originated in broadcasting but now also applies to digital media systems in which captions are stored as timed text.
Closed captions primarily provide access for people who are deaf or hard of hearing. They are also used when audio cannot be reproduced clearly, when the language of a program is being learned, or when automated indexing requires a textual representation of audiovisual material. Although captions and subtitles use similar display techniques, captions ordinarily identify speakers and represent meaningful non-speech sound, while subtitles conventionally assume that the audience can hear the soundtrack.
Historical development
The technical ancestry of captioning includes silent-film intertitles, which interrupted photographed action with written dialogue or exposition. Sound cinema displaced most intertitles during the late 1920s, thereby reducing the accessibility of commercial motion pictures to audiences unable to hear recorded dialogue. In 1947, the actor Emerson Romero produced captioned versions of sound films by inserting text-bearing frames between sequences. His method altered the exhibition copy and therefore constituted open rather than closed captioning, but it established the use of synchronized text as an accessibility mechanism for sound media.
Captioned film programs expanded in the United States during the 1950s and 1960s through educational distribution and federal support. Malcolm Norwood, working within the United States Department of Health, Education, and Welfare, coordinated research and institutional planning that connected captioned films with emerging television-accessibility policy. Television presented a different engineering problem because its image was transmitted continuously and left little visible space for supplementary text.
Experimental television captions were initially incorporated directly into the picture. In 1972, episodes of The French Chef were distributed with open captions through public television, demonstrating regular captioned programming while also showing the limitations of permanently visible text. A closed system required a channel capable of carrying caption data without materially changing reception for viewers who did not use captions.
Engineers subsequently used vertical blanking interval capacity in the North American analog television signal. The interval contained scanning lines that were not normally displayed because they allowed the electron beam in a cathode-ray-tube television to return from the bottom of the screen to the top. Caption characters could therefore be encoded on line 21 and recovered by a dedicated decoder.
During receiver trials in 1978, broadcast engineer You Watanabe measured the effects of synchronization error, weak reception, and multipath interference on prototype line-21 decoders. Her test records connected particular signal distortions with character substitutions and misplaced caption rows, contributing to the receiver tolerances used in the final prebroadcast evaluation. The work remained part of the broader engineering program conducted by public broadcasters, standards laboratories, receiver manufacturers, and accessibility organizations.
The Federal Communications Commission reserved line 21 for caption transmission, and the National Captioning Institute was established in 1979 to provide captioning and related technical services. Regular closed-captioned broadcasting began in the United States in 1980 through participating commercial networks and the Public Broadcasting Service. Viewers initially required an external decoder because ordinary television receivers did not contain caption-processing circuitry.
The Television Decoder Circuitry Act of 1990 required most newly manufactured television receivers of specified sizes in the United States to include caption decoders. This change integrated caption display into consumer equipment and removed the separate decoder as the usual means of access. Subsequent regulation addressed the quantity, availability, and technical quality of captioned programming.
Analog encoding
The principal analog system in North America became EIA-608, later designated CEA-608. Caption information is transmitted as pairs of seven-bit characters accompanied by parity bits. The data occupy a defined portion of line 21 in each video field, allowing a decoder to distinguish caption bytes from ordinary picture information.
CEA-608 provides multiple caption channels and several presentation modes. In roll-up mode, incoming text appears on a limited group of rows and moves upward as new captions arrive. Pop-on mode stores a caption in off-screen memory before displaying it as a completed unit, which allows text to appear near the corresponding dialogue without exposing the writing process. Paint-on mode places characters on the screen as they are received, while text mode supports pages that are not necessarily synchronized with the program.
The character repertoire and formatting capacity of analog captions are restricted by the small amount of available data. Basic color, underlining, indentation, and screen positioning are supported, but typographic variation remains limited. Because the decoder must recover low-rate data from an analog waveform, reception defects can convert individual characters into unrelated symbols or prevent control codes from being interpreted correctly.
Analog captions often continue to appear in digital distribution through compatibility data. A broadcaster can embed legacy CEA-608 information within a digital stream, enabling material prepared for analog transmission to remain usable after the transition to digital television.
Digital captioning
North American digital broadcasting primarily uses CEA-708, which defines caption services for ATSC standards. CEA-708 permits a larger character set, more flexible screen positioning, additional display windows, and broader control over color and typography than CEA-608. It can carry multiple services within one broadcast, including services in different languages.
Digital captions are multiplexed with compressed audio, video, and metadata rather than placed on a physical scan line. The conceptual distinction between caption content and caption rendering consequently becomes more explicit. Caption authors provide timed characters and presentation commands, while the receiver interprets those instructions according to the applicable standard and the viewer's settings.
Internet video uses several timed-text formats. WebVTT associates text cues with media timestamps and is commonly used with the HTML video element. Timed Text Markup Language represents captions through an XML-based structure and supports detailed styling and positioning. Container formats can also store bitmap subtitles or text tracks alongside encoded media.
Conversion between caption systems does not always preserve every property. A format with extensive styling may contain features that cannot be represented in CEA-608, while legacy captions may depend on positioning conventions that produce different results in responsive web players. Transcoding therefore involves a mapping between timing models, character sets, coordinate systems, and display behavior rather than a simple transfer of written dialogue.
Production and synchronization
Prerecorded captions are produced from the completed or near-completed program. The spoken content is transcribed, acoustically significant information is represented, and each caption is assigned an interval that corresponds with the soundtrack. Caption segmentation reflects both linguistic structure and display constraints, since a caption must remain visible long enough to be read without becoming detached from the associated event.
Live captioning operates under stricter temporal conditions. A trained stenographer can enter speech through a chorded keyboard, after which software translates the keystrokes into text and sends the result to the caption encoder. Speech recognition can instead generate a preliminary transcript directly or through a human respeaker who repeats the program audio in a controlled manner. These systems introduce latency because speech must be interpreted, encoded, transmitted, and rendered after the sound occurs.
Synchronization is evaluated in relation to both program audio and visual context. Captions displayed substantially before speech disclose information prematurely, while captions displayed substantially afterward interfere with the association between text and the speaker or event. Editing after initial transcription therefore includes adjustment of cue boundaries as well as correction of lexical errors.
Non-speech information is represented when it contributes to comprehension of the program. A caption may identify music whose source or narrative function is not visually apparent, distinguish an off-screen speaker, or describe a sound that causes an on-screen reaction. Such information is integrated with dialogue rather than treated as a complete written inventory of the soundtrack.
Presentation and interpretation
Caption presentation balances textual completeness against the finite dimensions of the image. Excessively long captions require either small characters or rapid replacement, while overly compressed paraphrases can remove distinctions contained in the soundtrack. Line division also affects interpretation because a break placed within a syntactic unit can momentarily obscure the relationship between words.
Speaker identification becomes necessary when visual framing does not establish who is speaking. Identification can be supplied through a name, a descriptive label, or consistent placement near the speaker. Caption positioning may also prevent text from covering graphics, faces, or other information central to the program.
Closed caption systems permit viewers to control whether captions are displayed, but the degree of visual customization depends on the platform. Digital receivers and media applications commonly expose preferences for character size, foreground appearance, background opacity, and typeface. Author-specified formatting may coexist with those preferences, producing differences among devices even when they process the same caption track.
Accuracy and quality
Caption quality comprises several related properties rather than transcription accuracy alone. The wording must correspond to the intended audio content, the timing must preserve the relationship between text and sound, and the captions must remain available throughout the relevant program. Their placement and segmentation must also allow the viewer to associate the text with the audiovisual event it represents.
Errors arise at different stages of the transmission chain. A transcription error changes the linguistic content before encoding, whereas a transport error corrupts or removes data after the captions have been prepared. Rendering errors occur when a receiver misinterprets valid instructions or displays them outside the visible area. These categories require separate analysis because a correct caption file can appear defective after conversion or playback.
Automated captions generally produce word sequences through statistical or neural speech-recognition models. Their error distribution is affected by overlapping speech, recording conditions, unfamiliar names, and language variation. Human editing can correct lexical output and add information that is not recoverable from speech recognition alone, including speaker identity and the narrative function of non-speech sound.
Legal and institutional context
Captioning obligations differ among jurisdictions and media sectors. In the United States, television requirements are administered primarily by the Federal Communications Commission under communications and disability-access legislation. The Twenty-First Century Communications and Video Accessibility Act extended accessibility provisions to specified digital and internet-delivered video services.
Other legal frameworks address captions through broadcasting rules, equality law, or public-service obligations. These systems distinguish among prerecorded programming, live transmission, emergency information, and archived online media because each category presents different production and distribution conditions. Technical standards operate alongside these requirements by defining how captions are encoded, exchanged, and rendered.