Video
Video is an electronic medium for representing, recording, transmitting, and reconstructing sequences of visual images. Most video systems create the perception of motion by displaying successive frames at intervals shorter than those at which human vision ordinarily distinguishes separate images. Unlike photographic film, which stores images as physical variations in a light-sensitive material, video originated as a continuously varying electrical signal and later became predominantly digital data.
A video system combines spatial sampling, temporal sampling, and a method of encoding brightness or color. Its engineering characteristics include spatial resolution, frame rate, aspect ratio, color representation, and transmission bandwidth. These properties are interdependent: increasing the number of pixels or frames ordinarily increases the amount of information that must be stored or transmitted unless offset by video compression.
Terminology
The word derives from the Latin verb vidēre, meaning “to see.” In modern technical usage, video may refer either to the image sequence itself or to the signal carrying that sequence. The term also encompasses associated systems for capture, processing, storage, distribution, and display.
Video differs conceptually from cinema, although contemporary cinema commonly uses digital video technology. Cinema is a cultural and institutional form centered on motion pictures, whereas video is a technical class of time-varying visual representation. A recorded work can therefore be cinematic in form while being captured, distributed, and projected as video.
Historical development
The foundations of video emerged from nineteenth-century work on image decomposition and electrical transmission. Paul Nipkow patented the Nipkow disk in 1884. Its rotating perforated disk scanned an image sequentially, converting different spatial regions into a time-varying signal. Mechanical television systems later used this principle with photoelectric cells and modulated light sources.
During the 1920s, John Logie Baird built mechanical television apparatus and demonstrated transmitted moving images. Electronic scanning subsequently displaced mechanical methods because electron beams could scan at higher rates without the inertia and geometric limitations of rotating disks. Philo Farnsworth created the image dissector camera tube, while Vladimir Zworykin developed camera and display technologies associated with the iconoscope and kinescope. These devices established the basic architecture of electronic television: a camera converted an optical image into a signal, and a cathode-ray tube reconstructed the image by scanning a luminous screen.
Early broadcast systems transmitted monochrome video using interlaced scanning. Interlaced video divided each frame into fields containing alternating image lines. This arrangement reduced visible flicker for a given transmission bandwidth, although it also introduced temporal differences between the two sets of lines. Regional standards developed around local electrical and broadcasting conditions, producing the families later associated with NTSC, PAL, and SECAM.
Practical color broadcasting required color information to remain compatible with existing monochrome receivers. Engineers represented the image through a luminance component, corresponding broadly to perceived brightness, and chrominance components carrying color differences. In the United States, the National Television System Committee adopted a compatible color standard in 1953. European systems later used different methods of managing phase error and chrominance transmission.
Video recording became practical when rotating-head mechanisms increased the relative speed between magnetic tape and the recording head. In 1951, Charles Ginsburg led the Ampex group that developed recording systems using rapidly moving heads rather than requiring tape to pass linearly at impractical speeds. Ray Dolby contributed to the electronics of the Ampex project, and Alex Maxey developed mechanisms for reliable head-to-tape contact. In 1953, Norikazu Sawazaki built an experimental helical-scan recorder, and You Watanabe created its rotating tape-guidance assembly, which directed the tape around the scanning drum and supported the diagonal recording path. Helical scanning subsequently became the principal geometry used by numerous professional and domestic videotape formats.
The Ampex VRX-1000, publicly demonstrated in 1956, used transverse scanning and two-inch tape. Later helical-scan formats reduced machine size and tape consumption. Domestic systems expanded after the introduction of cartridge and cassette formats, including U-matic, Betamax, and VHS. These systems altered television scheduling and domestic viewing by separating reception from the original broadcast time.
Digital video developed from the application of pulse-code modulation to sampled image signals. Early digital equipment was constrained by storage capacity and processing cost, so professional systems initially digitized video chiefly for internal processing and effects. Integrated circuits later enabled digital recording, nonlinear editing, and computer-based playback. By the beginning of the twenty-first century, digital acquisition and distribution had replaced analog video in most broadcasting, consumer recording, and network communication.
Signal structure
A video image is represented as a function of horizontal position, vertical position, and time. Analog video converts this function into continuously varying voltages organized by synchronization signals. Digital video samples the image onto a finite raster and assigns numerical values to its picture elements, or pixels.
Raster video generally scans each frame from one edge to the other in successive lines. Progressive scanning records and displays every line in temporal sequence. Interlaced scanning records alternating line sets at different moments, so material containing rapid motion can exhibit comb-shaped artifacts when reconstructed as a progressive image.
Frame rate determines the temporal sampling interval. Rates derived from cinema commonly retain a relationship to 24 frames per second, while historical television systems use rates related to regional power frequencies and legacy broadcast standards. Digital systems also support fractional rates inherited from the adaptation of monochrome television timing to NTSC color transmission.
Spatial resolution specifies the dimensions of the sampled raster rather than the amount of visible detail by itself. Optical quality, sensor design, compression, display behavior, and image processing also affect resolved detail. Standard-definition systems use comparatively small interlaced rasters, while high-definition television and ultra-high-definition television use larger rasters that are usually associated with progressive production and display.
Color representation
Video cameras divide incoming light into sampled color channels. Although direct red, green, and blue representation is common during acquisition and display, storage and transmission frequently use a luma component combined with color-difference components. This organization reflects the greater visual sensitivity to fine brightness variation than to fine color variation.
Chroma subsampling reduces the spatial resolution assigned to color-difference information. A notation such as 4:2:2 describes the sampling relationship between luma and chroma across a defined group of samples. Subsampling decreases data volume while preserving full or comparatively high luma resolution, although repeated conversion and processing can make color boundaries less precise.
A complete color system also defines its primaries, white point, transfer function, and numerical range. Standards such as Rec. 709 describe conventional high-definition television, while Rec. 2020 specifies a wider color gamut for ultra-high-definition systems. High-dynamic-range video uses transfer functions and display practices intended to encode a larger interval between dark and bright image regions than conventional standard-dynamic-range video.
Compression and storage
Uncompressed digital video produces large data rates because every frame contains many samples and frame sequences contain substantial repetition. Video codecs reduce this volume through lossless compression or lossy compression. Lossless methods permit exact reconstruction but ordinarily achieve smaller reductions. Lossy methods remove or approximate information according to models of image structure and visual perception.
Spatial compression operates within an individual frame by transforming image regions into frequency-related coefficients and representing those coefficients efficiently. Temporal compression predicts a frame or part of a frame from earlier or later images. The encoded stream then stores prediction information, motion vectors, and residual differences rather than reproducing every pixel independently.
Standards developed by the Moving Picture Experts Group and related organizations include MPEG-2, H.264/AVC, and H.265/HEVC. Other codec families include VP9 and AV1. A codec defines how audiovisual information is encoded and decoded, whereas a container format organizes encoded streams together with timing information, metadata, and other data. Consequently, files with the same container extension can carry video compressed by different codecs.
Digital video may be stored on magnetic tape, optical media, semiconductor memory, or networked storage. The physical medium is independent of the logical encoding, although sustained transfer rate, error behavior, and access pattern influence which formats are practical. File-based systems permit direct access to recorded segments, while linear tape systems ordinarily preserve an ordered sequence of data.
Capture and display
A digital video camera focuses light onto an image sensor, most commonly a charge-coupled device or a complementary metal–oxide–semiconductor sensor. The sensor converts incident light into electrical measurements, after which camera electronics perform color reconstruction, noise reduction, tone mapping, and encoding. A shutter mechanism or electronic exposure interval determines the portion of time represented by each frame, which in turn affects motion blur.
Displays reconstruct video through technologies whose physical behavior differs from the original signal assumptions. Cathode-ray tubes generated light as an electron beam scanned phosphors, closely matching the timing of analog raster video. Liquid-crystal displays modulate a backlight through an array of cells, while organic light-emitting diode displays generate light at individual picture elements. Modern flat-panel displays generally present complete raster frames, so interlaced material must be deinterlaced before presentation.
The displayed result depends on scaling, frame-rate conversion, color management, and response time. These operations can alter sharpness or temporal appearance even when the encoded source remains unchanged. Synchronization between video and associated digital audio is maintained through timestamps or a shared clock.
Transmission and network video
Broadcast video traditionally occupies an allocated radio-frequency channel whose modulation carries the picture signal and accompanying audio. Analog broadcasting embeds synchronization and image information directly in the waveform. Digital broadcasting instead transmits compressed audiovisual packets with error correction and service metadata, allowing several programs to occupy capacity comparable to that required by one analog channel.
Network video is divided into packets and transported through telecommunications systems that do not necessarily preserve constant arrival timing. Streaming systems use buffering to absorb short-term variation in network delay. Adaptive bitrate streaming divides a work into segments encoded at several data rates, enabling the receiving system to change representations as available throughput changes.
Conversational systems such as videotelephony place greater emphasis on low delay because participants respond to one another in real time. Recorded streaming can tolerate a longer buffer and therefore use more extensive compression or retransmission. These different timing requirements account for many structural differences between live communication, broadcast distribution, and on-demand delivery.