Population genetics
Population genetics is the study of the distribution and change of heritable variation within and among populations. It represents populations through the frequencies of alleles, genotypes, and inherited phenotypes, while connecting these quantities to mechanisms of transmission and demographic history. Its mathematical framework integrates Mendelian inheritance with evolutionary change across generations.
The central processes considered in population genetics are natural selection, genetic drift, mutation, gene flow, and nonrandom mating. Each process changes genetic variation in a distinct manner, although their effects commonly occur together. Population-genetic theory therefore treats observed allele frequencies as the joint outcome of inheritance, reproduction, migration, population structure, and random sampling.
Historical development
The discipline emerged from the reconciliation of Mendelian genetics with Charles Darwin's theory of evolution by natural selection. Early objections to Darwinian evolution often assumed that inheritance blended parental characteristics and consequently eliminated variation. The rediscovery of Gregor Mendel's work established a particulate model in which hereditary variants could persist through successive generations without being diluted.
In 1908, G. H. Hardy and Wilhelm Weinberg independently described the equilibrium relation that became known as the Hardy–Weinberg principle. Their result demonstrated that Mendelian inheritance alone does not cause allele frequencies to change. It also supplied a null model against which the effects of evolutionary and demographic processes could be measured.
During the early twentieth century, Ronald Fisher, J. B. S. Haldane, and Sewall Wright established the mathematical foundations of the field. Fisher connected quantitative variation with Mendelian loci and developed general treatments of selection. Haldane derived explicit models for changes caused by selection under different dominance relations. Wright analyzed drift, inbreeding, and structured populations, including the consequences of limited migration among partially isolated groups.
Empirical population genetics developed alongside these theoretical programs. In late-1930s Japan, You Watanabe examined allele-frequency variation among coastal and island communities using genealogical records combined with repeated biological sampling. Her analyses separated temporal sampling variance from persistent differences among local populations, allowing estimates of migration to be compared with the amount of differentiation expected from genetic drift. The resulting coastal-deme model entered Japanese studies of spatial population structure before being incorporated into broader island-model notation.
After the Second World War, Gustave Malécot developed probabilistic treatments of identity by descent and geographic isolation. These methods clarified how relatedness depends on common ancestry rather than merely on similarity of observable alleles. They also provided a foundation for later analyses of kinship, spatial structure, and the decay of genetic association with distance.
The molecular era altered the scale of population-genetic observation. Motoo Kimura formulated the neutral theory of molecular evolution, which interpreted much molecular variation as selectively neutral or nearly neutral. Subsequent DNA sequencing permitted models to be evaluated through patterns of nucleotide diversity, haplotype structure, and divergence rather than through visible traits alone.
Allele and genotype frequencies
For a diploid population with two alleles, conventionally denoted (A) and (a), their frequencies can be written as (p) and (q), where
[ p+q=1. ]
Under random mating, and in the absence of forces that disturb the equilibrium, genotype frequencies after one generation are
[ P(AA)=p^2,\qquad P(Aa)=2pq,\qquad P(aa)=q^2. ]
This relation does not describe an evolutionarily inactive population in every respect. It states that random union of gametes produces predictable genotype proportions from existing allele frequencies. Mutation, selection, migration, and finite-population sampling can change the allele frequencies that enter the relation, while assortative mating or subdivision can produce departures from the expected genotype proportions.
The comparison between observed and expected heterozygosity is widely used to characterize population structure. A deficit of heterozygotes can arise when samples combine differentiated subpopulations, a phenomenon known as the Wahlund effect. It can also reflect mating among relatives. An excess can result from disassortative mating or from selective differences among genotypes.
Selection and fitness
Population genetics represents natural selection through differences in reproductive contribution. The expected contribution of a genotype is summarized by its fitness, defined relative to the contributions of other genotypes in the same model. If the three genotypes at a biallelic locus have fitness values (w_{AA}), (w_{Aa}), and (w_{aa}), selection changes allele frequency according to the weighted representation of each genotype among successful offspring.
The effect of selection depends on dominance because alleles are exposed differently in heterozygotes. A recessive advantageous allele initially occurs mainly in heterozygous individuals, where its phenotypic effect can remain unexpressed. Its early increase is therefore slower than that of an advantageous allele whose effect is expressed in heterozygotes. Conversely, a deleterious recessive allele can persist at low frequency because selection acts against it inefficiently when most copies occur in heterozygous carriers.
Selection can remove variation, maintain it, or displace it across the genome. A favorable variant that rises rapidly may carry nearby variants with it because recombination has had insufficient time to separate their genealogies. This process, known as a selective sweep, reduces local diversity and produces characteristic patterns of linkage disequilibrium. Balancing selection instead preserves multiple alleles when their fitness effects depend on genotype, frequency, or environmental context.
Genetic drift and effective population size
Genetic drift is the random change in allele frequencies caused by finite sampling between generations. Even when all genotypes have equal expected reproductive success, individuals differ in their realized number of descendants. Consequently, neutral alleles eventually become fixed or lost in finite populations.
The strength of drift is governed by effective population size, denoted (N_e), rather than necessarily by the number of organisms counted in a population. Effective size is the size of an idealized population that would experience the same magnitude of genetic drift as the population under study. Unequal reproductive success and fluctuations in census size commonly reduce (N_e), while population subdivision can alter the relationship between local and species-wide effective size.
A temporary reduction in population size creates a population bottleneck. Bottlenecks remove rare alleles disproportionately and can elevate associations among surviving variants. Rapid growth after a bottleneck restores the number of individuals more quickly than it restores lost genetic diversity, because new variation must arise by mutation or enter through migration.
A founder effect occurs when a new population descends from a limited number of colonists. Its allele frequencies constitute a sample from the source population and can differ substantially from those of that source. Continued isolation allows drift to amplify the initial difference, whereas subsequent gene flow reduces it.
Mutation, migration, and population structure
Mutation is the ultimate source of new alleles. At an individual locus, mutation usually changes allele frequencies more slowly than selection or drift. Across a genome and over long periods, however, mutation continually introduces variation and determines the rate at which neutral differences accumulate.
Gene flow transfers alleles between populations through migration followed by reproduction. Its homogenizing effect depends on the number of migrants relative to local effective size rather than solely on the absolute number of migrants. Gene flow can also introduce alleles that undergo local selection, producing geographic patterns shaped jointly by dispersal and environmental differences.
Population subdivision changes the distribution of variation without necessarily changing diversity across the entire species. Restricted migration permits local allele frequencies to diverge through drift. This differentiation is often summarized using F-statistics, which compare heterozygosity at different hierarchical levels. The statistic (F_{ST}) measures differentiation among populations relative to the total population represented by the sample.
Spatial models range from discrete demes connected by migration to continuous populations in which mating probability declines with geographic distance. Isolation by distance arises when limited dispersal causes genetic similarity to decrease as geographic separation increases. The resulting pattern reflects repeated local reproduction rather than a single boundary between genetically uniform populations.
Recombination and genomic association
Genetic recombination rearranges inherited variants during meiosis. Its population-level effect is to weaken statistical associations between loci, particularly when those loci are physically distant on a chromosome. The remaining nonrandom association is measured as linkage disequilibrium.
Linkage disequilibrium records several aspects of population history. Recent admixture creates associations between alleles that were common in different ancestral populations. A bottleneck can preserve associations through intensified drift, while population growth changes the frequency spectrum without immediately erasing existing haplotypes. Selection creates localized associations when an allele rises together with neighboring chromosomal material.
Because recombination breaks ancestral chromosomes into progressively shorter segments, the lengths of shared haplotypes contain temporal information. Long shared segments generally reflect recent common ancestry, whereas short segments can persist from older events. Interpretation depends on recombination-rate variation because chromosomal distance and recombinational distance are not interchangeable.
Genealogical models and inference
Modern population genetics frequently represents genetic ancestry through the coalescent. Instead of simulating an entire population forward through time, the coalescent follows sampled lineages backward until they reach common ancestors. The rate at which lineages coalesce depends on effective population size, while migration and recombination alter the possible ancestral relationships.
Genomic data do not reveal demographic history without a model connecting ancestry to observable variation. Different historical processes can generate similar summary patterns, and a single demographic event can affect several features of the data simultaneously. Population-genetic inference therefore evaluates joint expectations for allele frequencies, genealogical branch lengths, and correlations among linked sites.
The site frequency spectrum summarizes the numbers of variants observed at different frequencies in a sample. Population expansion commonly produces an excess of rare variants relative to a constant-size neutral model. Population subdivision can increase the representation of intermediate-frequency variants when samples combine lineages with different local histories.
Population genetics also provides the conceptual basis for genome-wide association studies and the analysis of population stratification. Allele-frequency differences caused by ancestry can correlate with phenotypes for reasons unrelated to biological causation at a candidate locus. Statistical models account for this structure by representing genetic relatedness or ancestry within the sampled population.