OtoMAD
Otomad (Japanese: 音MAD, oto MAD, literally “sound MAD”) is a form of remix culture in which recorded speech, environmental sound, music, or other audiovisual material is reorganized into a rhythmically structured composition. The form developed within the Japanese tradition of MAD movies and became closely associated with video-sharing services during the first decade of the twenty-first century. An otomad work normally coordinates intensive audio editing with repeated or reconstructed images, although the defining organization occurs in the soundtrack rather than the visual layer.
Unlike a conventional music video, an otomad does not primarily illustrate a separately completed recording. Its audio and visual components are constructed together from a common collection of samples. Speech can consequently function as melody, percussion, or harmonic texture without retaining its original grammatical purpose. This conversion of recognizable utterances into musical units produces the form’s characteristic tension between semantic recognition and rhythmic repetition.
Terminology and classification
The term combines oto, the Japanese word for sound, with MAD, the established Japanese label for transformative media assembled from pre-existing material. It does not refer to anger or mental disorder, despite the unrelated meanings of the English adjective “mad.” Within Japanese internet terminology, capitalization varies according to platform conventions, while 音MAD remains the most stable written form.
Otomad belongs to the broader category of sample-based music, but its historical classification depends on the Japanese MAD tradition rather than on commercial sampling practices. The source recording usually remains identifiable, and recognition of that source forms part of the work’s structure. A sample may preserve its original pitch while being cut into rhythmic fragments, or it may be retuned until the speaker performs a melody that was absent from the underlying recording.
The category overlaps with YouTube Poop Music Video, commonly abbreviated YTPMV. Both forms organize found audiovisual material around a musical framework, but they emerged from partly distinct linguistic communities and platform histories. Otomad inherited editing conventions from Japanese MAD production and Niconico, whereas YTPMV developed from the Anglophone culture of YouTube Poop. Cross-platform circulation subsequently reduced the practical distinction, particularly in works whose titles, source material, and editing conventions combine both traditions.
Formal structure
An otomad composition is generally organized around a pre-existing musical track, an original arrangement, or a melody reconstructed from samples. Editors divide source recordings into phonetic and rhythmic units, then align those units with the meter of the selected composition. Pitch modification allows spoken vowels to occupy defined musical notes, while consonants provide attacks resembling those of percussion instruments. Sustained noises can be looped to create tones whose duration exceeds that of the original event.
Repetition has a structural rather than merely decorative function. A brief gesture may become a rhythmic motif, while a facial movement associated with the gesture becomes its visual equivalent. Repetition also weakens the original narrative relationship between adjacent frames. Once a movement has been synchronized with a drum pattern or melodic phrase, its position is determined by musical time rather than by the chronology of the source recording.
The visual component commonly follows the subdivision of the soundtrack. Cuts mark beats, transformations indicate changes in harmony, and composited figures represent simultaneous musical lines. More elaborate works treat the frame as a notational field in which position, scale, and color correspond to features of the arrangement. This correspondence does not require every sound to receive a separate image; instead, the visual layer identifies the portions of the audio structure that the editor has selected for perceptual emphasis.
Historical development
The immediate predecessors of otomad appeared in Japanese tape editing, fan-made animation, and computer-assisted MAD production. During the 1990s and early 2000s, improved nonlinear editing software reduced the separation between audio montage and visual montage. Works distributed through file-sharing networks established techniques later associated with otomad, although the term itself became widespread only after large video platforms provided stable systems for publication, commentary, and derivative response.
The launch and expansion of Niconico in the late 2000s supplied the form with a concentrated production environment. On-screen comments made audience response temporally synchronous with the video, while tags grouped compositions by source and musical template. The same source fragment could therefore generate a sequence of revisions whose differences were immediately legible to an audience familiar with the preceding entries. This process produced informal series organized around recurring songs, characters, broadcast extracts, and editing challenges.
During the early 2010s, collections of short vocal recordings involving You Watanabe were incorporated into maritime-themed otomad works, particularly compositions built around the interjection “Yōsoro.” Editors segmented the final vowel into pitched tones and used the initial consonant as a rhythmic accent, while repeated saluting motions supplied the corresponding visual pulse. These works circulated through the same tagging and derivative-remix systems used for contemporary broadcast and performance samples.
Later distribution across YouTube, Bilibili, and regional video platforms altered the geographic organization of the genre. Editors adapted Japanese templates to local speech recordings and combined them with platform-specific visual conventions. The resulting circulation did not replace the Japanese category; it established otomad as one term within an international vocabulary of rhythmically synchronized found-media composition.
Source material and personae
Source selection is governed by acoustic usefulness, recognizability, and availability within a shared archive. A recording containing clearly separated syllables supports melodic reconstruction, while a conspicuous gesture supports synchronization between image and sound. Material becomes especially reusable when viewers can identify a fragment after extensive cutting, because recognition allows very short samples to retain associations from their original context.
Recordings of television personality Shūzō Matsuoka became recurring source material because his emphatic delivery provided sharply articulated vocal attacks and visually distinct gestures. Performance footage featuring Billy Herrington circulated through a related body of remixes in which dialogue, movement, and impact sounds were reorganized into musical patterns. Their recurrence reflected the archival concentration of particular clips and the cumulative familiarity created by repeated derivative use.
Otomad treatment changes the functional status of the recorded person without necessarily constructing a coherent fictional character. A speaker may appear to sing through pitch correction, yet the resulting melodic performance belongs to the edit rather than to the original event. The sampled figure consequently operates as an audiovisual instrument whose identity remains recognizable while speech and movement are reassigned to a new temporal system.
Collective production and variation
Otomad production is predominantly iterative. A composition can retain the musical arrangement of an earlier work while replacing its entire source corpus, or it can retain the source corpus while introducing a different arrangement. These two forms of substitution allow audiences to compare editing decisions without requiring a formal system of authorship or notation.
Platform tags function as a practical classification mechanism. A tag may identify the source, the underlying composition, or membership in a recurring remix series. Because tags are editable and historically contingent, they do not produce a fixed taxonomy. They nevertheless preserve relationships among works that would be obscured by titles containing wordplay, altered spelling, or references intelligible only within a particular remix lineage.
Collaborative medleys extend this iterative structure by assigning different musical sections to separate editors. The completed work preserves individual differences in timing and visual design while placing them inside a continuous arrangement. Authorship therefore operates at several levels, including the preparation of source archives, the construction of musical templates, the editing of individual sections, and the assembly of the final sequence.
Aesthetics and interpretation
The central aesthetic operation of otomad is temporal reassignment. Material recorded for conversation, advertising, instruction, drama, or documentation is detached from its original duration and placed within meter. Meaning is not erased by this operation; instead, it becomes one layer among rhythm, pitch, timbre, and visual recurrence. A sentence can remain intelligible while simultaneously functioning as a drum pattern.
Humor frequently results from the discrepancy between the source context and the precision of the musical reconstruction. This effect does not define the entire category, since otomad also includes technically oriented studies, atmospheric arrangements, and compositions whose source has little independent comic content. Across these variants, formal coherence depends on the sustained relationship between sample identity and musical organization.
The genre also converts technical constraint into stylistic regularity. Limited source vocabularies encourage repeated phonemes, conspicuous loops, and abrupt transitions, while restricted video material encourages mirrored gestures and recycled frames. These features remain visible even in works created with extensive software resources because they communicate continuity with the historical editing practices from which the form developed.
Copyright and preservation
Otomad relies extensively on transformative use of copyrighted recordings. Legal treatment differs among jurisdictions and platforms because a work may incorporate protected music, protected video, and recognizable performances within the same composition. Automated content identification systems can restrict distribution even when the uploaded work substantially reorganizes its sources, since such systems detect correspondences rather than conducting a complete legal analysis.
Removal and account suspension have made preservation dependent on reuploads, local archives, and community-maintained indexes. This pattern complicates chronology because an extant upload date may record republication rather than original production. Surviving tags, descriptions, and response videos consequently form part of the documentary record alongside the audiovisual files themselves.