An interactive companion

Music, Mathematics & Language

The book opens with a pianist's question: why are the black and white keys arranged as they are? Answering it takes 2,500 years of mathematics — and by the end of the road, music has become something a computer can parse like a sentence, reduce like an equation, and recombine like algebra.

This site teaches the core ideas of Music, Mathematics and Language: The New Horizon of Computational Musicology (Keiji Hirata, Satoshi Tojo & Masatoshi Hamanaka, Springer 2022) — slowly, from zero, with sound. Each chapter builds on the last: first what musical meaning could even mean (I), then the physics and arithmetic of tuning (II), then the idea that harmony has a grammar (III–IV), then two rival theories of how listeners actually parse music (V–VI), and finally what happens when you implement all of it (VII–IX). Read in order the first time; every lab is a thing to hear, not just read.

Temperament keyboardC4 – C5

Play the octave. Then switch the tuning system underneath it and play the same keys again — same note names, different mathematics. The readout shows each note's ratio to C and its deviation in cents from equal temperament.

Press a key…

The oscilloscope draws the sum of the sounding waves. Watch a just-intonation major third lock into a still, repeating shape — then hear the same third in equal temperament shimmer, because 2^(4/12) is close to, but never equal to, 5/4.

Don't worry about the vocabulary yet. Just do this: press E, listen. Switch to Just intonation, press E again. One of them wobbles; one sits perfectly still. The entire second chapter of this site exists to explain those two E's — and why your DAW, your hardware synths, and every piano on Earth chose the wobbly one.
All audio is synthesized live in your browser — nothing to load, but you will want speakers or headphones.
Chapter I

A machine that computes the meaning of music

Before any mathematics, the book has to answer a philosophical objection: "the meaning of music" sounds like the least scientific topic imaginable. Chapter I is about how to make it scientific — by being very careful about what you claim.

Why drag computers into musicology at all?

For most of history, music theory was apprenticeship knowledge. A composer learned harmony the way you learned compression: by imitating people who could do it, absorbing thousands of unwritten rules of thumb. Books existed, but they were collections of practice, not systems — closer to "mixing tips" threads than to physics. Rameau, in Bach's generation, was the first to try to systematize harmony, and even his theory leaned heavily on intuition.

The authors' proposal is simple to state: take those intuitions and force them through a compiler. Write the hypothesis "listeners hear phrase boundaries where the gaps are" as an actual program, run it on actual scores, and compare the output to what human listeners report. If they match, the hypothesis survives; if not, it's refuted and revised. This is Karl Popper's criterion for science — a claim is scientific exactly when something could prove it wrong — applied to musicology. A program is the perfect vehicle for it, because a program can't be vague. Prose can say a boundary goes where the gap is "relatively pronounced"; code has to say how pronounced, compared to what, normalized how. Every fudge becomes visible.

There's a second reason to use programs rather than only human experiments: humans are noisy. Any listening test is contaminated by the subject's training, taste, mood, and what they had for lunch. A simulation isolates the hypothesis itself.

The obstacle: Moravec's paradox

In AI there's a famous observation called Moravec's paradox: the things that are hard for humans are often easy for computers, and vice versa. A computer will tap a metronome-perfect beat forever, name the exact pitch of any note, and separate simultaneous voices — all things trained musicians struggle with. Meanwhile a three-year-old can tell a happy song from a sad one, recognize "that song from the car," and hum something like a tune she heard once — and fifty years of research haven't fully taught a machine any of it.

The difference is tacit knowledge: everything about value, feeling, association, and intuition that humans exchange without ever writing down. Music theory is soaked in tacit knowledge, which is precisely why it resisted computers for so long. The whole book is one long campaign to make pieces of that tacit knowledge explicit — and its honest admission is that you can only do this one careful layer at a time.

A detour that matters: does music "say" anything?

Before you can compute the meaning of music you need to decide what kind of thing that meaning is — and the nineteenth century fought a war over exactly this. On one side, program music: Liszt's term for music that deliberately evokes scenes and stories. Berlioz's Symphonie Fantastique walks you to the guillotine; Vivaldi's Four Seasons paints weather; Beethoven's Pastoral symphony imitates actual thunder. Richard Strauss boasted there was nothing music couldn't express. On the other side, the critic Eduard Hanslick and the idea of absolute music: music refers to nothing at all, and its beauty is the abstract beauty of structure itself — the way an abstract painting can be beautiful without depicting anything. Hanslick thought attaching stories to music was a category error, and he attacked Wagner and Liszt for it mercilessly.

The book's position is pragmatic and important for everything that follows. Music can carry external associations — but those are unstable, private, and hopeless to compute. What can be computed is the meaning music has as structure: the relations of notes to other notes. So the book sides, methodologically, with Hanslick: forget the guillotine, study the architecture. (It also notes, wryly, that history sided the other way — lyrics, opera, and above all film re-attached music to stories, and Hanslick would have hated the entire genre of film scoring.)

Gestalt: meaning that only exists in the whole

So what is "structural meaning"? Start with the oldest observation in perceptual psychology. A gestalt is a percept that appears only when parts combine — something genuinely more than the sum of its components. You see it constantly in vision: three dots become a face. The book's musical example is disarmingly humble: in Japanese elementary schools, teachers open and close each lesson with three chords played to the words "stand up — bow — sit down." Any one of those chords, alone, is just a sound. In sequence, the third chord is unmistakably an ending. Nothing in the physics of chord three carries "ending"; the ending exists only in the relationship.

Gestalt demo · stand up, bow, sit down

Hear each chord in isolation, then in sequence. Isolated, they're inert. Chained, the third chord sounds unmistakably like an ending — a perception that exists nowhere in the individual chords.

I → V⁷ → I  ·  the same progression closes Beethoven's Ode to Joy
Play the three chords in isolation first — really try to hear "conclusion" in chord 3 alone. You can't. Then play the sequence. The finality snaps into place on the last chord. What changed? Not the sound — the context. That difference is the entire subject matter of this book. (The same three chords close Beethoven's Ode to Joy, transposed.)

Now the crucial second observation: nearly everyone hears that third chord as an ending. Your perception is subjective — it happens privately, in your head — and yet it agrees with everyone else's. The book calls this intersubjectivity: subjective experiences that reliably line up across a community. This is the escape hatch from "taste is personal, so nothing can be studied." Taste is personal. But underneath taste there is a layer of perceptions — where phrases end, which chord feels like home, which note is an ornament — that behaves almost like objective fact, because virtually all enculturated listeners report the same thing. That layer is where a science can live.

Separate the layers, or drown

This suggests a strategy the book uses everywhere: hierarchical separation. Put the stable, intersubjective, score-derived structure in a fundamental layer, and treat sensitivity, preference, and artistic judgment as a higher layer built on top of it. Study the foundation first; don't let the penthouse collapse the analysis.

Why insist on this? Because theories that skip the separation — that try to model "how this music feels" in one leap — come out narrow, unstable, and inaccurate, and the book says so bluntly. There's even a general engineering principle behind it, the flexibility–usability tradeoff: any system that gives you precise control is complicated to operate, and any system that's simple hides control from you. You live this daily — a preset romplerish synth is instant but rigid; a full modular rig can make anything and takes years. The tradeoff exists because the lower layer (physics, wiring) and the upper layer (intention, taste) are genuinely different kinds of thing. Good tools separate them; so should good theories.

The same move — divide the system into comprehensible parts, understand the parts, reassemble — is just reductionism, the working method of all modern science. Chemistry didn't crack matter by contemplating whole objects; it found atoms. The bet of this book is that music has atoms too, and that they're relations between notes.

Signs: how anything means anything

If musical meaning is going to be relations, we need the general theory of how one thing points at another — semiotics. The American logician Charles Sanders Peirce sorted signs into three kinds, by how they connect to what they mean:

  • Icon — connected by resemblance. A hieroglyph of a bird means "bird" because it looks like one; the figures on restroom doors are icons.
  • Index — connected by physical causation or measurement. A thermometer means "30 degrees" because the temperature physically pushes the mercury there; smoke is an index of fire.
  • Symbol — connected by pure convention. Nothing about the letter "A" or a quarter-note glyph resembles or is caused by what it denotes; a community simply agreed. Most of language, math, and notation is symbols.

Signs also drift. The book's example: the DeLorean DMC-12 started as a sign for "that odd stainless-steel car" (1981), became a sign for "time machine" (Back to the Future, 1985), and by the 2015 anniversary events had become a sign for the film franchise itself. Meaning isn't in the object; it's in the observer's evolving web of associations.

Now aim this at melody. Music has no dictionary — a G doesn't denote anything — so what do its signs point at? The book's answer: other music. Hear do-re-mi-fa do-re-mi-fa and, by gestalt, you've chunked "do-re-mi-fa" into a unit. The second occurrence points backward at the first ("that again") and forward at a third ("presumably again"). While you listen you are ceaselessly doing two things: searching memory for connections to what you've already heard, and predicting what comes next. Every phrase is simultaneously a sign referring to other phrases and a target being referred to. A melody is a referential network — and where the connections are dense, you perceive a group; where they're sparse, a boundary. Two of the book's major theories fall straight out of this picture: analyze the network's hierarchical grouping and you get GTTM (Chapter VI); analyze its chains of reference and expectation and you get the Implication–Realization model (Chapter V).

Two footnotes from linguistics complete the picture. Saussure split every sign into signifier (the word-sound "apple") and signified (the red fruit in your mind), joined only by convention — which is why the same chord can signify differently in different communities. And Wittgenstein added the warning that such conventions are never fixed contracts: meanings shift as communities use them, like house rules in a children's game — he called it a language-game. Keep both in mind whenever a music theory claims to have found the meaning of anything.

Meyer's map of musical meaning

The music philosopher Leonard Meyer organized all of this into a two-by-two grid. Ask two questions of any claimed musical meaning: does it live inside the music or point outside it? And is its aim emotion or form?

Intrinsic · Emotion

Meaning inside the music: phrases raise expectations, which are realized or denied, and the denial stirs feeling. This is the computable heart of musical meaning — Chapter V is entirely about it.

Extrinsic · Emotion

"Darling, they're playing our tune": music wired to a memory, a scene, a person. Completely real, completely private, computationally hopeless — and the book politely sets it aside.

Intrinsic · Form

Pleasure in structure as such: minimalism's slow permutations, or James Brown's Sex Machine — nearly one chord (E7) for the whole track, meaning generated by pattern alone.

Extrinsic · Form

Structure that refers outward: Xenakis composing from probability formulas (what matters is that the score satisfies the math), arabesques whose melodic outline echoes mosque ornament.

The book's program targets the left column — intrinsic meaning — and mostly the top-left cell. That's not a limitation to apologize for; it's the layer-separation strategy doing its job.

Schenker: every piece hides a skeleton

One more giant, and he's the bridge to everything computational that follows. The Viennese theorist Heinrich Schenker (1868–1935) proposed the reduction hypothesis: in any tonal piece, some notes are structural and some are decoration, and if you repeatedly strip the decoration you descend through layers to a simple skeleton that is the piece's deep identity.

His terms, worth learning because Chapter VI inherits them: the note-for-note surface is the foreground (Vordergrund); stripping ornaments takes you through several middleground layers (Mittelgrund); at the bottom lies the background (Hintergrund) — the Ursatz, or fundamental structure. And Schenker claimed the Ursatz is nearly universal: an upper voice stepping down to the tonic (the Urlinie, e.g. 3̂–2̂–1̂) over a bass that rises a fifth and falls home — I–V–I. The stripping operation is reduction; the reverse — growing a rich surface from a skeleton by inserting compatible ornaments and harmonies — is elaboration. Think of it as the difference between a full arrangement and its chord-track: reduction is bouncing the arrangement down to the sketch it grew from; elaboration is producing the arrangement up from the sketch.

What makes prolongation psychologically real is worth pausing on: when a structural chord sounds, it doesn't stop existing when the next note starts. It keeps governing — everything that follows is heard against it until something equally strong replaces it. Schenker called this persistence prolongation, and it's why an ornamental run doesn't destroy your sense of "we're still on the I chord." A long reverb tail is a decent physical metaphor for a fundamentally mental phenomenon.

And here is the problem that motivates half this book: Schenker's method works — analysts produce compelling reductions — but he described it in intuitive, almost mystical prose. Which notes are "important"? He couldn't say algorithmically, and neither could anyone after him. A brilliant theory that no computer can run is exactly the kind of tacit knowledge this book exists to externalize.

Whose structure is it? Creation vs. reception

A last conceptual tool. The semiologist Jean Molino pointed out that every piece of music involves three distinct things: a creation process (the composer's structures and intentions), a reception process (the listener's), and between them the physical trace — the score or recording, the only part you can put on a table and analyze. Crucially, creation and reception need not agree. The composer may build structure the listener never perceives; the listener may perceive structure the composer never intended. Communication is not guaranteed — the trace just sits there, letting both sides project into it.

Retrograde · creation vs. reception

Composers since the Middle Ages have hidden melodies played backwards (Bach's crab canon; Beethoven's Hammerklavier fugue). But listeners almost never hear a retrograde — most only spot it in the score. Try it honestly: play the melody, then its reverse. Would you have noticed the relation blind?

Molino's trichotomy: the composer's intention and the listener's reception meet only in the physical trace — the score — and they need not agree.
Play the melody a couple of times, then its retrograde. Be honest: blind, would you have caught the relation? Composers from the medieval crab canon to Bach's Musical Offering to the Hammerklavier fugue have hidden exact reversals — structure that is fully real on the creation side and almost never crosses to the reception side. Analysis has to say which side it's describing.

There's even a hard number underneath this. Human short-term memory holds about 7±2 chunks for a handful of seconds. Convert that to music at 80 BPM in 4/4 and you get roughly 10–20 beats — about two and a half to five bars. A four-bar phrase fits in the buffer, which is why four-bar phrases feel like natural units everywhere from Mozart to techno. A forty-bar span does not fit; whatever structure lives at that scale (like Beethoven's buried retrograde) belongs to the composer's side of the trace, not to moment-to-moment hearing. So when theory books analyze forty-bar arcs "as heard," be suspicious — a different cognitive mechanism, or the composer's blueprint, is doing the work.

Cognitive reality: the admission standard

The chapter closes by setting the bar every later theory must clear. A concept or model has cognitive reality when it's needed to rationally explain what listeners actually perceive — the way "short-term memory" earns its keep by explaining why short pitch sequences are cognized stably. It's not enough for a model to be elegant or to compress the data; it has to correspond to something the mind actually does. The book's bet, backed by a line of experiments (Dibben's hierarchical-hearing studies, Krumhansl's tonal hierarchies, Bod's memory-plus-rules model), is that tree structures over notes have cognitive reality: when you listen, you really do build something tree-shaped. The rest of the book takes that bet and runs.

Music's private associations can't be computed, but its structure can: listeners chunk notes by gestalt, nearly everyone chunks the same way (intersubjectivity), phrases refer to other phrases like signs, decoration hangs off skeleton (Schenker), and the whole arrangement is plausibly a tree in the mind. Make each of those claims precise enough to run as a program, and musicology becomes refutable — a science.

Check yourself
Why do the authors insist on separating a "fundamental layer" from taste and artistic quality?

Taste varies per listener; the structures grasped through gestalt and intersubjectivity are stable across listeners. Only the stable layer supports refutable, reproducible claims — the subjective layer can then be modeled on top. Theories that skip the separation come out narrow, unstable, and inaccurate.

In Peirce's terms, is a crotchet (quarter note) an icon, an index, or a symbol? What about a VU meter?

The crotchet is a symbol: pure convention connects the glyph to "one quarter of a bar." A VU meter is an index — the signal physically drives the needle. A waveform display drawn to look like the sound it represents leans iconic.

What does Molino's trichotomy say about a hidden retrograde in a fugue?

It's structure on the creation side, embedded in the trace (the score), that mostly fails to arrive on the reception side — listeners don't hear it, they discover it by reading. Analysis must state which of the three perspectives it describes, because they genuinely differ.

Why is the four-bar phrase such a universal unit?

Short-term memory holds 7±2 chunks for seconds — at typical tempos, about 2.5–5 bars of material. A four-bar phrase fits inside the buffer, so it can be grasped as one gestalt; longer spans require a different mechanism (or the composer's blueprint) to hold together.

Chapter II

The mathematics of ebony and ivory keys

Why twelve notes? Why those twelve? Why does your synth's oscillator read 440.0 and not some other number? This chapter answers the pianist's question from the top of the page — and the answer is a 2,500-year negotiation with one stubborn arithmetic fact.

Start from what you already know: waves

You work with this physics every day, so briefly: a vibrating string (or air column, or speaker cone) displaces air in a repeating cycle. Cycles per second is frequency — pitch. Displacement size is amplitude — loudness. And a real instrument's tone isn't one sine wave but a stack of them — which is timbre, and the subject of the overtone bench below. For now we only need one question: when do two simultaneous frequencies sound good together?

The classical answer: when their waves line up often. If string B vibrates exactly twice per cycle of string A — frequency ratio 1:2 — then every single cycle of A, the two waves re-synchronize: their zero-crossings coincide, their peaks interlock, and the combined waveform is short and repeating. The ear hears this locking as smoothness, so much smoothness that we give the two pitches the same name: the ratio 1:2 is the octave. C4 at 261.6 Hz, C5 at 523.3 Hz. The next-simplest ratio, 2:3, locks almost as tightly: that's the perfect fifth — C4 at 261.6 Hz against G4 at 392.4 Hz. The rougher the ratio (larger numbers), the longer the combined wave takes to repeat, and the rougher it sounds. Simple ratios = consonance. That single sentence is the physics under all of harmony.

Go back up to the hero keyboard, pick Just intonation, and hold C with the "Play C–G fifth" button while watching the oscilloscope. The trace is a short, stable, repeating shape — that's 2:3 locking. This visual will be your reference for everything that follows.

Pythagoras: an entire scale from two numbers

Legend says Pythagoras heard consonant intervals in a blacksmith's hammering and went home to experiment with strings (the story is probably apocryphal — the discovery isn't). His school held as doctrine that all numbers were ratios of integers, and he built a scale on the two smallest useful ones: 2 (the octave) and 3 (the fifth, once you fold it with 2 into 3/2). The recipe:

  1. Start at C, call its frequency 1.
  2. Multiply by 3/2 to get a new note (up a fifth).
  3. If the result left the octave (≥ 2), halve it to fold it back in. Repeat.

Walk the first steps by hand, because the pattern matters. C = 1. One fifth up: 3/2 — call it G. Another fifth: 3/2 × 3/2 = 9/4, which is past the octave, so halve it: 9/8 — call it D. Again: 9/8 × 3/2 = 27/16 — A. Again: 81/64 — E. Five notes in — C, G, D, A, E — rearrange them by pitch and you have C–D–E–G–A: the pentatonic scale, the five-note scale found in folk music on every continent. (The note names are hindsight, by the way — the generation order is fifths; the alphabet came later.) Astonishingly, Jing Fang in Han-dynasty China derived the same scale the same way at nearly the same time, by cutting bamboo pipes to 2/3 length — the method called sanfen sunyi. Two civilizations, no contact, same arithmetic: this construction is something you discover, not invent.

Keep stacking — B, F♯, C♯, G♯, D♯, A♯, E♯ — and after the twelfth fifth you arrive at B♯, which by all rights should be C again, closing the circle. Should be.

Stack of fifths · find the comma

Stack fifths one at a time. Each new note is (3/2)ⁿ folded back into one octave. The twelfth note should land back on C — listen to how close it gets.

0 fifths stacked — C = 1
Stack all twelve, watching the fractions grow. The twelfth note lands on 2.02787… where a perfect octave demands 2 — about a quarter of a semitone sharp. Then press Hear B♯ against C. That slow pulsing is two frequencies 3.6 Hz apart interfering — the same beating you'd hear from two oscillators slightly detuned. You are listening to the Pythagorean comma: the error at the bottom of Western tuning.

Why can't the circle ever close? Twelve fifths is (3/2)¹² = 3¹²/2¹²; seven octaves is 2⁷. Closing would require 3¹² = 2¹⁹ — a power of 3 equal to a power of 2. But powers of 3 are always odd and powers of 2 always even, so no stack of pure fifths, however long, ever returns exactly home. Twelve is simply the first count that comes close (the overshoot, 531441/524288 ≈ 1.0136, is about 23.5 cents — small enough that scales pretend it away). Want closer? The next good approximation closes after 53 fifths, missing by only 3.6 cents — and yes, a 53-key-per-octave harmonium was actually built. After that, 665. Every tuning system in history is a policy decision about this one leftover error: where do you hide the comma?

Pythagorean practice hid it by building fifths upward from C to G♯ and downward from C to D♭ and letting the leftover land between G♯ and E♭ — an interval so sour it was named the wolf, after the howl. There was no principled reason for that spot; it just inconvenienced the fewest pieces. A second blemish: in the Pythagorean scale the "half" tone isn't half. The whole tone is 9/8 (204¢) but the diatonic semitone comes out 256/243 (90¢) — two of them fall short of a whole tone, and sharps and flats split into two different sizes. The tidy keyboard picture was already cracking.

Just intonation: inviting the number 5

The Pythagorean major third — four stacked fifths, 81/64 — is mathematically pedigreed and sounds harsh: 81:64 is not a simple ratio, and thirds were long treated as dissonances partly for this reason. But notice: 81/64 = 1.2656, and just below it sits 5/4 = 1.25 — a genuinely simple ratio, involving the next prime, 5. The gap between them, 81/80 (about 22¢), is the syntonic comma — the second famous comma, not to be confused with Pythagoras's.

Just intonation is what you get by swapping the 3-heavy thirds for 5-based ones. The payoff is spectacular: the major triad becomes exact small integers, C:E:G = 4:5:6, and every scale interval locks — C:E = 4:5, E:G = 5:6, C:A = 3:5. Chords stop shimmering entirely. This is why barbershop quartets, string ensembles, and choirs — anyone who can bend pitch freely — drift toward just intervals when they sustain a chord: the beating vanishes and the sound "rings."

Where does the shimmer physically come from? From overtones. C4's 5th harmonic is 1308 Hz; a pure E4's 4th harmonic is the same 1308 Hz — aligned, silent. Tune that E the equal-tempered way (329.63 Hz, as your DAW does) and its 4th harmonic sits at 1318.5 Hz: the two partials fight at about 10 Hz, and you hear it as the familiar gentle churn of a tempered third. Detuned-oscillator beating and "out-of-tune interval" roughness are literally the same phenomenon at different addresses in the spectrum.

So why didn't just intonation win? Because its beauty is bought with unevenness. Check two adjacent whole tones: C→D is 9/8 = 1.125 but D→E is (5/4)/(9/8) = 10/9 = 1.111. Two different sizes of whole tone. The semitone, 16/15, fits evenly into neither. The scale is a beautifully consonant crystal — with an irregular lattice. And that irregularity is fatal to one thing musicians increasingly wanted to do…

The modulation problem

Modulation means changing key mid-piece — restating material from a new starting note. Say your melody begins C→D, an interval of 9/8, and you want to restate it beginning on D. The parallel move D→E needs to also be 9/8 — but in just intonation D→E is 10/9. On a violin you'd just bend; on a keyboard, where every pitch is pre-tuned metal, the transposed melody comes out warped, and its accompanying chords worse. It's exactly the problem of a sampler patch mapped from a single sample: perfect at the root key, increasingly wrong as you play away from it. Fixed-pitch instruments plus uneven scales = broken transposition. Something had to give.

Meantone: sacrifice the fifth to save the third

The Renaissance compromise, quarter-comma meantone, chose thirds. The reasoning is lovely: four stacked fifths make a third (plus octaves), and that third comes out one syntonic comma sharp. So shave each fifth by a quarter of that comma, and after four of them the errors sum to exactly one comma — landing the third on a pure 5/4. Solve it as algebra: you need a fifth x with x⁴ = 5 (a pure third two octaves up), so x = ⁴√5 ≈ 1.4953 — audibly close to 1.5, just 5.4¢ flat. Bonus: all whole tones become equal (√5/2, hence "mean tone"). Costs: fifths everywhere slightly impure, semitones still uneven — and the comma, squeezed out of eleven fifths, piles undiluted into the twelfth: the meantone wolf between G♯ and E♭, 35¢ too wide and genuinely unusable. Composers of the era simply avoided keys that touched it; Kepler, Mersenne, Euler and Rousseau all published rival tweaks; the "well-tempered" tunings of Bach's time (Werckmeister and friends) were subtler redistributions that made every key playable while leaving each a slightly different color.

Comparison bench · four temperaments

One interval, four tunings. The cents column shows deviation from just (pure) intonation — the smaller, the smoother. Then meet the wolf.

Pythagorean → Just → Meantone → Equal, in order
Choose the major third and play all four systems in a row: Pythagorean (harsh, +22¢), just (pure, still), meantone (pure again — that was the whole point), equal (+14¢, gently churning). Then press the wolf button and hear where meantone hid the bill.

Equal temperament: perfectly, uniformly wrong

The endgame abandons purity outright. Demand only this: twelve semitones, all identical, closing exactly at the octave. Then each semitone must be the ratio r with r¹² = 2, i.e. r = ¹²√2 ≈ 1.05946 — an irrational number, which would have scandalized Pythagoras. No interval except the octave is pure anymore: the fifth is 700¢ against a true 702, the major third 400 against a true 386. But every key is now exactly the same shape, so any melody transposes perfectly to any starting note, the wolf is smeared invisibly across all twelve fifths, and one keyboard plays everything. Equal temperament trades maximal beauty in one key for guaranteed adequacy in all of them.

To compare systems on one ruler we use cents: 1200 per octave, 100 per equal semitone, logarithmic — cents are to frequency ratio exactly what dB are to amplitude ratio, a log scale that turns multiplication into addition. The four systems, side by side:

Scale stepCDEFGABC
Equal020040050070090011001200
Meantone019338650369789010831200
Just020438649870288410881200
Pythagorean020440849870290611101200

Green = pure interval; red = the wide Pythagorean third. Equal temperament's 400¢ third is 14¢ sharp of pure — the shimmer in the hero keyboard.

A useful mental picture from the book: put the twelve pitch classes on a clock face, C at noon, one semitone per hour. Going around the dial once doubles the frequency — but be careful what the dial hides. Positions add (a semitone plus a semitone is two semitones) while frequencies multiply (1.0595 × 1.0595). The clock works because pitch perception is logarithmic: equal multiplicative steps feel like equal distances. Same reason faders are marked in dB. Roll the clock upward in a spiral — each turn one octave higher — and every C stacks on a vertical line: that vertical line is what "pitch class" means.

Now the pianist's question has its answer. Why are the black and white keys arranged as they are? Because the keyboard is a map of history: seven white keys for the diatonic scale the Greeks distilled from stacked fifths, five black keys wedged in where the chromatic notes were later squeezed, all of it finally tuned by ¹²√2 so the pattern means the same thing from any starting key. (Even the colors are convention — 17th-century harpsichords often reversed them.) And the book adds a sharp observation about what standardization cost: the five-line staff still encodes pre-equal-temperament pitch spacing — E–F and B–C look like every other step but aren't — and equal temperament's power to absorb any ethnic scale into its grid quietly erases the microtonal identities it approximates. Every quantizer discards something.

What is an overtone?

Everything above kept appealing to harmonics, so meet the series properly. A vibrating string doesn't only vibrate whole: it simultaneously vibrates in halves, thirds, quarters… producing frequencies at integer multiples of the fundamental. Partial 2 is the octave; partial 3, an octave plus a fifth (Pythagoras's brick); partial 4, two octaves; partial 5, two octaves plus a pure major third (just intonation's brick); partial 6, two octaves plus a fifth. Then it gets interesting: partial 7 lands near a minor seventh but 31¢ flat — the sweet "barbershop seventh" that no 12-note keyboard can play; partials 11 and 13 fall between the cracks entirely. The harmonic series is nature's chord, and every tuning system is a policy about which of its rungs to honor with a key.

Overtone bench · additive drone on C2

Toggle partials of a C2 fundamental (65.4 Hz) and start the drone — you are doing additive synthesis, and simultaneously touring the raw material of harmony. The table shows where each partial falls against equal temperament.

no partials selected
nHzNearest ET noteDeviationBook’s note
Start the drone with partials 1–5 (first preset): that's a root-position major chord assembled from raw physics — no scale involved. Then add partial 7 and hear the blue seventh your keyboard doesn't own. Then try the odd-partials preset — odd harmonics only is the recipe for a square-ish, clarinet-hollow timbre, which is why this bench is also just additive synthesis wearing a theory hat.

Tuning theorists map these prime ingredients on the Euler lattice — powers of 3 (fifths) running horizontally, powers of 5 (pure thirds) vertically — so each tuning system becomes a shape: Pythagorean tuning is a single horizontal line, just intonation a compact block, and a cadence a leftward drift. Eitz's notation compresses the same idea into superscripts: just-intonation E is written E−1, "Pythagorean E lowered by one syntonic comma."

What equal temperament unleashed — and provoked

The historical correlation is hard to ignore. Keyboard tuning converges on workable temperaments; Bach writes the Well-Tempered Clavier (1722) — a demonstrative lap through all 24 keys; and within the next 150 years nearly every canonical name appears: Haydn, Mozart, Beethoven, Schubert, Schumann, Wagner, Brahms, Mahler. The book argues the completed 12-tone system plus the completed keyboard was the enabling technology — a stable, universal platform, and platforms breed golden ages. Then the twentieth century spent itself trying to escape the platform: Debussy's whole-tone scale (six equal steps, tonality dissolved into haze), Bartók pouring Balkan folk modes into the grid, Stravinsky's polytonality, and Schoenberg's dodecaphony — rows using all twelve tones exactly once, tonality abolished by bookkeeping. Schoenberg predicted people would whistle twelve-tone rows within a century. They don't. The gravitational pull of the tonal system — of these particular compromises — turned out to be enormous.

Symmetry: the octave becomes algebra

Equal temperament's deepest gift is invisible: because all twelve semitones are identical, pitch relationships become pure arithmetic, and the right mathematics for structured arithmetic is group theory. A group is any set of operations that (1) compose without leaving the set, (2) associate, (3) include a do-nothing identity, and (4) can each be undone. Rotating an equilateral triangle by 120° generates a tiny group: three rotations, then you're back. Musical operations form groups constantly — the three inversions of a triad (ρ³ = identity), the inversion-and-retrograde operators of twelve-tone music.

Now let ρ = "raise by one semitone." Apply it twelve times and you're home: ρ¹² = e. The twelve pitch classes under ρ form the cyclic group ℤ/12 — the integers with arithmetic mod 12, i.e. clock arithmetic, which is why the clock-face picture works. Immediate consequences: ρ⁷ = ρ⁻⁵ (up a fifth is the same place as down a fourth), and every complementary interval pair — P4↔P5, M3↔m6, M2↔m7 — is just two exponents summing to 12. Define ξ = ρ⁷ and watch: e, ξ, ξ², ξ³… visits C, G, D, A… — all twelve notes in circle-of-fifths order before closing. ξ generates the whole group, and its orbit is Pythagoras's construction, rediscovered as algebra. Run it backwards (ξ⁻¹) and you get the circle of fourths. Twenty-five centuries between the blacksmith legend and the group axioms, and it's one object.

Cyclic group ℤ/12 · generators ρ and ξ

Apply an operator repeatedly and watch it walk the octave. ρ (semitone) visits notes chromatically; ξ = ρ⁷ (fifth) visits all twelve in circle-of-fifths order before returning home — both generate the whole group.

at C — ρ⁰(C)
Apply ξ twelve times and count: you visit every pitch class exactly once — G, D, A, E, B, F♯… — and land back on C. Then reset and do the same with ρ. Two different orderings of the same twelve notes, two generators of the same group. When a producer friend talks about "moving around the circle of fifths," they are — in the precise mathematical sense — iterating ξ.

Consonance is simple frequency ratios locking. Stacking the simplest ratio (3/2) generates the scale but overshoots the octave by an irreducible comma, because 3ⁿ never equals 2ᵐ. Just intonation buys pure chords at the cost of broken transposition; meantone saves thirds and breeds a wolf; equal temperament spreads the error evenly — every interval slightly wrong, every key identical — and turns pitch into clock arithmetic: the cyclic group ℤ/12, whose fifth-generator ξ is the circle of fifths.

Check yourself
Why can't a stack of pure fifths ever return exactly to C?

It would require (3/2)ᵐ·(1/2)ⁿ = 2, i.e. 3ᵐ = 2ⁿ⁺ᵐ⁺¹ — impossible, since powers of 3 are odd and powers of 2 even. Twelve fifths merely come close, overshooting by the Pythagorean comma (~23.5¢); every tuning system is a strategy for hiding that error.

Where does the "shimmer" of an equal-tempered major third physically come from?

From colliding overtones: the root's 5th harmonic and the third's 4th harmonic coincide exactly when the third is pure (5/4), but sit ~10 Hz apart when the third is tempered (2^(4/12)). The partials beat against each other — the same interference as two slightly detuned oscillators.

What does just intonation buy you, and what does it cost?

Buys: maximally consonant chords (C:E:G = 4:5:6) — the ringing choirs and barbershop quartets get. Costs: two unequal whole tones (9/8 vs 10/9), so melodies can't be transposed on fixed-pitch instruments — modulation breaks. Equal temperament makes exactly the opposite trade.

Why are cents (and dB) logarithmic?

Perception of pitch (and loudness) tracks ratios, not differences: each octave is a doubling, each equal semitone a multiplication by ¹²√2. A log scale converts those multiplications into equal additive steps — so intervals can be added like distances. 1200¢ per octave is to frequency what dB is to amplitude.

Why is 7 ≡ −5 (mod 12) musically meaningful?

Rising a fifth (7 semitones) lands on the same pitch class as falling a fourth (5 semitones). Complementary interval pairs summing to 12 are octave-inversions of each other: P4↔P5, M3↔m6, M2↔m7 — the symmetry table of the chapter, stated as modular arithmetic.

Chapter III

Music as formal language

Darwin suspected our ancestors sang before they spoke — that language and music share one origin. If that's true, music should have something like a grammar. This chapter builds the machinery to ask that question precisely: what kind of grammar, running on what kind of mental machine?

Songs, birds, and the common root

We casually say birds "sing," but from the bird's side chirping is communication — mostly males advertising to females. It strikes us as song because it has contour and rhythm, the two things a beak-limited animal can vary. Humans, with soft lips and agile tongues, added consonants and vowels on top of contour — and got both speech and song from the same equipment. The book endorses the now-widespread view that music and language grew from one proto-communication; Darwin put it beautifully in 1871: our progenitors, "before acquiring the power of expressing their mutual love in articulate language, endeavored to charm each other with musical notes and rhythm." The two remain entangled — chanson only really works in French, Beethoven built the accents of "Muss es sein? Es muss sein!" directly into a string-quartet theme, and every language's stress patterns are, note for note, musical specifications: high–low, long–short, strong–weak.

But there's an obvious difference too. Language has visible units — words, sentences, the full stop. A melody is just… notes. No spaces, no punctuation. So if music has grammar, the grammar is latent, and we'll need real machinery to expose it. Enter Chomsky.

What a grammar actually is

Forget schoolbook grammar; in this chapter a grammar is a machine for generating sentences: a finite set of rewriting rules that, applied repeatedly, produce every valid sentence and nothing else. A miniature English:

S → NP VPa sentence is a noun phrase then a verb phrase
NP → Det Na noun phrase is a determiner then a noun
VP → TV NPa verb phrase can be a transitive verb plus its object
N → cat, girl, … · Det → a, the · TV → bites, …vocabulary

Start from S and keep rewriting: S → NP VP → Det N VP → Det N TV NP → … → "A cat bites the girl." The history of rewrites forms a parse tree — S at the root, words at the leaves, phrases as branches. Chomsky's 1957 bombshell was that a finite rule set plus recursion generates infinitely many sentences — and his stranger, bolder claim was that human children are born with a language acquisition device, an innate universal grammar that experience merely configures. His framework went through five decades of revisions (deep vs. surface structure, X-bar theory, government and binding, the minimalist program), but the core picture — sentences as trees generated by rules — is the piece music theory borrows.

Machines that recognize languages

Every class of grammar has a twin: an abstract machine that can check whether a string belongs to the language. The simplest is the finite automaton: a handful of states, transitions triggered by input symbols, some states marked "accepting." It reads left to right, one symbol at a time, and it has no memory beyond its current state — like a switch with a few positions.

Finite automaton · accepts strings with ≥ 3 a's

The book's example machine: four states, transitions on a and b. Type any string of a's and b's and step through it. It reaches the accepting state ③ exactly when the string contains at least three a's — equivalent to the regular expression b*ab*ab*a(a+b)*.

state 0 · nothing read yet
Run the default abaa and watch the state pointer: b's loop in place, each a advances one state, and reaching state ③ means "accepted." Then try to design an input it wrongly accepts — you can't; the machine exactly recognizes "strings with at least three a's." Now try to imagine a machine like this counting matched things — every a later answered by exactly one b. You can't do that either, and the reason why is the hinge of this whole chapter.

Languages these machines can recognize are called regular, and they have a compact notation you already know from programming: regular expressions. The automaton above is b*ab*ab*a(a+b)*. Regular languages are the bottom rung of a famous ladder:

TypeGrammarMachineCan handle
3RegularFinite automatonlocal patterns — birdsong
2Context-freePush-down automaton (stack)nesting: aⁿbⁿ — human syntax, cadences
1Context-sensitiveLinear-bounded automatonrules that depend on surroundings
0Phrase structureTuring machineanything computable

The rung that matters is type 2, the context-free grammar (CFG): rules whose left side is a single symbol, rewritten regardless of context. The canonical example is two rules — S → aSb and S → ε (nothing) — which generate ab, aabb, aaabbb… every aⁿbⁿ. Watch the recursion: S → aSb → a(aSb)b → aa(ε)bb = aabb. Each application opens an a that the same application promises to close with a b. A finite automaton can't recognize this language, because checking the counts match requires remembering an unbounded number — and a finite automaton's entire memory is which of its few states it's in.

The stack: why nesting needs memory

What's the minimal upgrade? A push-down stack: a first-in-last-out memory, like a stack of plates — you can only add to the top (push) or remove from the top (pop). An automaton with a stack can recognize aⁿbⁿ trivially: push each a, pop for each b, accept if the stack empties on time. And here's the deep connection to human language: dependencies in a sentence nest without crossing. In "We go to Broadway to see a musical," "go" binds "to Broadway" and "to see" binds "a musical," and the bindings tuck inside one another — you can't say "We go to see to Broadway a musical." Nesting-without-crossing is exactly the discipline of balanced parentheses, and balanced parentheses are exactly what a stack checks. Three statements, one fact:

  • dependencies never cross;
  • parentheses must close innermost-first;
  • working memory behaves like a push-down stack.

So human language sits (with a few genuine exceptions — Dutch subordinate clauses and French clitic pronouns really do cross) at the context-free level, while birdsong, transition-diagrammable, is regular. The evolutionary claim hiding here is startling: what separates us from the finches may be, at bottom, a stack.

Now aim it at harmony

Listen to yourself listening to chords. When a dominant sounds, something is owed: the V chord opens an expectation that a coming I will close. In stack terms, hearing V pushes an open parenthesis; the resolving tonic pops it. That tension you feel in an unresolved progression is — the book argues — your syntactic working memory holding open brackets.

Make it concrete with the cadence, the formula of an ending: I ⋯ V–I (authentic) or I ⋯ IV–V–I, or the plagal IV–I. If cadences were mere fixed formulas, a finite automaton would suffice. But cadences embed. Take V–I and ask: how do you intensify V? Precede it with its own fifth — the chord a fifth above V, called the double dominant II (in C: D major, with its out-of-key F♯). Structurally, you've replaced V by (II–V), a little cadence onto V, nested inside the big cadence onto I: V–I → (II–V)–I. That's recursion, the S → aSb move, and it's precisely what regular grammars can't express. If you play jazz you've spent your life inside this rule — ii–V–I is literally a recursive grammar expansion, and the long pre-dominant chains of bebop (iii–VI–ii–V–I…) are the recursion applied again and again.

People have actually written these grammars

This isn't hand-waving; the chapter surveys fifty years of concrete chord grammars, each importing a different tool from computational linguistics:

  • Winograd, 1968. The first real CFG for harmony — cadence → plagal | authentic; authentic → dominant tonic; dominant → V; V → Vseventh | Vtriad… — implemented in LISP, taking scores in and printing scale-degree analyses, decades before "music informatics" had a name. Its honest limitation: real cadences are fuzzy and chord labels ambiguous, so a naïve parser can't guarantee the right reading.
  • Rohrmeier's Generative Syntax Model, 2011. A full generative apparatus for tonal harmony. A piece generates tonal regions (TR → DR t: a tonic region can be a dominant region then a tonic); regions generate Riemann's chord functions — tonic, subdominant, dominant and their "parallel" substitutes (t → tp, the vi-for-I move); functions map to scale degrees (t → I, d → V | VII); modulation rules let any region become a local key; secondary-dominant rules insert D(X) before any X — the ii–V machine again, fully general. Read a derivation top-down and you watch a key exhale into chords.
  • Probabilistic CFG. Attach a probability to each competing rule (summing to 1 per symbol) and two things happen. Parsing gets a tiebreaker: for an ambiguous sentence like "Book the flight through Houston" (did you book through Houston, or is it a flight through Houston?), each parse tree's probability is the product of its rules', and you prefer the likelier tree. And style becomes numbers: the book quotes a rule set fitted to Bach's major-key music — T → i [0.8] | iii [0.1] | vi [0.1], phrase shapes like T S1 D T [0.35]. Same rules with different probabilities = a different composer, a different era, a different genre.
  • HPSG — grammar with features. V, V⁷, and an inverted V⁶ are one category with different internal feature structures — think one chord "patch" with parameters, rather than three unrelated presets. Rules unify features of adjacent chords into parent categories; the book's own work (Tojo et al.) parses C–Am–F–Dm–G–C into a complete cadence tree this way, ii functioning as specifier to V, V to I.
  • CCG — jazz, categorially. In combinatory categorial grammar every word (or chord) is a function hungry for arguments: a transitive verb is (t\e)/e, "something that takes an object, then a subject, and yields a sentence." Steedman-school jazz analysis types each chord by what it demands next, so Dm7–G7–C parses two ways — a chain of falling fifths, or Dm7 as substitute for F (a IV–V–I). Both parses are legitimate hearings; the ambiguity is the point.
PCFG roulette · roll a Bach-flavored cadence

The book quotes a probabilistic grammar fitted to Bach's major-key pieces — rules like T → i [0.8] | iii [0.1] | vi [0.1] and phrase shapes like S → T S1 D T [0.35]. This bench actually rolls those dice: each press samples a phrase and a final cadence from the book's probabilities, shows the derivation with its probability, and plays the result in C major. Same rules every time — the numbers are what make it Bach-shaped.

press generate — every output is a fresh sample from the grammar
Generate five or six cadences in a row. Notice what stays constant (the T…D–T skeleton — the grammar's well-formedness) and what varies (which chords realize each function — the dice). Then look at the derivation line: every choice is annotated with its probability, and the product is the probability of the whole progression under "Bach." This is composition as sampling from a distribution — the 1970s ancestor of every neural music generator you've tried, with the advantage that you can read the entire model in ten lines.

One more voice belongs in this story: Leonard Bernstein, whose 1976 Harvard lectures (The Unanswered Question) brought the Chomsky–music analogy to a mass audience — mapping prosody onto chord progressions, morphemes onto notes, sentences onto phrases. Not all his mappings hold up, and the book says so, but the instinct — that the deep comparison is between syntaxes, not surfaces — was right, and this chapter is that instinct made rigorous. (Xenakis, for the record, dissented entirely: "music is much closer to the sub-structure of space and time" than to language. Keep his objection in your pocket.)

The asymmetry that changes everything

So music is context-free-ish. Time for the crucial difference, because it drives the rest of the book. In language, the surface pins down the tree: word categories (noun, verb…) constrain composition so hard that a sentence usually admits one parse — which, incidentally, is a big part of why statistical NLP works so well. In music, notes carry no such rigid categories. Lerdahl and Jackendoff, whose theory dominates Chapter VI, put it flatly: almost any passage is "vastly ambiguous — much easier to construe in a multiplicity of ways," because music, tied to no fixed meanings, is "pure structure, to be played with within certain bounds."

So a music grammar can't just define legal trees; it must also say which legal tree a listener prefers. Hence the two-tier rule system you'll meet in Chapter VI: well-formedness rules (what trees are possible) and preference rules (which tree an experienced listener actually hears). Preference rules have no counterpart in linguistics — and their existence explains why reduction matters in music but not syntax: an underdetermined surface leaves room to rank notes by importance and strip the lesser ones, which is exactly what Schenker was doing by hand all along.

A final practical note: expectation itself has an algorithmic face. Chart parsers (Earley's algorithm) read left to right maintaining explicit predictions — "an NP has begun, so a noun is expected" — realized or revised word by word, top-down meets bottom-up. That is startlingly close to a description of listening; and where global grammar is overkill, plain N-gram statistics over chord transitions (with Viterbi decoding over hidden keys) do serious practical work, as in the key-finding example that closes the chapter. Grammar and statistics aren't rivals; they're the same enterprise at different depths.

Grammars generate sentences by rewriting rules; machines recognize them — and the machine's memory determines the language class. Human language needs a stack (nested, non-crossing dependencies); harmony plausibly does too, because cadences embed recursively — ii–V–I is grammar recursion you can hum. Real chord grammars exist (Winograd, Rohrmeier, PCFG, HPSG, CCG). But music's surface underdetermines its tree, so musical grammar needs preference rules on top of well-formedness — the conceptual key to GTTM.

Check yourself
What musical experience corresponds to the "push" of a push-down automaton?

Hearing a chord that creates an obligation — a dominant demanding resolution, a half cadence opening a question. The unresolved tension is an open parenthesis on your mental stack; the resolving tonic pops it. Chained dominants stack several brackets deep.

Why can't a finite automaton recognize aⁿbⁿ — and what's the musical analog?

Checking that the b's match the a's requires counting arbitrarily high, but a finite automaton's only memory is its current state — finitely many. The musical analog is nested expectation: each opened dominant must eventually be closed by its resolution, innermost first, which requires stack-like memory.

In what sense is ii–V–I "recursive"?

Start from the cadence V–I. The rule "any goal chord may be preceded by its own dominant" rewrites V as (II–V) — a cadence onto V nested inside the cadence onto I. Apply the rule again and you get the longer pre-dominant chains of jazz. One rule, applied to its own output: recursion.

Why does music theory need preference rules when linguistics doesn't?

Musical surfaces underdetermine their structure: notes lack the rigid categories words carry, so many well-formed trees fit one passage. Preference rules rank the candidates by how an experienced listener hears them. A sentence's surface usually contains enough information to force (nearly) one tree — so linguistics never needed the second tier.

Chapter IV

The Berklee method

Write G7 over a melody and any competent player on Earth realizes it — different voicing every night, same harmony every time. This chapter is about that compression: what it means to turn harmony into a symbol system, what such systems buy, and what they quietly delete.

Music's long war with notation

Music happens in a two-dimensional space — time × pitch — and the whole history of notation is trial and error at flattening that space onto paper. The staff itself probably began as a picture of an instrument: horizontal lines standing for strings, blobs marking when to pluck which one, horizontal spacing standing for time. Every era's notation is a decision about which information deserves symbols and which stays tacit — and the Berklee method, developed at the Boston school founded in 1945 (the first in America to teach jazz; alumni run from Quincy Jones to John Scofield to Dream Theater), is one of history's most consequential such decisions: a complete symbolization of functional harmony.

First decision: legislate consonance

Chapter II ranked intervals by frequency ratio — the "classical sense of harmony." Berklee makes a colder move: it declares dissonance levels by axiom — the "modern sense" — and builds every subsequent rule on the declaration:

LevelCharacterIntervals
1ConsonantM3, m3, M6, m6
2Neutral, stableunison, P8, P5
3InorganicP4
4Moderately dissonantM2, m7
5Acutely dissonantm2, M7
6Unstabletritone (+4)

Notice what axiomatizing does: it ends the argument. People genuinely differ about consonance (by ear, training, culture); a working method can't wait for consensus, so it stipulates and moves on. The book calls the result a "puzzle gamification" of harmony — which tensions are legal where becomes a rule-game, like chess: arbitrary at the foundations, deep in play.

Second decision: every chord is stacked thirds

The construction rule: take a root, stack thirds above it — root, 3rd, 5th makes the triad; add the 7th for a tetrad; the optional 9th, 11th, 13th are tensions. Do it on C with a dominant flavor: C, up a major third to E, minor third to G, minor third to B♭ — C7. Two abstractions are then bolted on, and both are choices, not facts:

  • Octaves don't matter. Only pitch class counts — the E can sit in any register. One symbol covers every voicing.
  • Inversions don't matter. G–C–E is still "C." Naming an inversion identically is an explicit theoretical claim: that it carries the same harmonic function, even though (as the book concedes) it audibly differs.

The chord name is then assembled like a little formula — root letter, quality suffix, tension list: G7, E♭M7(9,13), Am7♭5. And one absence is the masterstroke: the key is not written. G7 is spelled G7 whether it's functioning as V of C or as a passing chord in another key entirely. Determining its function is deliberately left to the player's analysis — which is precisely what leaves room for reharmonization (new chords under the same melody) and ad lib (new melody over the same chords). The symbol system's silence is where the creativity lives. The book's sharpest illustration: the famous Tristan chord has accumulated thirty-three published functional interpretations since 1879 — and exactly one Berklee name, Fø7. What the method refuses to say is what keeps the argument (and the music) alive.

Chord constructor · name ↔ sound

Assemble a chord the Berklee way and hear it. The name is generated by the same rules a chart uses.

Build Dm7, then G7 with the 9 and 13 tensions, then CM7 — and play them in that order: a fully-dressed ii–V–I. Watch how the name assembles mechanically from your choices; then flip a tension and hear how much color one symbol adds. Everything you just did is the "puzzle game": moves licensed by tables, sounding like jazz.

What can't the notation say? Voicings (octave data is abstracted away), chords not built in thirds (quartal harmony, clusters), deliberately omitted tones (except by ad-hoc marks like "omit 3"), microtones. When chord tones are missing in real music, assigning a name means guessing the absent notes — harmonic analysis smuggled in through the back door. Every symbol system pays for its power with blind spots; the method's genius was choosing blind spots working musicians could live with.

The three great symbolizations

Jazz musician Naruyoshi Kikuchi and critic Yoshio Otani argue that music has been successfully "symbolized" three times, and the book adopts their frame — reinterpreting each event, in computer-science terms, as virtualization: abstracting away physical detail to expose clean, manipulable function, the way a cloud service virtualizes racks of hardware into "compute."

≈ 1722 · 12-TET

Virtualizes pitch: any melody reproducible by any instrument in any key, ensembles trivially in tune with each other, foreign scales mappable onto the grid. The Well-Tempered Clavier is its manifesto.

≈ 1945 · Berklee method

Virtualizes harmony: a monophonic staff plus chord names replaces the full score. Composition separates cleanly from performance, and popular music becomes mass-producible.

1981 · MIDI 1.0

Virtualizes performance: key, velocity, and timing as digital events. Score and rendition both become computable; recording exactness becomes free; technopop, dance music, and the bedroom producer follow.

The common trade

Each innovation navigates expressiveness vs. description cost. Controlling every detail of every note is expressive and expensive; abstract instructions are cheap but need skilled interpretation. A lead sheet and an orchestral score sit at opposite ends of the same dial.

You live inside the third one. MIDI's whole design — note-on, note-off, velocity, and nothing else unless you add controllers — is one point chosen on that expressiveness/cost dial, and every frustration you've ever had with MIDI (where's the breath? the bow pressure?) is the cost side of the trade. It's also worth saying what symbolization did at civilizational scale: over 99% of what the world now hears is Western tonal music or its descendants, and the book attributes that hegemony largely to drastic symbolization — the system could cheaply absorb, imitate, and recombine everyone else's music. Indian classical music, by contrast, kept performance deliberately unsymbolized and tacit; it stayed profound, and stayed local. Symbolization is power, and power is never neutral.

Berklee ↔ bebop: a method and a music co-evolve

The method and bebop were born in the same city in the same decade, and they grew by feeding each other. Bebop was already the most "linguistic" jazz — compositional, rule-dense, low-ambiguity, improvisation unfolding like competitive sport within known constraints. The Berklee method was built partly to analyze it; bebop then absorbed the method and became more logical and more expressible; eventually it became the first jazz style deemed teachable at universities. But the same feedback loop has a failure mode the book names "Berklee sickness": symbolization pursued for differentiation's sake — late bebop and 1970s fusion piling up symbol density until pieces became nearly unplayable and, for audiences, incomprehensible. "To perform symbolization," Kikuchi and Otani write, "is a game with no end."

Miles Davis, twice

Against that saturation, put the musician the chapter keeps returning to. At a 1987 banquet Miles Davis told the politician seated next to him, "I've changed music five or six times." The book takes two of those changes as structural lessons:

  • Modal jazz (Kind of Blue, 1959). Before mode, propulsion came from chord change — melody threading changes every bar or two. Davis's move: hold one scale (mode) for long spans and let propulsion come from the color of melody against a static field — tension and release inside stillness. If you make ambient music, this is your direct ancestor: harmonic stasis as a canvas is the modal idea, slowed to geological tempo.
  • Fusion (Bitches Brew, 1969). The book is precise about what "fusion" means, using software versioning: jazz-played-rock is just old syntax with borrowed vocabulary. Real fusion is a new syntax — one that can express everything the old syntaxes could (backward compatible) while containing music the old ones cannot express at all (no forward compatibility). Bitches Brew, drawing on Hendrix and James Brown, was version 2.0: traditional jazz fans heard it as betrayal precisely because their v1.0 parser genuinely couldn't read it.

Classical music ran the same arc a half-century earlier — impressionism (Debussy, Ravel) as its modal turn away from functional harmony; the folk-fueled primitivism of Bartók and Stravinsky as its fusion — and the book proposes a general cycle: symbolization complexifies until saturation → a reaction returns to radical simplicity → the two are sublated into a new syntax at a higher level. Each genre's cycle seems to run faster than the last. Where the cycle is in your genre, this decade, is left as an exercise.

The Berklee method symbolizes harmony: dissonance by axiom, chords as stacked thirds named by root + quality + tensions, octaves and inversions abstracted away — and the key deliberately unwritten, leaving function to the player. It's the middle member of three great virtualizations (12-TET for pitch, chord names for harmony, MIDI for performance), each trading description cost against expressiveness. Symbol systems enable mass creativity, breed complexity sickness, and get periodically overthrown by simplifications that become new syntaxes.

Check yourself
Why does a Berklee chord symbol deliberately omit the key?

Leaving function un-notated hands interpretation to the musician: the same chart supports different keys, reharmonizations, and improvisation. Notating degrees (I, IV, V) would fix one interpretation; notating letter-roots keeps them all open. The Tristan chord: 33 functional interpretations, one Berklee name — Fø7.

What is "virtualization," and why is MIDI an example?

Abstracting physical detail to expose clean function — like cloud services virtualizing hardware into "compute." MIDI virtualizes performance: a keystroke becomes {note, velocity, time}, playable by any synth anywhere, at the cost of every nuance the event format doesn't encode. 12-TET does the same for pitch; chord names for harmony.

By the book's compatibility test, why is "jazz-rock" not fusion but Bitches Brew is?

Playing rock tunes in jazz syntax (or vice versa) is old syntax with imported vocabulary — v1.0 reading v1.0 files. Fusion proper is a new syntax: it can express the old repertoires, but produces music the old syntaxes cannot express — v2.0 files that crash the v1.0 parser. That asymmetry is the test.

What does the Berklee method structurally fail to express?

Voicings (octave data abstracted), non-tertian chords (quartal, clusters), deliberately omitted tones, microtones — and when chord tones are absent in the music, naming requires guessing them. Every symbolization buys power with blind spots.

Chapter V

The Implication–Realization model

You hear two notes. Before the third arrives, your ear has already placed a bet on it. Narmour's Implication–Realization model says exactly which bet — and locates musical emotion precisely where the bet is lost.

Meyer's founding move: emotion is denied expectation

Recall from Chapter I that Leonard Meyer put musical emotion on a scientific footing with one idea: as you listen you are continuously predicting, and emotion arises when prediction fails. A phrase that continues exactly as implied slides by almost unnoticed — no surprise, no affect. A phrase that swerves — later than expected, higher than expected, darker than expected — produces the little jolt of surprise, tension, or delight that we experience as the music "doing something." Confirmed predictions are invisible; denials are felt. (Meyer is careful, and so is this site: "emotion" here means these fleeting, intersubjective jolts — surprise, tension, release — not aesthetic judgment, which stays upstairs in the subjective layer.)

If you produce music, you already exploit this daily without the vocabulary: a drop that arrives a bar early, a chord swap on the repeat, the drum fill that withholds its final hit. All of it is manufactured denial of implication. And ambient music is a fascinating limit case: radically slowed harmonic rhythm stretches the implication window until expectation itself becomes diffuse — which is arguably why it relaxes. This chapter gives you the theory of the machine you've been operating.

Where do the predictions come from? Gestalt again

Prediction needs regularities, and the bottom-up regularities come from the gestalt principles of Chapter I, now applied along a melody: proximity (notes close in pitch/time bind together), similarity (like continues with like), and common direction (a line in motion tends to stay in motion). The book's example: hear C–D–E and the small, same-direction steps imply F next, or F–G–A — the process wants to continue. But interpose a larger interval — a perfect fourth up to A — and the chain breaks: the old process closes at E, and A starts a new unit. That breaking point is called closure; hold the term, it becomes load-bearing shortly. There are also top-down gestalt forces — "good form," "good continuation" — but note the asymmetry the book flags: proximity and direction can be objectively measured; "good form" can't. A computational model builds on the measurable ones.

Narmour's formalization

Eugene Narmour took Meyer's insight and did to it what this whole book wants done to musical intuitions: made it mechanical. Where Schenker and GTTM understand melody by grouping notes into hierarchies, the Implication–Realization model understands it as a network of references between neighboring notes — and Narmour was openly hostile to reduction, insisting that the abstract melodies of high-level analysis are things nobody actually hears. His model stays on the sounding surface. Its entire engine is two hypotheses about what any two heard events imply:

  • X + X → X — hear a thing twice, expect it a third time (sameness implies more sameness);
  • X + Y → Z — hear a change, expect further change (difference implies more difference).

Refined onto actual intervals, these become the model's two principles, keyed to one threshold: an interval is small up to 5 semitones, large from 7 up (the tritone, 6, sits ambiguously on the fence):

PRD · Principle of Registral Direction

A small interval implies continuation in the same direction. A large interval implies reversal: after a leap, the line should turn back.

PID · Principle of Intervallic Difference

A small interval implies a similar-sized interval next (within ±2 semitones). A large interval implies a smaller one — fill the gap the leap opened.

Work the classic examples by hand. C–D: small, upward — so PRD says "keep going up," PID says "by about a step": prediction E. If E arrives, implication realized, nothing felt. If C arrives instead (C–D–C): direction denied — a flicker of the unexpected, though such returns are so common they're their own pattern. If F arrives (C–D–F): direction confirmed but the interval widened — PID denied, another flicker. If E♭ arrives: right direction, right size, wrong scale — denial at the learned level. Each denial type is a distinct, nameable event, and by Meyer's principle each is a little generator of affect, whether or not you consciously notice.

The pattern alphabet

Classify every three-note move by (1) was the first interval small or large, (2) did direction continue or reverse, (3) was the second interval similar, same, or different — and each viable combination earns a name. The core alphabet: D (duplication — repetition confirmed), P (process — small interval continuing as implied), R (reversal — leap answered by a turn back, both principles satisfied), the exact return aba/ID, and the partial-denial forms IP, IR, VR where one principle is honored and the other refused. Inverted (upside-down) versions count as the same pattern.

Pattern gallery · the basic I-R shapes

Each basic pattern is a three-note shape classified by what the third note does to the implication of the first two. One typical realization of each (patterns also occur inverted):

pick a pattern to hear it and see which principles it satisfies
Play P, then R, then IR, a few times each. P should feel like nothing — that's correct, it's the invisible default (and the most common pattern in real melody). R feels shaped, almost narrative: leap out, settle back. IR has a small thrill — the leap that keeps going against gravity. You're calibrating your ear to feel implications as forces, which is the whole point of the model.

Two structural details complete the machinery. Closure answers "where do patterns start and end?": at any point where implication weakens — a long note, a rest, a strong beat, a change of direction, a dissonance resolving. Closure both terminates one pattern and opens the next, so the parsing is self-organizing. And patterns overlap: a single note can be simultaneously the last note of one pattern and the first of another, so a melody is a chain-mail of linked implications, not a string of beads. Narmour's own analyses (the book reproduces his reading of the Pastoral Symphony's fifth movement) annotate melody, duration, and harmony this way at once — and note how strongly this contrasts with a tree: I-R yields a network, dense with local links, agnostic about deep hierarchy. Keep both pictures; Chapter VI takes the other branch.

Narmour's chess metaphor is worth keeping too: a semitone is a pawn — you always know roughly where it's going; an augmented sixth is a queen — free in principle, yet in any actual position her options collapse to a few. Melodic analysis, he says, is reading the game: which pieces moved, and why those moves won.

The test: Carlsen's experiment

A model of expectation makes a checkable claim: give people two notes, ask them to continue, and the continuations should distribute as the model predicts. James Carlsen ran exactly this in 1981 with 91 first-year music students in Hungary, Germany, and the USA. Protocol: a metronome at 60; two stimulus notes; the subject sings whatever continuation feels natural for seven beats (deliberately not "make a good melody" — that instruction, piloted earlier, just produced cadence clichés). The first sung note defines the response interval. Tabulate everything and you get, for each stimulus interval from −12 to +12 semitones, the empirical distribution of what human ears expect.

Expectation lab · bet against 91 music students

Choose an opening interval. The model states its implication; press play to hear the two stimulus notes and then the model's predicted continuation. Then reveal what Carlsen's subjects (91 students, Hungary/Germany/USA, 1981) actually sang — the top three responses and how often.

Set the stimulus to +2 (up a major second), read the model's bet, play it — then reveal the data: 64.3% of subjects continued up a step. The model's best case. Now set +8 (a minor sixth up): the top responses all reverse direction — R, as PRD demands. Then probe the edges: at −10, why is the top answer a further descent of 2? At ±6 (the tritone), what do people do with an interval the model can't classify? The mismatches are the most instructive part.

The verdict: strong but not clean. Small intervals overwhelmingly beget small same-direction intervals (P); large leaps overwhelmingly reverse (R); after an octave leap every top response turns back. But systematic deviations appear exactly where learned tonality — what the book calls the scale-step — rides on top of raw gestalt: descending sevenths behave like leading tones resolving upward to complete the octave; tritones resolve by semitone as tonal voice-leading trains us to; and one asymmetry (up-a-major-third continues, down-a-major-third turns) has no bottom-up explanation at all. So the picture is two-layered, and honestly so: inherited, unconscious gestalt principles underneath; acquired, style-specific expectation on top. Carlsen's subjects' responses "markedly reflected the styles and culture they had experienced" — while register and proficiency barely mattered.

The other half of Narmour's ambition

One more idea deserves rescue from the footnotes: Narmour held that music theory has two legitimate aims. One is scientific — universal laws of hearing, the PRD/PID kind of thing, testable à la Carlsen. The other is the opposite of universal: characterizing what makes a particular artist's structures theirs. He coined idiostructure for that fingerprint — the recurring structural choices a musician discovers in themselves and then cultivates work after work. The methodology for finding idiostructure can be scientific even though the thing found is irreducibly individual. It's a generous frame for thinking about your own catalog: your idiostructure is whatever survives across your tracks when genre, tempo, and gear all change.

Listening is continuous prediction; emotion lives in prediction's failures (Meyer). Narmour mechanized the predictions: small intervals imply same-direction, similar-size continuation; large intervals imply reversal and gap-fill (PRD + PID). Three-note realizations and denials form a pattern alphabet (P, D, R, ID, IP, IR, VR) chained between points of closure into a network — not a tree. Carlsen's sung-continuation data largely confirm the bottom-up principles, with deviations that map the learned tonal layer sitting on top.

Check yourself
Why does the model predict emotion at C–D–F but not at C–D–E?

C–D (small, up) implies same-direction, similar-interval continuation: E. E realizes the implication — invisible. F keeps the direction but widens the interval to a minor third: PID denied, and denials are where affect is generated, conscious or not.

What is closure, and why does the model need it?

Closure is any point where implication weakens — long note, rest, accented beat, direction change, resolving dissonance. It answers "where do three-note patterns begin and end?": each closure terminates one pattern and starts the next, letting the melody parse itself into overlapping implicative units.

Which Carlsen results show learning riding on top of gestalt?

Descending sevenths answered by upward semitones (heard as leading tones completing an octave), tritones resolving by semitone (tonal voice-leading), and the up/down asymmetry after major thirds. Raw PRD/PID can't produce these; the subjects' tonal enculturation — the scale-step — can.

How does I-R differ from Schenker/GTTM in what counts as "the melody"?

Reduction theories posit abstract melodies made of chronologically distant structural notes. Narmour objects that nobody hears those; I-R stays on the sounding surface and models the network of expectations between locally adjacent notes. Sets-and-hierarchies vs. relations-and-networks — the two mathematical styles of understanding, both applied to the same tunes.

Chapter VI

GTTM & Tonal Pitch Space

The Generative Theory of Tonal Music is Schenker's intuition rebuilt as an explicit system: four cascaded analyses that turn a score into a tree with the piece's most important event at the root. Tonal Pitch Space then supplies the one thing GTTM lacks — a way to measure harmony.

What Lerdahl & Jackendoff were trying to do

In 1983 a composer (Fred Lerdahl) and a Chomskyan linguist (Ray Jackendoff) published A Generative Theory of Tonal Music — the most serious attempt ever made to state, formally, what an experienced listener understands when they understand a piece. "Generative" is borrowed from linguistics but redirected: the grammar doesn't generate pieces, it generates structural descriptions of pieces — the tree you unconsciously build while listening. And because music's surface underdetermines its tree (Chapter III's asymmetry), every module of the theory splits its rules in two: well-formedness rules (what structures are legal) and preference rules (which legal structure an experienced listener actually hears). That second tier is GTTM's signature invention.

The theory is a pipeline of four analyses, each feeding the next:

  1. Grouping analysis — carve the surface into motives, phrases, sections.
  2. Metrical analysis — find the grid of strong and weak beats.
  3. Time-span reduction — combine 1 + 2 to rank every note's structural importance in a tree.
  4. Prolongational reduction — re-derive the tree as a story of tension and relaxation.

Grouping: where do phrases begin and end?

The well-formedness side is pure common sense made explicit: a group is a contiguous stretch of notes; the whole piece is a group; groups nest inside larger groups; and they may never partially overlap. If that sounds like regions in a DAW arrangement — clips inside sections inside the song, never half-inside — that's exactly the right picture, and it's also just Chapter I's layer discipline again.

The interesting rules are the preferences, because they name the actual perceptual triggers of segmentation. GPR2 says boundaries fall at gaps — a rest or slur-end (2a), or a longer-than-neighbors inter-onset interval (2b). GPR3 says boundaries fall at changes — of register, dynamics, articulation, or note length. GPR4: where those effects are strongest, put the boundaries of larger groups. GPR1 pushes back against degenerate solutions (avoid single-note groups); GPR5 prefers symmetric halves; GPR6 says parallel material should get parallel structure — if the phrase repeats, the analysis should repeat. Individually banal, collectively sharp: they routinely disagree, and the disagreements are the hard part (Chapter VIII lives there). GPR7 even refers forward to analyses that haven't run yet — prefer groupings that will make the later reductions stable — making the whole system quietly circular, which is one reason mechanizing it took twenty-five years.

Meter: the grid under the groups

The metrical analysis recovers the hierarchy of beats — the thing notated as a dot-grid: every eighth gets a dot, every quarter two, every half three, and so on upward; more dots = metrically stronger. Well-formedness: strong beats come every 2 or 3, evenly spaced per level, each level's beats belonging to the level below. Preferences: align strong beats with note onsets (an onset without a beat is syncopation — notated with the asterisk in the book's figures), with stresses, with long notes, with parallel structure, and put the strong beat early in a group. You've internalized all of this from years of programming drums; the only news is seeing it written as a rule system — and noticing that meter and grouping are independent structures that the next stage must reconcile.

Time-span reduction: the tournament

Now the centerpiece. Intersect grouping and meter and you get time-spans — the natural segments of the piece at every scale. Within each span, adjacent events face off, and one wins: it becomes the head of the span, the event the span is "about." Winners advance to face the winners of neighboring spans, tournament-style, until one event heads the entire piece. Draw every match and you have the time-span tree — a full ranking of every note's structural importance, with Schenker's skeleton now derivable by machine-checkable steps: lop off the tree's lower branches and the surviving notes, stretched over the spans they head, are the reduction at that depth. That's the strong reduction hypothesis: listeners really do maintain something like this simplified skeleton as they listen.

Who wins a match? Preference rules again: prefer the event on the stronger beat (TSRPR1); the more consonant event, or the one closer to the local tonic (TSRPR2); respect parallelism (4); prefer heads that make the meter and the coming prolongational analysis stable (5, 6); and treat phrase edges specially (8, 9). One special mechanism deserves its own sentence: cadential retention. A cadence V–I works as a unit — the book's image is a thrown ball and its catch: the phrase's structural beginning [b] throws, the cadence [c] catches — so the tree is allowed to glue V–I together and carry the pair upward as one super-event rather than forcing an artificial winner between them. The book's worked example is the opening of Mozart's K.331: get the [b]/[c] correspondence wrong and the fourth bar's half cadence parses absurdly; get it right and the tree matches what every musician feels.

Time-span reduction · hear the hierarchy

One plausible time-span analysis of the first phrase of Ah, vous dirai-je, maman ("Twinkle, Twinkle"). Slide the depth control: each level strips away less-salient events, and the survivors stretch to fill the vacated time-spans. Level 1 is the phrase's single head — Schenker's skeleton, made audible.

Play level 4, then 3, then 2, then 1 — listening for what survives each cut. Level 3 removes the repeated-note echoes; level 2 keeps one head per half-phrase (C·G·E·C — the tonic triad, unrolled); level 1 is the phrase's single head. This is Schenker's foreground → middleground → background as an audible slider. Also worth noticing: each survivor stretches to fill its span — reduction isn't deletion, it's promotion.

Prolongational reduction: the tension tree

The time-span tree ranks events by rhythmic-structural salience. But your experience of harmony is a drama of departure and return — tension out, relaxation home — and that drama ignores barlines (a chord's influence prolongs past them; Schenker again). So GTTM derives a second tree over the same events, re-attached top-down by harmonic logic: a right-branch means "this event departs from its parent" — tension; a left-branch means "this event arrives back" — relaxation. Junctions come in three strengths: strong prolongation (same chord, same bass, same melody — pure persistence), weak prolongation (same root, new dress), and progression (genuinely new harmony). Two shapes then define tonal normalcy: the basic form — GTTM's Ursatz — is the piece's head tonic with a cadential dominant-and-tonic hanging left beneath it; and the normative structure demands tension then release: pieces that only relax, only tense, or relax before tensing all sound wrong. Even the interruption form I…V ∥ I…V–I (antecedent, consequent — the oldest phrase pattern in the repertoire) falls out as a specific branching. If you want one sentence: the time-span tree is what the piece is made of; the prolongational tree is what the piece feels like.

The rule system, condensed

For reference — and to appreciate exactly what Chapter VIII signed up to implement — the full apparatus:

Grouping rules — GWFR 1–5 · GPR 1–7

Well-formedness: a group is a contiguous note sequence; the whole piece is a group; groups may nest; a larger group containing part of a smaller one must contain all of it; a partitioned group is partitioned exhaustively. Consequence: no overlaps, clean hierarchy.

GPR1 avoid single-note groups · GPR2 boundaries at slur/rest gaps (2a) and long inter-onset intervals (2b) · GPR3 boundaries at changes of register, dynamics, articulation, duration · GPR4 where GPR2–3 effects are strongest, place larger-level boundaries · GPR5 prefer symmetric halves · GPR6 parallel passages get parallel structure · GPR7 prefer groupings that make later reductions stable (a meta-rule referencing analyses not yet run).

Metrical rules — MWFR 1–4 · MPR 1–10

Well-formedness: every attack sits on a smallest-level beat; each level's beats are beats at the level below; strong beats come every 2 or 3; equally spaced per level.

MPR1 parallel groups, parallel meter · MPR2 strong beat early in the group · MPR3 beats on onsets (an onset without a beat = syncopation) · MPR4 stress marks want strong beats · MPR5 long events (duration, slur, harmony) want strong beats · MPR6–10 meta-rules: stable bass, stable cadence, suspensions on strong beats, time-span optimality, binary regularity.

Time-span reduction — TSRWFR 1–4 · TSRPR 1–9

Well-formedness: every span has a head; leaves are pitch events; heads come from subtrees — with three licensed exceptions (fusion of arpeggios, transformation via hypothetical chords, and cadential retention: V–I glued into one unit).

Preferences: 1 strong beat · 2 consonant / near local tonic · 3 registral extremes (weakly) · 4 parallelism · 5 metrical stability · 6 prolongational stability · 7 cadential retention · 8 structural beginning [b] early in span · 9 prefer the cadence [c] over [b] as overall head.

Prolongational reduction — PRWFR 1–4 · PRPR 1–6 · stability

Well-formedness: one head per prolongation; each event elaborates another via strong prolongation (same root, bass, melody), weak prolongation (same root), or progression; elaborations are recursive; branches never cross.

Preferences: 1 prolongational importance follows time-span importance (relaxed by the interaction principle: choose from the top two time-span levels) · 2–3 attach the chosen event where the connection is most stable · 4 attach to the more important endpoint · 5 parallelism · 6 the normative structure: a tonic opening, tension branch, cadential preparation, cadence. Stability favors: right-branching prolongations / left-branching progressions, shared diatonic collections, small melodic intervals, short circle-of-fifths moves.

Tonal Pitch Space: giving GTTM a ruler

Look back at the preference rules and notice how often they lean on words like "consonant," "stable," "cadence" — harmonic judgments the theory never defines. GTTM assumes a sense of harmonic distance it doesn't supply. Lerdahl's Tonal Pitch Space (2001) is the missing instrument: a geometry in which every chord and key has a position, and harmonic intuition becomes distance.

The atom is the basic space of a chord: five nested levels, from most to least privileged. For C major's I chord — (a) octave: C alone; (b) fifth: C, G; (c) triad: C, E, G; (d) diatonic: the seven scale notes; (e) chromatic: all twelve. A pitch class is more central the higher the level it first appears on: C lives at every level, G at four, E at three, D at two, C♯ at one. This one diagram already encodes an enormous amount of tonal common sense — why the fifth is stable, why non-chord tones feel decorative, why chromatic notes feel foreign.

Distance between chords x and y in a key is then defined as δ(x→y) = j + k: j = steps between their roots around the chordal circle of fifths (I→V→ii→vi→iii→vii°→IV), and k = how many level-climbs pitch classes must make to convert x's basic space into y's. Work I→V by hand: j = 1 (one step on the circle). For k, compare spaces — G must rise from "fifth" to "root" (+1), D from "diatonic" to "fifth" (+2), B from "diatonic" to "triad" (+1): k = 4. So δ(I→V) = 5. Run all seven degrees from I and you get the key's harmonic map: V and IV nearest at 5, iii and vi at 7, ii and vii° at 8 — matching, number for number, which chords musicians treat as close substitutes. (One subtlety our calculator honors: vii° is diminished, so its "fifth level" holds its actual diminished fifth, not a perfect one.)

TPS calculator · δ(x→y) within C major

Pick two diatonic chords. The bench computes j (circle-of-fifths steps I→V→ii→vi→iii→vii°→IV), builds each chord's basic space, counts the level-climbs k, and plays the progression. Compare I→V (δ=5, the closest real move) with I→ii (δ=8, a distant neighbor a step away).

Compute I→V and read the k-breakdown against the paragraph above. Then I→ii — a chord one scale-step away, yet δ = 8, among the farthest in the key. Distance in TPS is functional, not keyboard-spatial: what matters is shared structure, not adjacency. This is also your first quantitative reharmonization tool: substitutes with small δ between them (V and vii°, I and vi) are the classic swap pairs.

The construction then telescopes outward. Add a term i for steps between keys on the circle of fifths and δ measures across modulations; keys arrange themselves into a regional space — a torus where moving one way changes key by fifths and the other way alternates relative/parallel keys; distant keys are reached through pivot regions at a fixed toll. And over this geometry Lerdahl states a physical-sounding law: of all analyses of a progression, prefer the one tracing the shortest path — harmony follows geodesics, like light through media. An ambiguous chord is heard as whatever interpretation minimizes total distance travelled.

Tension and attraction: the payoff

With distance defined, GTTM's hand-waves become equations. Sequential tension: the tension of arriving at chord y is the distance travelled to reach it, δ(xprev→y) (plus a term for the chord's own dissonance — inversions, added notes). Global tension: inherited down the prolongational tree — each event's tension is its local distance from its parent plus everything its ancestors accumulated, which is why the same chord can feel calm early and loaded late. And melody gets its own force law: melodic attraction, α(p₁→p₂) = (s₂/s₁)·1/n² — each tone is pulled toward stronger anchors (s from the basic space), with the pull falling off as the square of the distance in semitones. An inverse-square law for voice leading: B craves C (α = 2, the strongest pull in the key), F leans onto E, and stable tones barely move. Lerdahl closes by conceding the arithmetic is deliberately naïve — δ = i + j + k could carry fitted weights w₁i + w₂j + w₃k — locating exactly where hand-built theory shades into machine learning.

Tension plotter · distance travelled = tension felt

TPS's central claim, made audible: sequential tension is the δ distance from the previous chord. Build a progression (or load a preset), then play it — each chord lights up with the distance it cost to reach.

the red bars are δ from the previous chord — watch V→I: high arrival tension at V, cheap release into I
Load I–vi–ii–V–I and play it: watch the δ bars swell toward V and collapse into the final I — tension as literal distance, release as coming home cheaply. Then build a progression from your own music and see whether its felt arc matches its δ profile.
Melodic attraction · α(p₁→p₂) = (s₂/s₁) · 1/n²

Melody has gravity too. Each scale note has an anchoring strength s from the basic space (C=4, E=3, G=4; other diatonic notes 2); attraction to a neighbor scales with the strength ratio and falls with the square of the distance in semitones — leading-tone B pulls to C with α = 2, the strongest attraction in the key.

Rank B's attractors, then F's, then G's. B and F — the two notes of the tritone — are the needy ones, each desperate for its semitone neighbor; G barely cares. Now you know why the V7 chord (which contains both B and F) resolves so inevitably: it's holding the key's two most attracted tones at once.

GTTM: group the surface (gap and change rules), find the beat grid, then let events tournament for headship within each time-span — producing a tree whose prunings are Schenker's reductions — and re-derive the tree as tension (right branches) and relaxation (left branches). Well-formedness defines the legal; preference rules pick what listeners hear. TPS supplies the missing ruler: chords as five-level basic spaces, distance δ = circle-of-fifths steps + level-climbs, keys on a torus, analyses choosing shortest paths — from which tension curves and an inverse-square law of melodic attraction follow.

Check yourself
What's the difference between a time-span tree and a prolongational tree?

The time-span tree encodes rhythmic-structural salience, built bottom-up from grouping and meter. The prolongational tree encodes harmonic tension and relaxation, built top-down by re-attaching events — its branches ignore surface barlines because prolongation does. Same events, two hierarchies: what the piece is made of vs. what it feels like.

Why does GTTM need cadential retention?

A cadence V–I functions as one gesture — the "catch" [c] answering the phrase-opening "throw" [b]. Forcing a winner between V and I would misrepresent that; retention glues them into one super-event carried up the tree, which is what makes the K.331 half-cadence parse correctly.

Walk δ(I→V) = 5.

j = 1: V is one step from I on the chordal circle of fifths. k = 4: rebuilding the basic space, G climbs fifth→root (+1), D climbs diatonic→fifth (+2), B climbs diatonic→triad (+1). δ = j + k = 5 — tied with IV for the nearest chord in the key.

Why does TPS complete GTTM rather than compete with it?

GTTM's preference rules appeal repeatedly to "stability," "consonance," and "cadence" without defining chordal proximity. TPS supplies the metric: δ identifies cadences, arbitrates head choices, and quantifies the tension/relaxation the prolongational tree depicts — the ruler for the drawing GTTM already made.

Why is the melodic-attraction law written like gravity?

α(p₁→p₂) = (s₂/s₁)·1/n²: pull grows with the target's anchoring strength and dies off with the square of the semitone distance — nearby strong anchors dominate. It explains leading-tone resolution (B→C: α=2) and why stable tones (C, G) feel no urge to move at all.

Chapter VII

An algebra of trees

If a time-span tree really captures a melody's structure, then operations on trees are operations on music. This chapter builds the arithmetic — a distance, a union, an intersection — and checks it against human ears.

The dream: programming music

Why formalize further? Because of a gap the book calls putting intention into music. Every existing way of making a computer help you compose sits somewhere awkward on Chapter I's flexibility–usability tradeoff: an editor (note-by-note entry — full control, maximal labor); case-based methods and machine learning ("make it like this example" — cheap, but control is indirect); generate-and-test and genetic algorithms (can't specify what you want, only recognize it when it appears). What's missing is the way a programmer works: a small set of primitive operations whose behavior you understand so well that you can compose them, predict the result, and build arbitrarily far. All computation reduces to a few arithmetic operations; all compass-and-straightedge geometry to two tools. Is there an equivalent basis set for melody? Donald Norman's design principle states the target: the complexity of the tool should be the complexity of the task — no more. This chapter proposes the basis set: reduction, meet, and join over time-span trees.

Analysis and rendering: trees aren't scores

First, a housekeeping fact with consequences. Analysis (score → tree) is one-to-many: a melody genuinely admits several trees — slur it differently, interpret it differently, and different analyses result; identical notes at the start vs. the end of a piece can mean different things. Rendering (tree → score) is also one-to-many: a tree's leaves aren't literal notes with onsets and durations — and trees don't even contain rests — so realizing a tree as a playable score involves genuine choices. Keep this asymmetry in mind: the algebra below operates on trees, and a final rendering step (with freedom in it) turns results back into sound. It will matter in Chapter IX.

The unit of information: maximum time-span

To do arithmetic we need a measure — some way to say how much of a melody a note carries. The proposal: an event's maximum time-span is the longest span it ever heads in the tournament. The phrase's head owns the whole phrase's duration; a passing tone owns only its own little slot. Concretely, with a quarter note = 12 ticks: take four quarter-note events where e1 beats e2, e3 survives locally, and e4 heads everything. Then mt(e2) = its own 12 ticks; mt(e1) = its 12 plus e2's span = 24; mt(e4) = the whole 48. The general law: a head's maximum time-span is the concatenation of all the time-spans below it. Now the maximum time-span hypothesis: prune an event and you lose exactly mt-worth of information. Important notes literally carry more of the piece — the intuition Schenker never quantified, quantified.

Reduction paths, orderings, lattices

Armed with the measure, reduction becomes precise. You may prune only leaves — events heading no surviving subordinate (you can't remove a parent and leave its ornaments orphaned). Prune one leaf at a time and you trace a reduction path from the full melody down toward its skeleton; at every step the pruned tree stands in a subsumption relation to its ancestor (σ′ ⊑ σ: everything structural in σ′ is in σ). Since at most steps several leaves are prunable, many paths descend from the same melody — and the set of all reductions of all melodies forms a partially ordered set: some trees are comparable (one reduces to the other), most simply aren't.

Mathematics has a name for the nice case of such structures: a lattice, a poset where every pair has a greatest common lower bound and least common upper bound. Translated: for two melodies σA and σB,

  • meet(σA, σB) — the largest tree both reduce to: their shared skeleton. Think: what two variations of one theme have in common is (roughly) the theme.
  • join(σA, σB) — the smallest tree containing both: their structural union, a melody carrying the content of the two at once.

One asymmetry to remember: meet always exists (worst case, the empty tree ⊥), but join is partial — it exists only when the two trees' time-spans are compatible enough to overlap. You can always ask what two melodies share; you cannot always merge them. The expected algebraic identities (absorption: (σA ⊔ σB) ⊓ σA = σA, and friends) all hold — this is a real algebra, not a metaphor.

Distance — and the parallelogram

Now the metric. Define the distance from σA to a reduction σB as the sum of the maximum time-spans of everything pruned along the way. First non-obvious theorem: the sum doesn't depend on pruning order — every reduction path between the same two trees costs the same, so d is well-defined. Extend to any two melodies by routing through the lattice: go up to the join and down, or down to the meet and up — and the second theorem (the distance lemma) says both routes agree. The triangle inequality holds too: d is a true metric, and melody-space becomes a geometry.

join(σA, σB) — union · 822 ticks σA · Variation No. 2 (744) σB · Variation No. 5 (654) meet(σA, σB) — common skeleton · 576 ticks 168 78 78 168
Mozart, Twelve Variations on "Ah, vous dirai-je, maman" K.265 — variations No. 2 and No. 5. Numbers are total maximum time-spans in ticks (♩ = 12); edge labels are reduction distances. Both routes between σA and σB measure 246: the distance lemma, drawn.

Read the picture with real numbers. Mozart's Variations on "Ah, vous dirai-je, maman" (you know the theme as Twinkle, Twinkle): variation No. 2's tree totals 744 ticks of maximum time-span, No. 5's totals 654. Their meet — the shared skeleton, which is essentially the Twinkle theme itself — totals 576; their join, 822. Distance No. 2 → No. 5 via the meet: (744−576) + (654−576) = 168 + 78 = 246. Via the join: (822−744) + (822−654) = 78 + 168 = 246. Opposite sides equal: the four trees literally form a parallelogram in melody-space. When an abstraction hands you Euclid-like structure for free, you're usually onto something.

Does the geometry match the ear?

Time to cash the cognitive-reality check from Chapter I. The authors took the Twinkle theme plus twelve variations (arranged to monophony), built and cross-checked their trees, and computed all pairwise tree distances. Then, separately: eleven listeners heard every pair of melodies and rated similarity on a five-point scale — both orders, to cancel order effects — yielding a purely psychological distance matrix. Both matrices were flattened to 2-D maps with multidimensional scaling (MDS: position items so map distances best reproduce matrix distances) and compared.

Result: the maps cluster alike. Theme with variations 5, 8, 9 in one cluster; 3 with 4; 2 with 7; matching neighborhood relations almost throughout (variation 6 the honest outlier). Now the remarkable part: the tree distance ignores pitch entirely — it sees only structural shape, essentially rhythm-in-hierarchy — yet it reproduces much of human similarity judgment. (Partly because pitch sneaks in: pitch influenced which tree was built in the first place.) And the mismatches were all pitch-shaped — a variation in the minor feels distant even when its tree is near — which points exactly at the announced next step: fold TPS's pitch distances (Chapter VI) into the tree metric.

Reduction game · prune the tree yourself

The Twinkle phrase from Chapter VI, now under your knife. Click notes to prune them — but the algebra only allows pruning a note that heads no surviving subordinate (a leaf of the current tree). Each cut costs its maximum time-span in beats; the running total is the reduction distance d from the surface. Play the remains at any point: survivors stretch into the vacated spans.

Prune your way from 14 notes to the single head, watching d tick up by exactly each pruned note's maximum time-span. Try illegal moves — the game refuses parents with living dependents, which is the well-formedness of reduction. Then reset and take a different pruning order to the skeleton: same total, 30 — the path-independence theorem, verified by your own clicks. Play at each stage; around d ≈ 14 you can still hear "Twinkle" perfectly, which is the reduction hypothesis working.

Measure a note by the longest span it heads (maximum time-span); then reduction is leaf-pruning with a well-defined cost, melodies form a lattice under subsumption with meet (shared skeleton) and join (structural union), and distance — path-independent, triangle-obeying — makes melody-space a geometry where two Mozart variations and their meet and join form a literal parallelogram. Human similarity ratings track this geometry surprisingly well given that it ignores pitch — and its announced missing ingredient is exactly TPS.

Check yourself
Why must reduction distance be independent of pruning order to count as a distance?

A metric needs one value per pair of points. If different reduction paths between the same trees summed differently, "distance" would depend on the route. The maximum time-span hypothesis guarantees each pruned event costs the same whenever it's pruned — so all paths agree, and d is well-defined.

What do meet and join mean musically, and why is join partial?

Meet: the largest common reduction — what two melodies share (two Mozart variations meet in the Twinkle skeleton). Join: the smallest tree containing both — the two at once. Join requires the trees' time-spans to be compatible (overlapping); incompatible melodies simply have no union, though every pair has a meet.

What did the MDS experiment show, and what was the telling failure?

Tree distance — computed from structure alone, no pitch — reproduced the clusters of human similarity judgments among Mozart's variations. The failures were all pitch-driven (e.g. minor-mode variations feeling distant despite similar trees), pointing to TPS pitch distance as the missing term.

Chapter VIII

Implementing GTTM: exGTTM & ATTA

Lerdahl & Jackendoff said it themselves: "Our theory cannot provide a computable procedure." This chapter is the twenty-year engineering campaign — by this book's own authors — to build one anyway, and what its successes and failures teach about formalizing intuition.

Two kinds of ambiguity — only one of them a bug

Before the obstacles, one clarification the whole project rests on. Musical analysis is ambiguous in two totally different ways. First: music itself is genuinely multivalent — "almost any passage is potentially vastly ambiguous," and a passage really can support two hearings (you'll meet a Mozart bar with two legitimate groupings below). That ambiguity is a feature of the domain; a faithful system must be able to produce every human-valid answer, not one canonical answer. Second: the theory's prose is ambiguous — concepts needed for computation left undefined or implicit. That one is a bug, and the entire method of this chapter is to hunt it down case by case. Success means: all of the first kind of ambiguity preserved, none of the second kind left.

Why GTTM wouldn't run

  • Ambiguous rules. GPR4 places larger boundaries where GPR2–3 effects are "relatively more pronounced." Pronounced how much? Compared across rules how? Normalized how? Prose survives these questions; code doesn't.
  • Conflicting rules with no referee. The book's own figure: a melody where GPR3a (a leap) wants a boundary after note 3 while GPR6 (parallelism) wants one after note 4 — and accepting both would create a one-note group, violating GPR1. Something must decide, and GTTM is silent.
  • Missing algorithms. The rules describe properties of good analyses, never procedures for finding them. GPR6 says parallel segments prefer parallel structure — but detecting parallelism (how similar? pitch-similar or rhythm-similar? weighted where?) is itself a hard, unspecified subproblem. Worse, local rules (GPR2–3) build hierarchy bottom-up while global rules (GPR5–6) shape it top-down, and nothing says how to reconcile the two directions.

For contrast, the chapter reviews the one serious precedent: David Temperley's preference-rule systems and the Sleator–Temperley Melisma analyzer, which made preference rules computable via dynamic programming — score every candidate analysis numerically, search for the global optimum, prune the exponential candidate space as you go. It works, and it's principled. But its rule weights are fixed for all pieces, resolving every conflict the same way — and the authors' hunch, borne out below, is that proper weights vary from piece to piece. Earlier attempts got less far: Stammen implemented only basic grouping without hierarchy; Nord applied rules by hand; segmentation research found local boundaries but no hierarchical structure of the GTTM kind.

exGTTM: externalize everything, parameterize everything

The authors' strategy has a bureaucratic name and a radical content: full externalization and parameterization. Every vague quantity in the theory becomes an explicit, adjustable parameter:

  • Each preference rule R gets a strength SR ∈ [0,1] — turning rule conflicts from contradictions into weighted votes, and every possible priority ordering into a point in parameter space.
  • Vague thresholds ("salient enough") become explicit T parameters; hidden modeling choices (does parallelism weigh pitch or rhythm? does a parallel segment's boundary attach to its start or end?) become W parameters — including parameters the original theory never even recognized it was assuming.
  • Every intermediate quantity is normalized to [0,1], so evidence from any rule can be combined with any other in cascaded weighted means — a clean, composable numeric architecture.

Grand total: 15 parameters for grouping, 18 for meter, 13 for time-span reduction. The working loop that produced them: whenever a human-correct analysis existed that the system couldn't generate, find the implicit assumption blocking it, externalize that assumption as a new parameter, repeat — until every correct analysis in the corpus was reachable at some parameter setting. If 46 knobs sounds inelegant, note what they buy: with parameters explicit and fixed, the system's output is objective and reproducible, which makes the theory testable — Temperley's own criterion — and it makes the domain's real ambiguity representable (different settings = different valid hearings). Constructing the hierarchy itself is then posed as a constraint-satisfaction problem, taking global and local information into account simultaneously rather than Melisma's middle-out sweep.

Equally deliberate is what they didn't build — "paradoxically speaking, what we have not implemented enables us to implement GTTM": monophony only; no harmony (so no key or chord analysis); ordinary heads only; no feedback loops between analyses; and no prolongational reduction at all, that theory being judged still unsettled (and needing TPS-style chord knowledge GTTM lacks). Scope discipline is the difference between a running system and a manifesto.

Mini-ATTA · a boundary detector with three knobs

A working, deliberately tiny slice of exGTTM: three grouping-rule strengths S you can tune, exactly as the real system's 15 parameters work. For every gap between notes of Ode to Joy it scores GPR2b (a longer inter-onset interval than its neighbors), GPR3a (a bigger leap), and GPR3d (a change of duration), takes the weighted mean, and calls anything over the threshold a boundary. Defaults find the phrase boundary from rhythm alone — now try muting S2b: in an all-stepwise melody, the pitch rule has nothing to say, and the boundaries vanish. Parameterization in miniature.

The real exGTTM adds the rest: GPR1's small-group penalty, GPR4's promotion of strong boundaries to higher levels, GPR5 symmetry, GPR6 parallelism (which would also mark the phrase repeat here), then builds the full hierarchy as a constraint-satisfaction problem.

This bench is a genuine, if tiny, exGTTM: three rule strengths and a threshold — four of the real system's forty-six knobs. Defaults find Ode to Joy's phrase boundary from rhythm alone (GPR2b, the attack-gap rule, does the work — note how the leap rule stays silent in an all-stepwise melody). Drag S₂ᵦ to zero and watch segmentation collapse; push the threshold down and watch spurious boundaries sprout. You are experiencing, hands-on, both why parameterization is powerful and why "just tune it per piece" is an unfinished answer.

ATTA and the missing ground truth

The implementation, ATTA (Automatic Time-span Tree Analyzer), is deliberately plumbing-forward: MusicXML in; GroupingXML, MetricalXML, TimespanXML out (linked back to the score note-by-note); a Perl analysis engine served over CGI; a Java GUI with an automatic-analysis mode and a manual-edit mode. That last item quietly solved the field's biggest gap: there was no ground truth to evaluate against. So the authors built it — 100 eight-bar monophonic classical pieces, each analyzed by a musicology expert using the manual editor, cross-checked by three more experts: the first GTTM analysis database, since grown into a community resource.

Scoring is by F-measure — the harmonic mean of precision (of the boundaries/dots/heads you predicted, how many were right?) and recall (of the right ones, how many did you find?) — a single number that punishes both over- and under-detection. With ~10 minutes of human parameter tuning per piece:

Results · F-measure over 100 pieces
Grouping — baseline
0.46
Grouping — tuned
0.77
Meter — baseline
0.84
Meter — tuned
0.90
Time-span — baseline
0.44
Time-span — tuned
0.60

Default parameters vs. per-piece tuning. Meter is nearly solved; grouping responds strongly to tuning; the full time-span tree remains genuinely hard.

Read the table as three verdicts. Meter: nearly solved — beat hierarchies are the most regular structure in music. Grouping: strongly parameter-sensitive (0.46 → 0.77) — exactly the analysis where the "correct weights vary per piece" hunch matters most. Time-span trees: genuinely hard (0.60 tuned) — errors compound up the hierarchy, and headship depends on the harmonic knowledge the system deliberately excluded.

What the failures teach

Three case studies, each a different moral. Mozart K.331: the opening bar legitimately groups two ways, and ATTA produces both under different parameter settings — domain ambiguity handled as designed; this is the success mode. Chopin, Ballade No. 1: the melody changes character mid-piece, and no single parameter setting fits both halves — optimize globally and the first half's analysis breaks. Fixed-per-piece parameters are still too rigid; parameters may need to be time-varying. Bizet, L'Arlésienne: the piece ends on a genuine single-note group — which hard-coded GPR1 ("avoid single-note groups") forbids absolutely. A preference had been implemented as a law; real music votes preferences down sometimes, and an implementation must leave them outvotable.

The through-line of all three: the remaining problem is no longer representing musical knowledge (exGTTM's parameters do that) but choosing parameter values — which is a learning problem. Hence the successors: FATTA attempted full automation; the ΣGTTM line learned structures via probabilistic grammars; deepGTTM learns the rule applications with multi-task neural networks. The arc of the whole enterprise — hand-built rules → parameterized rules → learned parameters — recapitulates, in miniature, the history of AI itself.

GTTM wouldn't run because its prose was underdetermined: vague magnitudes, unrefereed rule conflicts, missing algorithms. exGTTM's cure is total externalization — 46 explicit parameters, [0,1]-normalized evidence, hierarchy-building as constraint satisfaction — implemented as ATTA and evaluated against a hand-built, cross-checked 100-piece ground truth (F-measures: meter ≈ solved, grouping parameter-hungry, trees hard). The failures (Chopin's mid-piece shift, Bizet's forbidden ending) show the frontier moving from representing knowledge to learning the knobs.

Check yourself
Why is "full parameterization" scientifically valuable rather than a fudge?

With parameters explicit, a fixed setting makes output objective and reproducible — the theory becomes testable against human analyses (Temperley's criterion). And the domain's genuine ambiguity becomes representable: different settings generate different valid hearings. The fudge was the prose that hid these knobs; externalizing them is the honest move.

What are the two kinds of ambiguity, and which must an implementation preserve?

Music's own multivalence (a passage supporting several hearings — preserve it, and ATTA does, producing both K.331 groupings) versus the theory's representational vagueness (undefined magnitudes and procedures — eliminate it via externalization). Confusing the two leads to systems that force one "correct" answer onto genuinely ambiguous music.

What does each famous failure case teach?

Chopin's Ballade: one parameter setting can't fit a piece whose character changes — parameters may need to vary over time. Bizet's Farandole: implementing the preference GPR1 as a hard law breaks on music that legitimately violates it — preferences must stay outvotable. Both point past hand-tuning toward learning.

Why did the authors deliberately exclude harmony and prolongational reduction?

"What we have not implemented enables us to implement GTTM": prolongational theory was unsettled and requires chord-distance knowledge (TPS) outside GTTM proper. Scoping to monophonic grouping/meter/time-span analysis made a working, evaluable system possible at all — Chapter I's layer discipline applied to their own project.

Chapter IX

Melodic morphing

The payoff of the whole book: with distance, meet, and join in hand, "a melody 30% of the way from A to B" stops being a metaphor and becomes a computable object — with provable properties.

The brief: composition for people with deadlines

The authors are explicit about who this is for — not the auteur composing from silence, but the working arranger who must produce variants of existing material for games, films, and sessions: transparent process, predictable output, no surprises in the delivery. The scenario: you have melody A; melody B carries a feeling you want; give A some of B's nuance, to a controllable degree. Anyone who has said "make the alt mix 30% closer to the demo" knows the request. The contribution is making "30%" mean something exact. Four requirements pin it down: (1) the result C sits between A and B in distance; (2) if B = A then C = A (morphing toward yourself changes nothing); (3) the blend ratio is a real parameter; (4) monophonic in, monophonic out.

The algorithm: three steps on the parallelogram

The name "morphing" comes from image processing — the face-blend effect Michael Jackson's Black or White video made famous in 1991. Image morphing works by linking corresponding features (eyes to eyes), then interpolating everything between. The melodic version does exactly this, with Chapter VII supplying both the "corresponding features" (the meet — the notes two melodies structurally share) and the space to interpolate in (tree distance):

  1. Link: compute meet(σA, σB) — the common skeleton. Everything not in the meet is a private ornament of A or of B.
  2. Fade: partially reduce σA toward the meet, keeping the fraction of A's private ornaments the ratio demands — pruning in a defined order (fewest metrical dots first) — giving α. Symmetrically reduce σB to β with the complementary fraction.
  3. Merge: compute join(α, β) and render the tree back to a score.

Because reduction obeys the subsumption chain σA⊓σB ⊑ … ⊑ α ⊑ σA, the result inherits guarantees no ad-hoc blender has: the skeleton is preserved exactly, and — via the parallelogram identity — d(A,C) + d(C,B) = d(A,B): C lies on the line segment between A and B. That last condition is subtler than it looks: infinitely many melodies sit at distance ratio M:N from A and B (a whole locus of them — the Apollonius construction from geometry), and almost all of them wander off the A–B line, mixing in material foreign to both parents. The algorithm's guarantee is stronger: no detours, no foreign notes, pure interpolation.

The two implementation dragons

Real algebra meets real music in two places, and the book is candid about both. Dragon one: order. "Prune the note with fewest dots first" isn't always right (a structurally vital note can sit on a weak beat), and a unique pruning order may even be undesirable — multiple orders would yield more diverse morphs. Dragon two: collisions. Join can superimpose a left-branching node from α with a right-branching one from β — a ternary node, which renders as two notes at the same instant: your monophonic morph sprouts a chord, violating requirement (4).

Two responses, tried in sequence. The logically pure response: admit ternary nodes as superimpositions and render them as simultaneous notes — keeps the algebra intact, but listeners in the evaluation heard those moments as belonging to one parent, skewing the morph's perceived position. The practical response: make everything deterministic by ranking every branch — breadth-first through the tree by maximum time-span, main branch before off-branch — then prune by rank, and resolve every collision by keeping the note from the less-reduced parent. Given A, B, and a ratio, exactly one melody comes out. The trade is explicit: arbitrariness eliminated, some expressive variety (the passing tones and appoggiaturas a human morpher would improvise — an appendix documents a musicologist doing exactly that by hand) given up.

Morph bench · two variations, one slider

A toy model of the algorithm on a Twinkle-style skeleton. Variation A ornaments it with arpeggio figures; Variation B with neighbor tones. Their meet is the bare skeleton; intermediate positions keep a ratio of each side's private notes, exactly as the algorithm prescribes.

Set the slider to pure A and play; then pure B; then the meet — the skeleton both share. Now walk 3:1 → 1:1 → 1:3 and hear A's arpeggio ornaments hand over, slot by slot, to B's neighbor tones while the skeleton never flinches. In the roll: mint = shared skeleton, brass = A's private notes, ivory = B's. This is the entire algorithm, miniaturized — reduce both parents toward the meet, join the remainders.

Did it work? Ears, again

Same protocol as Chapter VII: generate 1:1, 1:3, 3:1 morphs among three Mozart variations (Nos. 1, 2, 5 — chosen because their joins all exist), have six subjects rate all pairwise similarities, MDS the matrix, and look. The morphs of 1&2 and of 5&1 land where they should: between their parents, near the midpoint for 1:1. The 2&5 morph lands off-center toward No. 5 — and the post-mortem traced it to dragon two: simultaneous-note renderings in the first half that listeners attributed to No. 5's texture. A clean quantitative failure with a known cause is a good day in research: it validated the deterministic-priority redesign, whose output (reproduced in the book's final figures) morphs smoothly with no chords at all.

Why this matters beyond the demo

Set this against the generative tools you already know. A Markov chain or a neural model gives you plausible surface with no handle on structure — you can't tell it "keep the skeleton, trade the ornaments, 30/70." The morphing algebra is the opposite trade: modest surface ambition, total structural control, and — the rare thing — proofs about its output: skeleton preserved, result on the A–B segment, ratio exact, deterministic. It's the difference between sampling from a distribution and programming in a well-behaved calculus — precisely the "music as arithmetic" dream Chapter VII announced. And the chapter's honest caveats (pitch still underweighted in the metric; rendering still ad hoc) double as its research agenda: fold TPS pitch distance into d; make rendering principled. The frontier, stated as homework.

Morphing = meet, two partial reductions, join: link the shared skeleton, fade each parent's private ornaments by the chosen ratio, merge. The output provably preserves the skeleton and sits exactly on the line segment between the parents — no foreign material — once two implementation dragons (pruning order, ternary-node collisions) are slain by breadth-first branch priority. Listening tests place the morphs about where the geometry says, with the one systematic miss tracing to the known collision issue. Structure-aware, controllable, provable generation — the book's thesis, cashed.

Check yourself
Why morph on time-span trees rather than directly on notes?

Note-level interpolation has no notion of importance — it degrades skeleton and ornament alike. Trees separate the two, so the morph preserves what the melodies share and interpolates only what distinguishes them; and tree distance is what makes "30% of the way" a well-defined location rather than a vibe.

Why isn't "same ratio of distances to A and B" a sufficient spec for the morph?

Infinitely many melodies satisfy any given ratio — a whole Apollonius locus of them — most containing material foreign to both parents. The algorithm's stronger guarantee, d(A,C)+d(C,B)=d(A,B), pins C to the line segment: pure interpolation, nothing imported.

What causes ternary nodes, and what are the two ways out?

Joining nodes whose branches point opposite ways superimposes left- and right-branching — rendering as two simultaneous notes. Way one: accept the superimposition (logically clean; listeners mis-hear the result as one parent's texture). Way two: prioritize branches deterministically and keep the less-reduced parent's note at collisions — monophony guaranteed, some variety sacrificed.

Epilogue

Ten questions, after Hilbert

In 1900 Hilbert posed the problems that would shape a century of mathematics. The authors close by posing theirs — for the future of music in the age of computation.

1 · Can a new temperament replace 12-tone equal temperament?

Computers retune at will and microtones are trivially playable — but our eardrums, fingers, and habits are sized for twelve. The current system is an optimum of compromises; the question is whether any pleasure lies beyond it that's worth the cost of leaving.

2 · When will the five-line staff be abandoned?

The staff fossilizes pre-equal-temperament pitch spacing — E–F and B–C look like every other step but aren't. Machines don't need it (MusicXML, piano rolls); only the training of human players keeps it alive. Which generation lets go?

3 · Will the black-and-white key arrangement survive?

The keyboard is music's QWERTY: an accident standardized too deeply to amend. Keys are switches for fixed pitches — hopeless for microtones. Does the piano's centrality fade with it?

4 · Can music gain elements beyond melody, rhythm and harmony?

Musique concrète brought unnotatable real-world sound; rap re-fused language and rhythm; film gives music external meaning. When computers compose, do algorithm and complexity themselves become audible elements?

5 · Does music survive without visual media?

Hi-fi culture died into convenience; music now arrives attached to video, drama, dance. Wagner hid his orchestra in a pit to serve the eye. Perhaps music was never destined to stand alone — vision is the stronger input bias.

6 · Will we still learn to play instruments?

Probably — for the reason amateurs still play tennis against the fact of Federer: the doing is the pleasure. Playing survives because it's fun, not because it's needed.

7 · What becomes of professional performers?

The catalogue of great recordings is complete and AI can blend an Argerich attack with a Brendel architecture on a sampled Steinway. Yet performance isn't an Olympic event — we may keep paying precisely for the human in it. The market, though, may not agree.

8 · Will concerts still be held?

Storr: music's primary function is "to bring and bind people together." We attend to share emotion, not to optimize acoustics — which robots cannot host. But training professionals requires a market large enough to pay for them.

9 · Which music will still be heard in 100 years?

Prediction has a bad record: the St. Matthew Passion was forgotten for decades until Mendelssohn revived it; Schoenberg promised we'd whistle twelve-tone rows; bebop's revolution became restaurant wallpaper; McCartney called The Beatles future classical — and may be right. Music and listeners co-evolve; valuation is capricious. Keep the masterpieces in digital form and give them chances to be rediscovered.

10 · Can a computer acquire creativity?

David Cope's reconstructed classics were praised until the composer was revealed to be a program. Dreyfus said computers merely recombine; but a system with no palate can invent a cocktail humans love. If a computed piece gives us the feeling of novelty — the authors suggest — we may have to call that creativity.

Reference

Glossary

gestalt
A perception that exists only in a combination of parts — three chords that individually mean nothing but together mean "the end."
intersubjectivity
Subjective perceptions that nevertheless agree across (nearly) everyone in a community — the stable ground for a science of musical meaning.
cognitive reality
The property a model has when it rationally explains what listeners actually perceive; the book's admission criterion for theory.
reduction
Removing less important notes step by step to expose a skeleton. Schenker's intuition; GTTM's engine; Chapter VII's algebra.
elaboration
Reduction's inverse: growing a surface from a skeleton by adding ornamental events.
Ursatz
Schenker's fundamental structure: a descending upper line over an I–V–I bass, claimed to underlie every tonal piece.
prolongation
An event's psychological persistence beyond its literal duration, governing what follows — the basis of the prolongational tree's tension and relaxation.
pitch class
A note name with octave ignored: all C's are one pitch class. The twelve of them form the cyclic group ℤ/12.
Pythagorean comma
The 23.5¢ gap by which twelve pure fifths overshoot seven octaves (531441/524288). The irreducible error every temperament must hide somewhere.
syntonic comma
81/80 — the gap between the Pythagorean third (81/64) and the pure third (5/4). Meantone shaves a quarter of it off each fifth.
cent
1/1200 of an octave on a log scale; 100 per equal semitone. The common ruler for comparing temperaments.
wolf fifth
The one badly-out-of-tune fifth where meantone's accumulated error is dumped — unusable, and howling.
cadence
The chord formula of an ending: V–I (authentic), IV–I (plagal). Recursively embeddable — the argument for harmony being context-free.
context-free grammar
Rewriting rules whose left side is a single symbol; recognized by a push-down automaton. Home of human syntax and, weakly, of tonal harmony.
preference rule
GTTM's ranking rule: among well-formed analyses, which one an experienced listener hears. No counterpart in linguistics.
time-span tree
The binary tree of salience produced by GTTM's third analysis: at each junction the more important event survives toward the root.
head
The event that survives a junction — the most important event of its time-span; the piece's head survives to the root.
maximum time-span
The longest interval an event dominates in the tree; the unit of information used to define reduction distance.
meet / join
Lattice operations on time-span trees: meet = the largest common reduction (shared skeleton); join = the smallest tree containing both.
basic space (TPS)
Lerdahl's five-level pitch representation (root, fifth, triad, diatonic, chromatic); chord distance counts fifth-steps plus level-climbs.
tonal tension
In TPS: distance travelled in chord space, accumulated down the prolongational tree. High δ arriving = tension; low δ home = release.
melodic attraction
α(p₁→p₂) = (s₂/s₁)·1/n²: notes gravitate toward stronger anchors nearby. The leading tone's pull to the tonic is the strongest in the key.
morphing
Generating a melody at a chosen ratio between two others: reduce each toward their meet, then join. Lands exactly on the line segment between them.
virtualization
The book's reframe of symbolization: abstracting physical detail to expose clean function — 12-TET for pitch, chord names for harmony, MIDI for performance.
idiostructure
Narmour's term for the structural features that make a work recognizably its composer's — the un-scientific half of music theory's double aim.
closure
In I-R: a point where implication weakens or stops — a long note, rest, strong beat, or change of direction — where patterns begin and end.

Built as a study companion to Keiji Hirata, Satoshi Tojo & Masatoshi Hamanaka — Music, Mathematics and Language: The New Horizon of Computational Musicology Opened by Information Science, Springer Nature Singapore, 2022 (ISBN 978-981-19-5166-4). All concepts, rules, data tables and examples paraphrased from the book for personal study; go to the source for the full derivations, figures and proofs.

Further primary sources: Lerdahl & Jackendoff, A Generative Theory of Tonal Music (MIT Press 1983) · Lerdahl, Tonal Pitch Space (OUP 2001) · Narmour, The Analysis and Cognition of Basic Melodic Structures (Chicago 1990) · Meyer, Emotion and Meaning in Music (Chicago 1956).