An interactive companion

Music, Mathematics & Language

The book opens with a pianist's question: why are the black and white keys arranged as they are? Answering it takes 2,500 years of mathematics — and by the end, music has become something a computer can parse like a sentence, reduce like an equation, and recombine like algebra.

This page teaches the core ideas of Music, Mathematics and Language: The New Horizon of Computational Musicology (Keiji Hirata, Satoshi Tojo & Masatoshi Hamanaka, Springer 2022). Almost every concept here can be heard, not just read — start below.

Temperament keyboardC4 – C5

Play the octave. Then switch the tuning system underneath it and play the same keys again — same note names, different mathematics. The readout shows each note's ratio to C and its deviation in cents from equal temperament.

Press a key…

The oscilloscope draws the sum of the sounding waves. Watch a just-intonation major third lock into a still, repeating shape — then hear the same third in equal temperament shimmer, because 2^(4/12) is close to, but never equal to, 5/4.

All audio is synthesized live in your browser — nothing to load, but you will want speakers or headphones.
Chapter I

A machine that computes the meaning of music

Can "the meaning of music" be studied scientifically? The book's answer: yes — if you separate the layers, and only claim the ones a computer can check.

Music theory was, historically, a body of tacit knowledge passed from master to apprentice. The authors propose making it explicit: state a hypothesis about how listeners hear music, implement it as a program, run it, and see whether the output matches human judgment. In Karl Popper's sense, that makes musicology refutable — and therefore a science. The obstacle is Moravec's paradox: computers tap perfect beats and name absolute pitches (hard for humans), yet cannot tell a happy song from a sad one (easy for an infant). What infants have and machines lack is tacit knowledge.

Gestalt: meaning that only exists in the whole

A gestalt is a perception that appears only when parts combine — meaning greater than the sum of the components. The book's example: in Japanese elementary schools, teachers open class with three chords played to "stand up — bow — sit down." Each chord alone is just a sound; in sequence, every student hears "the lesson has begun." And nearly everyone hears it the same way — that shared subjectivity is called intersubjectivity, and it's what lets a science of music-meaning get off the ground without pretending taste is objective.

Gestalt demo · stand up, bow, sit down

Hear each chord in isolation, then in sequence. Isolated, they're inert. Chained, the third chord sounds unmistakably like an ending — a perception that exists nowhere in the individual chords.

I → V⁷ → I  ·  the same progression closes Beethoven's Ode to Joy

Semiotics: how notes point at other notes

Peirce split signs into three kinds: icon (resembles its object — a hieroglyph), index (physically caused by it — a thermometer), and symbol (pure convention — the letter A, a quarter note). Music has no pronouns and no dictionary, so what do its signs point at? The book's move: phrases point at other phrases. When you hear do-re-mi-fa do-re-mi-fa, the second occurrence refers back to the first while implying a third. A melody builds a referential network of memory and prediction — and that network is a large part of its meaning.

Meyer's four kinds of musical meaning

Intrinsic · Emotion

Meaning inside the music: expectations raised by phrases and either realized or denied, stirring emotion. This is the book's home territory — and Chapter V's whole subject.

Extrinsic · Emotion

"Darling, they're playing our tune": music tied to memories and scenes outside itself. Real, but unstable — poor material for computation.

Intrinsic · Form

Pleasure in pure structure: minimalism, or James Brown's Sex Machine, built almost entirely on one E⁷ chord and a rhythm pattern.

Extrinsic · Form

Structure referencing something outside music: Xenakis composing from probability formulas, arabesque melodies tracing mosque ornament.

Schenker's reduction hypothesis

Heinrich Schenker claimed every tonal piece hides a skeleton. Strip away the less important notes step by step — reduction — and you pass from the surface (Vordergrund) through middle layers to a fundamental structure (Ursatz): a descending line over a bass that rises a fifth and falls back (I–V–I). The reverse operation, growing a surface from the skeleton, is elaboration. Schenker described this in intuitive prose; nobody could state it as an algorithm. That gap — between a brilliant idea and an executable procedure — is what the rest of the book exists to close.

Retrograde · creation vs. reception

Composers since the Middle Ages have hidden melodies played backwards (Bach's crab canon; Beethoven's Hammerklavier fugue). But listeners almost never hear a retrograde — most only spot it in the score. Try it honestly: play the melody, then its reverse. Would you have noticed the relation blind?

Molino's trichotomy: the composer's intention and the listener's reception meet only in the physical trace — the score — and they need not agree.

The lesson generalizes. Short-term memory holds 7±2 chunks for a few seconds — roughly a four-bar phrase. Structure spanning forty bars can't live in a listener's working memory; it belongs to the composer's side of the trace. A theory must say whose structure it is describing. The book closes the chapter with its criterion for all that follows: a model has cognitive reality when it rationally explains what listeners actually perceive — and experiments suggest that tree structures over notes do.

Check yourself
Why do the authors insist on separating a "fundamental layer" from taste and artistic quality?

Taste is subjective and varies per listener; the score-level structures grasped through gestalt and intersubjectivity are stable across listeners. Only the stable layer supports refutable, reproducible claims — the subjective layer can then be modeled on top of it. Skipping this separation yields theories that are narrow, unstable, and inaccurate.

In Peirce's terms, is a crotchet (quarter note) an icon, an index, or a symbol?

A symbol: nothing about the glyph physically resembles or is caused by "one quarter of a bar." The connection is pure convention — like Morse code or the alphabet.

Chapter II

The mathematics of ebony and ivory keys

Every tuning system is a compromise between two facts that cannot both be true: octaves are powers of 2, fifths are powers of 3, and no power of 3 is ever a power of 2.

Pythagoras: building a scale from 2 and 3 alone

Two strings sound consonant when their frequencies form a simple ratio — their waves' nodes keep coinciding. Ratio 1:2 is the octave; 2:3 the perfect fifth. Pythagoras (and, independently, Jing Fang in Han-dynasty China) generated a whole scale from just these: multiply by 3/2 to get a new note, halve whenever you overflow the octave. Five steps give the pentatonic scale; keep going and something goes wrong at step twelve.

Stack of fifths · find the comma

Stack fifths one at a time. Each new note is (3/2)ⁿ folded back into one octave. The twelfth note should land back on C — listen to how close it gets.

0 fifths stacked — C = 1

Just intonation: inviting 5 to the party

The Pythagorean major third 81/64 is harsh. The next prime, 5, offers 5/4 — barely flatter (the gap, 81/80, is the syntonic comma) and far sweeter. Replacing the third-related notes with ratios of 5 gives just intonation: C:E:G = 4:5:6, and every pair in the scale locks into a simple ratio. The price: two different whole tones (9/8 and 10/9), which makes transposing a melody impossible on fixed-pitch instruments — the intervals come out warped.

Meantone: sacrifice the fifth to save the third

Quarter-comma meantone flattens each fifth by a quarter of the syntonic comma, so that four of them stack to a pure 5/4 third: the fifth becomes ⁴√5 ≈ 1.4953 instead of 1.5. Lovely thirds, uniform whole tones — but the leftover error piles into one interval, the notorious wolf fifth, which howls. Kepler, Mersenne, Euler and Rousseau all took turns at this problem. There is no perfect answer; there are only different places to hide the comma.

Comparison bench · four temperaments

One interval, four tunings. The cents column shows deviation from just (pure) intonation — the smaller, the smoother. Then meet the wolf.

Pythagorean → Just → Meantone → Equal, in order

Equal temperament: the great compromise

Abandon purity entirely: divide the octave into twelve identical semitones of ratio ¹²√2. Every interval is now slightly wrong and every key is equally wrong, so modulation — restating a melody from any starting note — becomes free. Measured in cents (1200 per octave, 100 per equal semitone), the systems compare like this:

Scale stepCDEFGABC
Equal020040050070090011001200
Meantone019338650369789010831200
Just020438649870288410881200
Pythagorean020440849870290611101200

Green = pure interval; red = the wide Pythagorean third. The pure third is 386¢; equal temperament's 400¢ third is 14¢ sharp — that's the shimmer you heard in the hero keyboard.

Historically, the completion of 12-tone equal temperament and the modern keyboard immediately preceded the explosion of classical music — nearly every canonical composer from Haydn to Mahler appears within 150 years of Bach's Well-Tempered Clavier. Much of what came after (Debussy's whole-tone scale, Bartók's folk modes, Schoenberg's dodecaphony) can be read as attempts to escape the system's gravity. And a warning the authors add: equal temperament absorbs every ethnic scale into its grid — a versatility that quietly erases the microtonal diversity it approximates.

Symmetry: the octave as a cyclic group

Because all twelve semitones are identical, transposition becomes pure arithmetic. Let ρ = "raise by one semitone." Then ρ¹² = e (the identity), and the pitch classes form the cyclic group ℤ/12 — arithmetic modulo 12. A fifth up equals a fourth down: ρ⁷ = ρ⁻⁵. Define ξ = ρ⁷ and the powers of ξ generate the circle of fifths — which is exactly Pythagoras's construction, rediscovered as group theory. Chord inversion is a group; the retrograde and inversion operators of twelve-tone music form one too.

Cyclic group ℤ/12 · generators ρ and ξ

Apply an operator repeatedly and watch it walk the octave. ρ (semitone) visits notes chromatically; ξ = ρ⁷ (fifth) visits all twelve in circle-of-fifths order before returning home — both generate the whole group.

at C — ρ⁰(C)
Check yourself
Why can't a stack of pure fifths ever return exactly to C?

It would require (3/2)ᵐ·(1/2)ⁿ = 2, i.e. 3ᵐ = 2ⁿ⁺ᵐ⁺¹ — impossible, since 3 and 2 are coprime. Every tuning system is a strategy for hiding this irreducible error.

What does just intonation buy you, and what does it cost?

Buys: maximally consonant chords (C:E:G = 4:5:6). Costs: two unequal whole tones (9/8 vs 10/9), so melodies can't be transposed on fixed-pitch instruments — modulation breaks. Equal temperament makes exactly the opposite trade.

Why is 7 ≡ −5 (mod 12) musically meaningful?

Rising a fifth (7 semitones) lands on the same pitch class as falling a fourth (5 semitones). Interval pairs summing to 12 are octave-complements: P4↔P5, M3↔m6, M2↔m7 — the symmetry table of Chapter II, stated as modular arithmetic.

Chapter III

Music as formal language

If music and language share one origin — Darwin thought our ancestors sang before they spoke — then music should have a grammar. The question is: which kind?

Chomsky's ladder of grammars

A formal language is a set of sentences generated by rewriting rules. Chomsky ranked grammars by power, each with an equivalent machine:

TypeGrammarMachineExample
3RegularFinite automatonab*a — birdsong
2Context-freePush-down automatonaⁿbⁿ — human language, cadences
1Context-sensitiveLinear-bounded automatonaXb → aYZb
0Phrase structureTuring machineanything computable
Finite automaton · accepts strings with ≥ 3 a's

The book's example machine: four states, transitions on a and b. Type any string of a's and b's and step through it. It reaches the accepting state ③ exactly when the string contains at least three a's — equivalent to the regular expression b*ab*ab*a(a+b)*.

state 0 · nothing read yet

Why human language — and harmony — need a stack

In "We go to Broadway to see a musical," dependencies nest; they don't cross. Nesting without crossing is exactly balanced parentheses, and recognizing balanced parentheses requires a push-down stack: open a dependency → push; resolve it → pop. That working memory is what puts human language at the context-free level — finches, whose songs are regular, don't have it. (A few exceptions exist: Dutch subordinate clauses and French clitic pronouns really do cross.)

Now the transfer to music. A cadence — the chord formula that signals an ending, I ⋯ V–I or I ⋯ IV–V–I — can be embedded in another cadence: precede V with its own fifth (the double dominant II) and you get II–V–I, a cadence inside a cadence. Recursion of that kind is precisely what context-free grammars capture and regular grammars cannot. When you hear a dominant chord and feel something owed, that's your push-down stack holding an open parenthesis.

Grammars people have actually written for chords

  • Winograd (1968) — the first serious CFG for harmony: cadence → plagal | authentic; authentic → dominant tonic; … — implemented in LISP, decades before "music informatics" existed.
  • Rohrmeier's Generative Syntax Model (2011) — a full rule system: pieces generate tonal regions (TR → DR t, DR → SR d), regions generate Riemannian functions (tonic, dominant, subdominant and their parallels), functions map to scale degrees (t → I, d → V | VII), with modulation and secondary-dominant rules. Recursive structure of tonal harmony, stated formally.
  • Probabilistic CFG — attach a probability to each rule and a corpus can vote: a rule set fitted to Bach's chorales (e.g. T → i [0.8] | iii [0.1] | vi [0.1]) encodes his stylistic fingerprint in the numbers. Same rules, different probabilities: different composer.
  • HPSG for harmony — chords as feature structures: V, V⁷ and V⁶ are one category with different internal features, unified by head-driven rules. The book (Tojo et al.) parses C–Am–F–Dm–G–C into a full cadence tree this way.
  • CCG for jazz — Steedman-school categorial grammar where Dm7–G7–C parses two ways (chain of fifths, or Dm as substitute for F). Ambiguity isn't a bug; it's two hearings.

One deep asymmetry, and it drives the whole rest of the book. In language, a sentence usually gets one parse tree; grammatical categories pin it down. In music, Lerdahl & Jackendoff observed, almost any passage is "vastly ambiguous" — many trees fit. So a music grammar needs two kinds of rules: well-formedness rules (which trees are possible) and preference rules (which tree an experienced listener actually hears). Preference rules have no counterpart in linguistics — and they are why reduction matters in music but not in syntax. Bernstein's 1976 Harvard lectures drew the language–music analogy in public; GTTM made it precise.

Check yourself
What musical experience corresponds to the "push" of a push-down automaton?

Hearing a chord that creates an expectation — a dominant demanding resolution, a half cadence opening a question. The unresolved tension is an open parenthesis on your mental stack; the resolving tonic pops it.

Why does music theory need preference rules when linguistics doesn't?

Musical surfaces underdetermine their structure: many well-formed trees fit one passage, because notes don't carry rigid categories the way words do. Preference rules rank the candidates by how an experienced listener would hear them. In language the surface typically contains enough information to force (nearly) one tree.

Chapter IV

The Berklee method

Write G7 above a melody line and any competent player on Earth can realize it. That compression — harmony as a symbol system — is one of the three great symbolizations of music.

The method, born at Berklee in 1940s Boston alongside bebop, builds every chord the same way: stack thirds on a root — root, 3rd, 5th (triad), add the 7th (tetrad), then optionally the tensions 9, 11, 13. Octave position is ignored (only pitch class matters), inversions get the same name, and — crucially — the name records the root's letter, not its scale degree. G7 is G7 in any key. The key, and therefore the chord's function, is left to the performer's interpretation. That deliberate underspecification is what makes reharmonization and ad-lib possible: same melody, new chords; same chords, new melody.

Chord constructor · name ↔ sound

Assemble a chord the Berklee way and hear it. The name is generated by the same rules a chart uses.

Consonance by decree

Where the classical ear ranked intervals by frequency ratio, the Berklee method simply legislates dissonance levels as an axiom of the system — thirds and sixths consonant, seconds and sevenths dissonant, the tritone unstable — and builds its tension rules on top. The book calls the result a "puzzle gamification" of harmony: which tensions are legal where becomes a rule-game, like chess.

LevelCharacterIntervals
1ConsonantM3, m3, M6, m6
2Neutral, stableunison, P8, P5
3InorganicP4
4Moderately dissonantM2, m7
5Acutely dissonantm2, M7
6Unstabletritone (+4)

Three symbolizations, one arc

Jazz musician Naruyoshi Kikuchi and critic Yoshio Otani identify three moments when music was successfully turned into symbols — the book reframes each as virtualization, in the computer-science sense of abstracting away physical detail to expose clean function:

≈ 1722 · 12-TET

Equal temperament virtualizes pitch: any melody reproducible on any instrument in any key. Bach's Well-Tempered Clavier is its manifesto.

≈ 1945 · Berklee method

Virtualizes harmony: chord charts separate composition from performance, enabling mass production of popular music — and giving bebop its rule-system.

1981 · MIDI 1.0

Virtualizes performance: key, timing and velocity as digital events. Score and performance become computable; technopop and the bedroom producer follow.

The trade-off

Each innovation navigates expressiveness vs. description cost. A lead sheet is cheap to write and rich to interpret; an orchestral score is the reverse. (You live on the MIDI side of this trade every day.)

The chapter ends on a cycle the authors see across genres: symbolization complexifies (late bebop, fusion — "Berklee sickness") until the music saturates; a reaction returns to simplicity (modal jazz, impressionism); and the two are eventually sublated into a new syntax. Miles Davis appears twice — Kind of Blue (1959) making mode, not chord change, the engine of motion; Bitches Brew (1969) fusing jazz, rock and R&B into a syntax that could express things the old ones couldn't. The book's test for a genuinely new syntax is software versioning: v2.0 can express v1.0's pieces, but not vice versa.

Check yourself
Why does a Berklee chord symbol deliberately omit the key?

Leaving function un-notated hands interpretation to the musician: the same chart supports different keys, reharmonizations, and improvisation. Notating degrees (I, IV, V) would fix an interpretation; notating letter-roots keeps it open. Wagner's Tristan chord has 33 published functional interpretations — but exactly one Berklee name, Fø7.

What can the Berklee method not express?

Anything needing octave information (voicings are abstracted to pitch classes), chords not built in thirds, deliberately omitted chord tones (except by heuristics like "omit 3"), and microtonal or non-12-TET material. Symbolization always trades away something.

Chapter V

The Implication–Realization model

You hear two notes. Your ear has already bet on the third. Narmour's model says which bet — and emotion lives exactly where the bet is lost.

Meyer's founding idea: musical emotion arises from expectation. Confirmed predictions pass unnoticed; denied ones move you. Eugene Narmour formalized this into the Implication–Realization model: where GTTM understands music by grouping notes into hierarchies (sets), I-R understands it by the network of references between notes (relations). Two hypotheses drive everything — X + X → X (hear a thing twice, expect it again) and X + Y → Z (hear a change, expect further change) — refined by two principles applied to the first two notes' interval:

PRD · Registral direction

Small interval (≤ 5 semitones) → expect the melody to continue in the same direction. Large interval (≥ 7) → expect a reversal.

PID · Intervallic difference

Small interval → expect a similar-sized interval next (±2 semitones). Large interval → expect a smaller one: after a leap, fill the gap.

Satisfying combinations get pattern names: D duplication, P process (stepwise continuation), R reversal (leap, then turn back), plus intermediate forms ID, IP, IR, VR and the returning figure aba. Patterns start and end at points of closure — long notes, rests, strong beats, changes of direction — and they can overlap.

Expectation lab · bet against 91 music students

Choose an opening interval. The model states its implication; press play to hear the two stimulus notes and then the model's predicted continuation. Then reveal what Carlsen's subjects (91 students, Hungary/Germany/USA, 1981) actually sang — the top three responses and how often.

The data largely vindicate the model: after a major second up, 64% of subjects continued stepwise in the same direction (pattern P); after an octave leap, the top responses all reverse direction (R). The misses are as instructive as the hits — descending sevenths behaved like leading tones resolving to the octave, and tritones resolved by semitone — showing the scale-step: learned tonal schemas riding on top of the bottom-up gestalt principles. Narmour's chess metaphor: a semitone is a pawn (you know where it's going); an augmented sixth is a queen (it could go anywhere, but position constrains it).

Check yourself
Why does the model predict emotion at C–D–F but not at C–D–E?

C–D (small, up) implies same-direction, similar-interval continuation: E. E realizes the implication — no emotional jolt. F continues the direction but widens the interval to a minor third: a denial of PID, and (per Meyer) denials are where affect is generated — whether or not you're conscious of it.

How does I-R differ from Schenker/GTTM in what counts as "the melody"?

Reduction theories posit abstract melodies at higher levels, made of chronologically distant notes. Narmour objects that nobody hears those; I-R stays on the actual sounding surface and models the network of expectations between locally adjacent notes.

Chapter VI

GTTM & Tonal Pitch Space

The Generative Theory of Tonal Music turns Schenker's intuition into an explicit system: four analyses that end in a tree — with the piece's most important event at its root.

Lerdahl & Jackendoff's GTTM (1983) analyzes a piece in four stages, each governed by well-formedness rules (what's structurally legal) and preference rules (what an experienced listener actually hears):

  1. Grouping analysis — segment the surface into motives, phrases, sections. Boundaries fall where gaps, leaps, or changes of dynamics/articulation/duration occur (GPR 2–3); larger boundaries where those effects are strongest (GPR 4); groups prefer symmetry (GPR 5) and parallelism (GPR 6). Groups nest without overlapping.
  2. Metrical analysis — find the hierarchy of strong and weak beats (the dot grid): strong beats every 2 or 3 beats, aligned with note onsets, stresses, and long events.
  3. Time-span reduction — combine both: within each time-span, adjacent events compete and the more salient becomes the head, tournament-style, up to a single head for the piece. Cadences get special treatment: V–I is retained as a fused unit, and a phrase's structural beginning [b] "throws a ball" that its cadence [c] catches.
  4. Prolongational reduction — re-derive the tree from the psychology of tension and relaxation: right-branching = departure from the tonic (tension), left-branching = return (relaxation), with three junction strengths (strong / weak prolongation, progression). Tonal music demands tension then release — the normative structure.
Time-span reduction · hear the hierarchy

One plausible time-span analysis of the first phrase of Ah, vous dirai-je, maman ("Twinkle, Twinkle"). Slide the depth control: each level strips away less-salient events, and the survivors stretch to fill the vacated time-spans. Level 1 is the phrase's single head — Schenker's skeleton, made audible.

GTTM's honesty about its own limits matters: its rules say what to prefer but often not how strongly, rules conflict without a resolution procedure, and detecting the cadence in the first place (the half cadence in Mozart's K.331 is the book's running example) requires knowing chord relationships — which GTTM doesn't supply. Lerdahl's Tonal Pitch Space (2001) fills that hole with geometry.

Tonal Pitch Space: measuring the distance between chords

TPS models a chord as a five-level basic space: octave root, fifth, triad, diatonic scale, chromatic scale. A pitch class is more central the higher the level it first appears on. The distance from chord x to chord y in one key is δ(x→y) = j + k: j steps around the chordal circle of fifths, plus k, the number of basic-space levels that pitch classes must climb. Across keys a term i (steps in the circle of fifths between keys) is added, keys arrange themselves on a torus, and chord progressions become paths — preferred paths being shortest, "as objects move in a gravitational field." Tension then falls out for free: sequential tension is distance travelled; global tension accumulates down the prolongational tree; and melodic notes are "attracted" to their stable neighbors like anchored masses.

TPS calculator · δ(x→y) within C major

Pick two diatonic chords. The bench computes j (circle-of-fifths steps I→V→ii→vi→iii→vii°→IV), builds each chord's basic space, counts the level-climbs k, and plays the progression. Compare I→V (δ=5, the closest real move) with I→ii (δ=8, a distant neighbor a step away).

Check yourself
What's the difference between a time-span tree and a prolongational tree?

The time-span tree encodes rhythmic/structural salience, built bottom-up from grouping and meter. The prolongational tree encodes harmonic/psychological tension and relaxation, built top-down by re-attaching events — its branches ignore surface bar-lines because a chord's influence prolongs past them. Same events, two hierarchies.

Why does TPS complete GTTM rather than compete with it?

GTTM's preference rules repeatedly appeal to "stability" and "cadence" without defining chordal proximity. TPS supplies the metric: δ distances identify cadences, choose between candidate heads, and quantify the tension/relaxation the prolongational tree depicts.

Chapter VII

An algebra of trees

If a time-span tree captures a melody's structure, then operations on trees are operations on music. The authors build the arithmetic: distance, union, intersection.

First, a unit of information. Each event's maximum time-span is the longest interval it dominates in the tree — the phrase's head owns the whole phrase; an ornament owns only its own duration. The maximum time-span hypothesis: pruning a branch loses exactly that much information. From it:

  • Reduction paths. Prune terminal branches one at a time and each tree subsumes the last (σ′ ⊑ σ); the set of all reductions of all melodies forms a partially ordered set.
  • Distance. d(σA, σB) = the summed maximum time-spans of the pruned events. Remarkably, the order of pruning doesn't change it, the triangle inequality holds, and it's a true metric.
  • Meet ⊓ and join ⊔. The poset is a lattice: meet(σA, σB) is the largest common reduction — the shared skeleton of two melodies; join(σA, σB) is the smallest tree containing both — their structural union (when the time-spans are compatible; join is partial, meet is total). The four trees σA, σB, their meet and their join form a parallelogram in melody-space with equal opposite sides.
join(σA, σB) — union · 822 ticks σA · Variation No. 2 (744) σB · Variation No. 5 (654) meet(σA, σB) — common skeleton · 576 ticks 168 78 78 168
Mozart, Twelve Variations on "Ah, vous dirai-je, maman" K.265 — variations No. 2 and No. 5. Numbers are total maximum time-spans in ticks (♩ = 12); edge labels are reduction distances. Both routes between σA and σB measure 246: the distance lemma, drawn.

Does the theoretical distance match perception? The authors had eleven subjects rate the pairwise similarity of Mozart's theme and twelve variations, then compared the psychological distance matrix (via multidimensional scaling) against tree distances. The clusters largely coincide — striking, because the tree distance ignores pitch entirely and sees only rhythmic-structural shape. Where they diverged, pitch was the culprit (a variation in the minor feels far away even when its tree is near) — which is precisely why TPS-style pitch distances are the flagged next step.

The chapter's ambition deserves stating plainly: as any computation reduces to arithmetic and any construction to ruler-and-compass, the authors want composition itself to reduce to a small set of well-understood operators on trees — "programming music" the way Norman's design principle demands: the complexity of the tool should be the complexity of the task. Chapter IX cashes the promise.

Check yourself
Why must reduction distance be independent of pruning order to count as a distance?

A metric needs a well-defined value for each pair. If different reduction paths between the same two trees summed differently, d would depend on the route, not the endpoints. The maximum time-span hypothesis guarantees each pruned event costs the same regardless of when it's pruned — so all paths agree.

What do meet and join mean musically?

Meet: what two melodies share — their common structural skeleton (both Mozart variations reduce to the Twinkle theme's core). Join: a melody containing the structural content of both at once. Morphing (Ch. IX) is built from exactly these two operations plus reduction.

Chapter VIII

Implementing GTTM: exGTTM & ATTA

Lerdahl & Jackendoff admitted it themselves: "Our theory cannot provide a computable procedure." Hamanaka, Hirata & Tojo set out to build one anyway.

Three obstacles stand between GTTM's prose and a running program:

  • Ambiguous rules. GPR4 says boundaries go where other rules' effects are "relatively more pronounced" — pronounced by how much, normalized how?
  • Conflicting rules. A leap (GPR3a) and a repetition (GPR6) can each demand a boundary one note apart, while GPR1 forbids single-note groups. GTTM offers no arbitration.
  • Missing algorithms. The rules say which structures to prefer, never how to construct one — especially the hierarchy, where local bottom-up rules (GPR2–3) and global top-down rules (GPR5–6, symmetry and parallelism) must be reconciled.

Their solution, exGTTM, is full externalization and parameterization: every vague quantity becomes an explicit parameter — a strength S per preference rule to arbitrate conflicts, thresholds T for "salient enough," weights for parallelism's pitch-vs-rhythm balance — 15 parameters for grouping, 18 for meter, 13 for time-span reduction, every internal variable normalized to [0,1] so evidence composes cleanly. Whenever a human-correct analysis the system couldn't produce turned up, a new parameter was externalized until it could. The point isn't that 46 knobs are elegant; it's that the theory becomes testable — fix the parameters and the output is objective (Temperley's criterion), and hierarchy-construction becomes a constraint-satisfaction problem, deliberately scoped: monophony only, no harmony, no prolongational reduction.

The implementation, ATTA (Automatic Time-span Tree Analyzer — Perl engine, Java GUI, everything flowing through MusicXML into GroupingXML / MetricalXML / TimespanXML), was evaluated against a hand-built ground truth: 100 eight-bar monophonic classical pieces analyzed by a musicology expert and cross-checked by three more — the first GTTM analysis database. With ~10 minutes of per-piece parameter tuning:

Results · F-measure over 100 pieces
Grouping — baseline
0.46
Grouping — tuned
0.77
Meter — baseline
0.84
Meter — tuned
0.90
Time-span — baseline
0.44
Time-span — tuned
0.60

Default parameters vs. per-piece tuning. Meter is nearly solved; grouping responds strongly to tuning; the full time-span tree remains genuinely hard.

The failure cases teach as much as the scores. Mozart K.331's opening legitimately parses two ways, and ATTA can produce both — ambiguity handled by parameters, as intended. But Chopin's Ballade No. 1 changes character mid-piece, and no single parameter set fits both halves; and Bizet's Farandole ends on a genuine single-note group, which hard-coded GPR1 forbids. Rules that are usually preferences are sometimes wrong — the permanent tension of any preference-rule system. Successors (ΣGTTM's PCFG learning, deepGTTM's multi-task networks) push toward learning the parameters instead of hand-tuning them.

Check yourself
Why is "full parameterization" scientifically valuable rather than a fudge?

With parameters explicit, a fixed setting makes the system's output objective and reproducible — so the theory can be tested, per piece, against human analyses (Temperley's point). The fudge would be leaving the knobs implicit inside prose like "relatively more pronounced."

Why did the authors deliberately exclude harmony and prolongational reduction?

"What we have not implemented enables us to implement GTTM": prolongational reduction was still theoretically unsettled and needs chord-distance knowledge (TPS). Scoping to monophonic time-span analysis made a working, evaluable system possible at all.

Chapter IX

Melodic morphing

The payoff: with distance, meet and join in hand, "a melody 30% of the way from A to B" is a computable object — not a metaphor.

The target user is the working arranger, not the auteur: someone who needs melody A to take on the nuance of melody B, predictably. The algorithm is three algebra steps on the parallelogram from Chapter VII:

  1. Compute meet(σA, σB) — the common skeleton.
  2. Reduce σA partway toward the meet (keeping N:M of its private ornaments, in branch-priority order), giving α; symmetrically reduce σB to β.
  3. join(α, β) — and render the tree back to a score.

The result lands on the line segment between A and B in tree-distance space — d(A,C) + d(C,B) = d(A,B) — not merely somewhere equidistant. Two implementation wrinkles: joining branches of opposite direction yields ternary nodes (superimposed left- and right-branching — rendered as two simultaneous notes), and overlapping notes need a winner; the automated version resolves both by prioritizing branches breadth-first by maximum time-span, making the whole pipeline deterministic. In listening tests with Mozart's variations, subjects placed the 1:1 morphs about midway between their parents on the MDS similarity map — except where those simultaneous notes leaked in and dragged the percept toward one parent.

Morph bench · two variations, one slider

A toy model of the algorithm on a Twinkle-style skeleton. Variation A ornaments it with arpeggio figures; Variation B with neighbor tones. Their meet is the bare skeleton; intermediate positions keep a ratio of each side's private notes, exactly as the algorithm prescribes.

Notice what kind of tool this is. A Markov babbler or a net sampling a latent space gives you plausible surface with no controllable structure; morphing gives you a proof about the output — it preserves the shared skeleton and interpolates the rest, with a dial. That's composition as programming: intention expressed through operators whose behavior you can predict, the design goal set out in Chapter VII.

Check yourself
Why morph on time-span trees rather than directly on the notes?

Note-level interpolation (averaging pitches, crossfading) has no notion of importance: it degrades structure and ornament alike. Trees separate skeleton from elaboration, so the morph preserves what the melodies share and interpolates only what distinguishes them — and the tree distance makes "30% of the way" well-defined.

Epilogue

Ten questions, after Hilbert

In 1900 Hilbert posed the problems that would shape a century of mathematics. The authors close by posing theirs — for the future of music in the age of computation.

1 · Can a new temperament replace 12-tone equal temperament?

Computers retune at will and microtones are trivially playable — but our eardrums, fingers, and habits are sized for twelve. The current system is an optimum of compromises; the question is whether any pleasure lies beyond it that's worth the cost of leaving.

2 · When will the five-line staff be abandoned?

The staff fossilizes pre-equal-temperament pitch spacing — E–F and B–C look like every other step but aren't. Machines don't need it (MusicXML, piano rolls); only the training of human players keeps it alive. Which generation lets go?

3 · Will the black-and-white key arrangement survive?

The keyboard is music's QWERTY: an accident standardized too deeply to amend. Keys are switches for fixed pitches — hopeless for microtones. Does the piano's centrality fade with it?

4 · Can music gain elements beyond melody, rhythm and harmony?

Musique concrète brought unnotatable real-world sound; rap re-fused language and rhythm; film gives music external meaning. When computers compose, do algorithm and complexity themselves become audible elements?

5 · Does music survive without visual media?

Hi-fi culture died into convenience; music now arrives attached to video, drama, dance. Wagner hid his orchestra in a pit to serve the eye. Perhaps music was never destined to stand alone — vision is the stronger input bias.

6 · Will we still learn to play instruments?

Probably — for the reason amateurs still play tennis against the fact of Federer: the doing is the pleasure. Playing survives because it's fun, not because it's needed.

7 · What becomes of professional performers?

The catalogue of great recordings is complete and AI can blend an Argerich attack with a Brendel architecture on a sampled Steinway. Yet performance isn't an Olympic event — we may keep paying precisely for the human in it. The market, though, may not agree.

8 · Will concerts still be held?

Storr: music's primary function is "to bring and bind people together." We attend to share emotion, not to optimize acoustics — which robots cannot host. But training professionals requires a market large enough to pay for them.

9 · Which music will still be heard in 100 years?

Prediction has a bad record: the St. Matthew Passion was forgotten for decades until Mendelssohn revived it; Schoenberg promised we'd whistle twelve-tone rows; bebop's revolution became restaurant wallpaper; McCartney called The Beatles future classical — and may be right. Music and listeners co-evolve; valuation is capricious. Keep the masterpieces in digital form and give them chances to be rediscovered.

10 · Can a computer acquire creativity?

David Cope's reconstructed classics were praised until the composer was revealed to be a program. Dreyfus said computers merely recombine; but a system with no palate can invent a cocktail humans love. If a computed piece gives us the feeling of novelty — the authors suggest — we may have to call that creativity.

Built as a study companion to Keiji Hirata, Satoshi Tojo & Masatoshi Hamanaka — Music, Mathematics and Language: The New Horizon of Computational Musicology Opened by Information Science, Springer Nature Singapore, 2022 (ISBN 978-981-19-5166-4). All concepts, rules, data tables and examples paraphrased from the book for personal study; go to the source for the full derivations, figures and proofs.

Further primary sources: Lerdahl & Jackendoff, A Generative Theory of Tonal Music (MIT Press 1983) · Lerdahl, Tonal Pitch Space (OUP 2001) · Narmour, The Analysis and Cognition of Basic Melodic Structures (Chicago 1990) · Meyer, Emotion and Meaning in Music (Chicago 1956).