AI/ ai · interpretability · music-ai · sparse-autoencoders

Researchers Map Chords and Keys Inside AI Music Models

A new interpretability method shows music AI models organize pitch concepts like chords and keys as structured groups, not single isolated features.

AI models that make or analyze music may organize their internal knowledge the way music theorists do: not as scattered facts, but as structured families like chords and keys.

A new paper proposes a way to look inside music foundation models using sparse autoencoders (SAEs), a common interpretability tool that normally hunts for single, isolated features. The researchers fed models pitch-shifted versions of the same musical input and aligned the resulting SAE representations across those shifts. That alignment surfaced what the authors call "orbits," organized groups of features that track concepts like the 12 transpositions of a chord or the notes within a key. Tested on two state-of-the-art music foundation models, the method reliably recovered structures for chords, keys, and melodic patterns, needing only a handful of anchor examples to interpret an entire concept family at once.

That's a meaningful shift from how interpretability work usually goes. Probing and standard SAEs treat a model's knowledge as a pile of separate switches. This result suggests that at least some models have internalized music theory's own relational logic, transposition, scale membership, without being told to do so.

Two models and a method that builds its own assumption, pitch transposition, into the search is a modest sample size for a sweeping claim about how AI "understands" music. Worth watching whether this generalizes beyond pitch to rhythm or timbre, areas with no clean mathematical orbit to search for.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →