Findings
No universal proxy. Across 12 models, 3 pre-training paradigms and 15 MIR tasks, no single label-free property predicts layer quality everywhere.
Geometry works outside tonal tasks. Intrinsic dimension is the strongest general proxy (mean ρ = 0.76), and straighter temporal trajectories transfer better, most consistently for beat tracking.
Standard metrics fail on tonal tasks. Key, chord and pitch show weak, sign-inconsistent correlations: static geometry and augmentation-induced invariance are the opposite of what these tasks need.
PTE closes the gap. This self-supervised circle-of-fifths diagnostic is the only metric with a consistent tonal signal (mean ρ = 0.80), and the only one that holds both within and across models.
A few proxy-ranked layers are enough. The top-3 land 0.4 pp from the per-task oracle and match it on 58% of model–task pairs, using 4–16× fewer probe runs; the default final layer is 4.3 pp worse.
Metric-guided fusion beats trainable fusion. Averaging those three layers is the strongest fusion method (p < 0.001): 0.3 pp from the oracle, exceeding it in 41% of model–task pairs, while every trainable baseline trails, and the margin grows when labels are scarce.
Metric values do not rank models. Effective rank tracks hidden size (ρ = 0.82), not quality (ρ = 0.02): these metrics choose layers within a model, not winners across architectures.
Results by model, layer and task
The pages below give the per-layer evaluations behind the paper's aggregates, to help choose a model and a layer for a task. They cover the 12 models of the paper and models evaluated since under the same protocol (marked +); the raw records are available as JSON.
Abstract
Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi-layer fusion, with limited understanding of why certain layers transfer better across downstream tasks or how representation quality varies with depth and pre-training paradigm. We conduct a systematic layer-wise analysis of 12 music foundation models spanning three pre-training paradigms (masked modeling, autoregressive modeling, and contrastive learning), characterizing their hidden representations through intrinsic geometric and transformation-based properties. Correlating label-free representation-quality metrics with layer-wise performance across 15 downstream tasks, we find that several metrics track layer quality for genre classification, emotion recognition, automatic tagging, and beat tracking, albeit with varying strength across tasks and pre-training paradigms. However, all metrics fail on tonal tasks such as key estimation and chord recognition, indicating that no single property serves as a general proxy for representation quality across music information retrieval tasks. To address this gap, we introduce a pitch-transposition equivariance measure that captures properties missed by these standard metrics, providing a consistent indicator of tonal quality across model families. Finally, we show that intrinsic metrics can serve as effective proxies for layer selection, matching or outperforming trainable multi-layer fusion methods, particularly in limited-data settings.
BibTeX
@inproceedings{kanatas2026goodlayer,
title = {What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models},
author = {Kanatas, Angelos-Nikolaos and Kong, Yuexuan and Alonso-Jim{\'e}nez, Pablo and Serra, Xavier and Bogdanov, Dmitry},
booktitle = {Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)},
year = {2026}
}