Selection & Fusion

Does combining layers beat picking the right one? Per model and task: the best single layer (the reference row, in bold), label-free proxy-guided fusion over a few chosen layers, and trainable multi-layer fusion. Small green/grey numbers show the gap to the best single layer.

← Results table · Layer cheat sheet · Layer explorer · Correlations · Paper · results as of 2026-08-04

As published retrospective protocol

Across the paper’s models and tasks: probing only the top-3 proxy-ranked layers comes within 0.4 pp of the best single layer (matching it on 58% of model–task pairs); averaging those three layers lands at 0.3 pp and beats every trainable fusion method (p < 0.001). The per-task proxy metric was chosen retrospectively on the same models (see the paper).

Held out non-circular

Fixing each task family’s proxy on all other models and applying it to a held-out model still places the true best layer in the proxy’s top-3 39.6% of the time (chance 15.1%; top-1: 16.7% vs 5.0%), and 41.7% for tonal tasks via the pitch-transposition equivariance measure.
MethodMERT-v1-95MMERT-v1-330MMusicFM (MSD)MuQ (iter)OMAR-RQ (base)MusicGen-SMusicGen-MMusicGen-LYuE-s1-0.5BMyna-Base
Reference
Oracle (best single layer)63.465.264.664.562.268.665.567.266.647.8
Middle layer61.164.364.361.955.666.355.254.154.347.8
Last layer61.058.462.258.051.163.263.464.966.646.7
Proxy-guided fusion (label-free layer choice)
Top-3 avg61.8-1.663.6-1.659.5-5.164.4-0.166.2+4.067.8-0.865.8+0.367.0+0.445.5-2.3
Top-3 concat62.8-0.663.7-1.558.5-6.163.3-1.265.3+3.166.9-1.763.5-2.065.0-1.645.1-2.7
Top-5 avg62.1-1.365.2+0.060.8-3.863.7-0.866.6+4.467.7-0.966.6+1.168.2+1.067.4+0.847.0-0.8
Top-5 concat60.4-3.061.8-3.561.3-3.361.2-3.360.2-2.066.4-2.263.9-1.666.5-0.762.5-4.146.7-1.1
All-layer fusion (non-trainable)
All-layer avg64.8+1.465.4+0.263.5-1.164.5+0.063.3+1.168.0-0.663.7-1.865.8-1.568.0+1.446.7-1.1
All-layer concat59.7-3.726.7-38.561.0-3.663.3-1.261.7-0.566.0-2.662.3-3.262.4-4.861.8-4.840.6-7.2
Trainable fusion
Weighted sum62.0-1.462.9-2.362.2-2.363.1-1.464.8+2.666.2-2.465.6+0.165.6-1.664.4-2.247.6-0.2
HConv60.6-2.860.1-5.161.8-2.863.0-1.563.8+1.662.9-5.762.4-3.160.6-6.661.4-5.248.0+0.2
Attentive57.2-6.261.5-3.759.8-4.861.7-2.860.6-1.662.0-6.658.1-7.460.2-7.061.7-4.939.4-8.4