Six frozen 24-layer speech encoders, eight datasets, five languages, and one question: when a probe reads emotion out of a pretrained model, is it hearing how something was said — or quietly reading what was said?
| Encoder | Pretraining | CREMA-D best layer | Fixed-lexicon acc. | Acc. loss, clean → 0 dB | Text bias (EMIS) |
|---|
Text bias is measured on utterances where the spoken words and the vocal delivery carry different emotions. A positive number means the probe followed the words. Whisper is the only encoder that stays with the voice — and it is also the most noise-robust. Supervised multilingual pretraining produces a different representation from self-supervised pretraining, at every depth we tested.