Where does emotion live in a speech transformer?

Six frozen 24-layer speech encoders, eight datasets, five languages, and one question: when a probe reads emotion out of a pretrained model, is it hearing how something was said — or quietly reading what was said?

Every number below is measured. Use the tabs to move between the three experiments that answer it.

Encoder scorecard

EncoderPretrainingCREMA-D best layer Fixed-lexicon acc.Acc. loss, clean → 0 dBText bias (EMIS)

Text bias is measured on utterances where the spoken words and the vocal delivery carry different emotions. A positive number means the probe followed the words. Whisper is the only encoder that stays with the voice — and it is also the most noise-robust. Supervised multilingual pretraining produces a different representation from self-supervised pretraining, at every depth we tested.

CSCI 535, USC · six encoders × eight datasets × five languages Project page · Repository · kelidari.com