Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
Published 18 Sept 2026arXiv:2609.20081
Updated 4 h ago · first seen 18 Sept 2026
paper_01M2SEGH8FP28PGT5G25F0XD8K
Abstract
SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperReading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition: arxiv announce type changed from new to cross
Arxiv announce typenew→crossarxiv - New paperPaperReading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
New paper: Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.