ConceptioArchivearXiv CS
arXiv CSopen access

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models

arXiv:2605.22732v1 [cs.AI] 21 May 2026

Jürgen Dietrich0000-0002-5494-3499 Democracy Intelligence gGmbH, Germany [email protected] May 2026 Abstract We investigate whether acoustic emotion recognition models can serve as proxies for the Pathos dimension in political speech analysis, as operationalised by the TRUST multiagent large language model (LLM) pipeline. Using a Bundestag plenary speech by Felix Banaszak (51 segments, 245 s) as a case study, we compare three analysis modalities: (1) emotion2vec_plus_large, an acoustic speech emotion recognition (SER) model whose continuous Arousal and Valence values are derived via post-hoc Russell Circumplex projection; (2) Gemini 2.5 Flash, an LLM analysing the full speech audio together with its transcript in an open-ended, context-aware fashion; and (3) TRUST-Pathos scores from a three-advocate LLM supervisor ensemble. Spearman rank correlations reveal that Gemini Valence correlates strongly with TRUST-Pathos (ρ = +0.664, p < 0.001), whereas emotion2vec Valence does not (ρ = +0.097, p = 0.499). We further demonstrate, via a systematic quality evaluation of the Berlin Database of Emotional Speech (EMO-DB) using Gemini in an open-ended annotation paradigm, that standard SER benchmark corpora suffer from acted speech, cultural bias, and category incompatibility. Our results suggest that LLM-based multimodal analysis captures semantically defined political emotion substantially better than acoustic models alone, while acoustic features remain informative for low-level Arousal estimation. Future work will extend this approach to video-based analysis incorporating facial expression and gaze. Keywords: speech emotion recognition, political communication, pathos analysis, large language models, emotion2vec, multimodal analysis, EMO-DB, Russell Circumplex

1

1

Introduction

The analysis of emotional expression in political discourse has received increasing attention from communication scholars, political scientists, and computational linguists. Within the TRUST framework [1, 2], political statements are evaluated along three rhetorical dimensions inspired by Aristotelian rhetoric: Logos (logical argumentation), Ethos (credibility), and Pathos (emotional appeal). While Logos and Ethos can be assessed through structured fact-checking and credibility scoring, Pathos poses a particular challenge: it requires the recognition of affective states embedded not only in lexical content but also in prosody, rhythm, and rhetorical strategy. Automatic SER has advanced substantially in recent years, with self-supervised models such as wav2vec 2.0 [3] and dedicated emotion encoders like emotion2vec [4] achieving competitive performance on standard benchmarks. However, these benchmarks are dominated by acted corpora—most prominently the Interactive Emotional Dyadic Motion Capture database (IEMOCAP) [5] and EMO-DB [6]—whose ecological validity for naturalistic political speech remains unexplored. In parallel, LLMs with multimodal capabilities have begun to demonstrate remarkable competence in understanding emotional nuance, irony, and rhetorical intent from audio and text. This raises the question of whether LLM-based analysis can serve as a more valid proxy for the politically relevant Pathos dimension than acoustic SER. This paper makes three main contributions: 1. We compare emotion2vec Arousal/Valence, Gemini multimodal Arousal/Valence, and TRUSTPathos scores on 51 segments of a naturalistic Bundestag (German Federal Parliament) speech, demonstrating that LLM-based emotion analysis correlates substantially better with LLM-based Pathos scoring than acoustic SER. 2. We provide a systematic quality evaluation of EMO-DB using Gemini as an open-ended, non-forced-choice annotator, revealing structural limitations of this widely used corpus, including undocumented discrepancies in gender coding and sentence transcription (see Appendix A). 3. We introduce the concept of post-hoc Russell Circumplex projection as an operationalisation of continuous Arousal/Valence from discrete SER class probabilities, and discuss its assumptions and limitations. The remainder of this paper is structured as follows. Section 2 reviews related work. Section 3 describes our methods and data. Section 4 presents all results, comprising the EMO-DB quality evaluation (Section 4.1) and the Banaszak speech analysis (Section 4.2). Section 5 discusses implications and limitations. Section 6 concludes.

2

Related Work

This section situates our work within three research streams: automatic SER, multimodal LLMbased affect analysis, and computational political communication.

2.1

Speech Emotion Recognition

The field of SER has been shaped by a small number of acted corpora. EMO-DB [6], recorded in 1999 at the Technische Universität Berlin, comprises 535 utterances from ten professional actors (five female, five male) covering seven emotion categories across ten fixed German sentences. Despite its age, EMO-DB remains a standard benchmark. The EmoBox framework [7], which evaluates ten pre-trained speech models across 32 datasets in 14 languages, reports that WavLM

2

Large achieves 92.67 % weighted accuracy (WA) on EMO-DB, while wav2vec 2.0 base reaches 83.14 % WA. The emotion2vec model family [4] employs self-supervised pre-training on 262 hours of opensource emotional speech data (internally termed Emo-262 ), followed by fine-tuning. The composition of Emo-262 in terms of language, culture, and acted vs. naturalistic speech is not publicly documented, constituting a known limitation for cross-cultural application. Text-independent SER evaluation has been identified as substantially harder than speakerindependent evaluation [8]; the gap is typically 10–20 percentage points in WA. Crucially, EMODB uses ten fixed sentences spoken by all actors, making genuine text-independent evaluation structurally impossible: any test sentence has appeared during training.

2.2

LLM-Based Affect Analysis

Recent work has explored LLMs as emotion annotators, with findings suggesting that LLMs match or exceed crowdsourced human annotation on categorical emotion tasks [9]. Gemini’s multimodal architecture enables joint processing of audio and text, allowing the model to integrate prosodic, semantic, and contextual cues simultaneously. Unlike forced-choice paradigms used in most SER corpora, open-ended LLM annotation avoids demand characteristics—the tendency of annotators to select one of the presented options even when none fits well, leading to artificially inflated agreement with predefined categories.

2.3

Computational Political Communication

The TRUST pipeline [1, 2] operationalises Pathos as the societal impact of emotional language: from +2 (unifying across party lines) through 0 (neutral or group-internal) to −2 (actively divisive). This definition differs substantially from the valence-arousal circumplex used in affective computing, raising the question of which computational approach better approximates TRUSTPathos.

3

Methods

This section describes the data, models, and analysis pipeline used in this study. All components are part of the TRUST Multimodal Pipeline (v1.0), an open research system developed at Democracy Intelligence gGmbH.

3.1

Speech Data

We analyse a Bundestag plenary speech delivered by Felix Banaszak (Bündnis 90/Die Grünen, co-chair), recorded on 5 March 2026 during the 62nd session of the 21st German Bundestag (agenda item ZP 3/4, energy policy). Banaszak was selected as a representative case of oppositional parliamentary rhetoric in the current German legislative period: as co-chair of a party transitioning from government to opposition, his speech combines high emotional intensity with complex rhetorical strategies (sarcasm, irony, appeal), providing a demanding test case for emotion recognition systems. The study is intentionally limited to a single speaker to enable controlled comparison of modalities; multi-speaker generalisation is addressed in ongoing work. The speech was retrieved as a Full HD video stream (1920×1080, 8000 kbps, H.264) from the official Bundestag media library [10] and converted to a mono 16 kHz WAV file using FFmpeg, excluding the first 12 seconds of procedural opening. Total duration: 232 s. The speech was segmented into 51 utterances using WhisperX [11] with pyannote speaker diarization [12]. Segment boundaries were determined by pause-based and syntactic criteria, yielding utterances of typically 3–15 seconds. 3

3.2

TRUST-Pathos Scoring

Each segment was submitted to the TRUST pipeline in API mode. TRUST employs three advocate LLMs—gemini-2.5-flash (critical), gpt-5.2 (balanced), and claude-sonnet-4-6 (benevolent)— whose Pathos scores are aggregated by a supervisor LLM using median consensus. Pathos scores are integers on a five-point scale: {−2, −1, 0, +1, +2}, where −2 denotes actively divisive and +2 denotes societally unifying emotional language. Of the 51 segments, 10 were excluded by the TRUST relevance filter (procedural utterances, greetings, and closing statement), leaving 41 segments with valid Pathos scores.

3.3

Acoustic Emotion Analysis: emotion2vec

Acoustic emotion features were extracted using emotion2vec_plus_large (FunASR implementation [13]) at utterance granularity. The model outputs class probabilities over eight categories: angry, disgusted, fearful, happy, neutral, other, sad, surprised. Post-hoc Russell Circumplex Projection. emotion2vec does not natively output continuous Arousal or Valence values. We derive these via a weighted sum of class probabilities, using weights adapted from Russell [14] and Warriner et al. [15]: X (1) Arousal = pk · wkA k

Valence =

X

pk · wkV

(2)

k

where pk is the predicted probability for class k and wkA , wkV are the Arousal and Valence weights listed in Table 1. This projection rests on three unverified assumptions: (1) Russell weights transfer to German language; (2) they apply to naturalistic political speech; (3) emotion2vec categories map onto Circumplex dimensions. None of these assumptions has been empirically validated. Table 1: Arousal and Valence weights for post-hoc Russell Circumplex projection of emotion2vec class probabilities. Sources: Russell [14]; Warriner et al. [15]. wA

wV

0.75 0.60 0.80 0.65 0.00 0.10 −0.30 0.70

−0.75 −0.80 −0.65 0.90 0.00 0.00 −0.85 0.20

Class angry disgusted fearful happy neutral other sad surprised

Abbreviations: wA = Arousal weight; wV = Valence weight.

3.4

LLM-Based Multimodal Analysis: Gemini

We submitted the full speech audio together with the complete 51-segment transcript (including segment IDs and timestamps) to Gemini 2.5 Flash (model: gemini-2.5-flash) via the Google GenAI API (v1.74.0). The system prompt instructed the model to evaluate each segment in terms of (a) primary and secondary emotion (open-ended, no forced choice), (b) Arousal on [−1, +1], (c) Valence on [−1, +1], (d) rhetorical function (open-ended), and (e) confidence on 4

[0, 1]. No predefined emotion categories were supplied—the model named emotions freely. This open-ended paradigm avoids the demand characteristics inherent in forced-choice annotation.

3.5

Statistical Analysis

Correlations between Arousal/Valence estimates from different modalities and TRUST-Pathos scores were computed using the Spearman rank correlation coefficient (ρ), appropriate for the ordinal TRUST-Pathos scale. Statistical significance was assessed at α = 0.05.

4

Results

This section reports all empirical findings in two parts. Section 4.1 presents the EMO-DB quality evaluation, establishing Gemini’s annotation behaviour on acted German speech and identifying structural corpus limitations. Section 4.2 presents the main comparative analysis on the Banaszak plenary speech.

4.1

EMO-DB Quality Evaluation

We evaluate Gemini as an open-ended annotator on all 535 EMO-DB utterances. This serves two purposes: it characterises Gemini’s emotion recognition on acted German speech, and it reveals structural limitations of this benchmark corpus. Appendix A provides the complete Speaker×Emotion matrix, which also documents discrepancies between the published EMO-DB documentation [6] and the actual corpus files, including differences in gender coding and sentence transcription found during our manual listening evaluation. 4.1.1

Corpus Structure

EMO-DB comprises 535 utterances from ten speakers (six female: IDs 03, 08, 09, 10, 11, 12; four male: IDs 13, 14, 15, 16) across seven emotion categories (Anger, Boredom, Disgust, Fear, Happiness, Neutral, Sadness) and ten fixed German sentences. The distribution is imbalanced: Anger is the most frequent class (n = 127, 23.7 %), Disgust the least frequent (n = 46, 8.6 %). Speaker 08 produced no Disgust utterances, and speaker 09 produced only one Fear utterance, creating systematic gaps in the Speaker×Emotion matrix. 4.1.2

Gemini Open-Ended Annotation

Each WAV file was submitted individually to Gemini without predefined category options. Gemini returned a primary emotion label, a secondary label (or null), confidence ([0, 1]), recording quality ([0, 1]), and a brief acoustic justification. Ground-truth (GT) matching used semantic mapping: e.g., Sachlichkeit (factuality) was mapped to Neutral, Verärgerung (annoyance) to Anger. Table 2 reports match rates per emotion category. 4.1.3

Key Findings

Three findings from Table 2 merit discussion. Disgust: 0.0 % match. Gemini consistently fails to identify Disgust as a distinct category, labelling these utterances as annoyance, contempt, or resignation. This suggests that Disgust is not acoustically distinguishable without visual context (facial expression). Boredom: 12.3 % match. Boredom is systematically misidentified as Neutral or factual speech. Notably, Gemini’s confidence is high (0.81) when mislabelling Boredom utterances— confident but wrong. Manual listening by the author confirmed that Boredom is readily iden-

5

Table 2: Gemini open-ended annotation results on EMO-DB (n = 535). Match rates reflect semantic matching between Gemini primary labels and EMO-DB GT. Emotion

n

Match (%)

Avg. Conf.

Neutral Sadness Happiness Anger Fear Boredom Disgust

79 62 71 127 69 81 46

65.8 35.5 29.6 29.1 27.5 12.3 0.0

0.83 0.80 0.83 0.86 0.77 0.81 0.81

Total

535

30.1

0.82

Abbreviations: n = number of utterances; Avg. Conf. = mean Gemini confidence score; GT = ground truth. tifiable to human listeners, suggesting a representation gap for low-arousal German speech in Gemini’s training data. High confidence, low match. The overall pattern—mean confidence 0.82 yet only 30.1 % semantic match—indicates that Gemini’s confidence is a poor predictor of correctness. This is consistent with a conceptual mismatch between EMO-DB’s forced-choice taxonomy and Gemini’s open-ended vocabulary. 4.1.4

Text Independence

EMO-DB uses ten fixed sentences spoken by all actors, making genuine text-independent evaluation structurally impossible. Our manual evaluation further revealed that emotion-specific prosodic patterns are confounded with sentence identity: Sadness is consistently produced with longer pauses across all speakers, while Anger is produced staccato. A model can exploit these text-specific rhythmic cues without learning emotion universals, inflating text-dependent accuracy estimates.

4.2

Banaszak Speech Analysis

This section presents the comparative analysis on the 51-segment plenary speech [10]. We report descriptive statistics, Spearman correlations, temporal dynamics, and Gemini’s rhetorical classification. 4.2.1

Descriptive Statistics

Table 3 summarises Arousal and Valence estimates from both modalities. Gemini assigns substantially higher Arousal (mean = 0.59) than emotion2vec (mean = 0.36) and strongly negative Valence (mean = −0.56), whereas emotion2vec Valence is near zero (mean = 0.04). TRUSTPathos scores cluster at −1 (n = 18) and 0 (n = 22), with one segment at −2 and one at +1. 4.2.2

Correlation Analysis

Table 4 reports Spearman rank correlations between all modality pairs. The central finding is a strong and significant correlation between Gemini Valence and TRUST-Pathos (ρ = +0.664, p < 0.001), and a moderate negative correlation between Gemini Arousal and TRUST-Pathos (ρ = −0.535, p < 0.001). By contrast, emotion2vec Valence shows no significant association 6

Table 3: Descriptive statistics for Arousal, Valence, and TRUST-Pathos across 51 segments of the Banaszak speech [10]. Measure

Mean

SD

Min

Max

Gemini Arousal Gemini Valence emotion2vec Arousal emotion2vec Valence TRUST-Pathos

0.59 −0.56 0.36 0.04 −0.37

0.28 0.44 0.21 0.32 0.56

0.00 −1.00 0.04 −0.74 −2.00

1.00 0.60 0.75 0.78 1.00

Abbreviations: SD = standard deviation; TRUST = Transparent Rhetorical Understanding and Scoring Tool. with TRUST-Pathos (ρ = +0.097, p = 0.499), and emotion2vec Arousal correlates weakly and non-significantly (ρ = −0.155, p = 0.278). Cross-modal agreement is also low: emotion2vec and Gemini Arousal correlate at ρ = +0.239 (p = 0.091), and Valence at ρ = +0.200 (p = 0.159), confirming that the two approaches capture substantially different dimensions. Segment-level scores are listed in Appendix B. Table 4: Spearman rank correlations between emotion modalities and TRUST-Pathos for 51 segments of the Banaszak speech [10]. Comparison Gemini Valence ↔ TRUST-Pathos Gemini Arousal ↔ TRUST-Pathos e2v Valence ↔ TRUST-Pathos e2v Arousal ↔ TRUST-Pathos e2v Arousal ↔ Gemini Arousal e2v Valence ↔ Gemini Valence

ρ

p

0.664 −0.535 0.097 −0.155 0.239 0.200

<0.001 <0.001 0.499 0.278 0.091 0.159

Abbreviations: ρ = Spearman rank correlation coefficient; e2v = emotion2vec; TRUST = Transparent Rhetorical Understanding and Scoring Tool. Bold rows indicate p < 0.001; remaining rows p ≥ 0.05 (not significant). 4.2.3

Temporal Dynamics

Figure 1 illustrates the temporal evolution of Gemini Valence, emotion2vec Arousal, and TRUSTPathos across the 51 segments. Gemini Valence tracks the rhetorical arc of the speech closely, remaining strongly negative throughout the main body (segments 6–47) and returning to neutral only for the closing statement (segment 49). emotion2vec Arousal shows high variability with no discernible correspondence to TRUST-Pathos. The single positive TRUST-Pathos segment (s0042: “Es gibt einen Morgen—”, “There is a tomorrow—”) coincides with the only positive Gemini Valence value in the body of the speech, confirming that Gemini captures this rhetorical turn correctly. Note that e2v Valence and Gemini Arousal are omitted from the figure for clarity; their temporal profiles show no correspondence to TRUST-Pathos (Table 4). 4.2.4

Rhetorical Analysis

Gemini’s open-ended rhetorical classification reveals a distribution consistent with oppositional parliamentary discourse: Criticism (n = 16, 31 %), Sarcasm (n = 14, 27 %), None (n = 9, 18 %), Appeal (n = 7, 14 %), Metaphor (n = 2, 4 %), Accusation (n = 1, 2 %), Rhetorical Question (n = 1, 2 %), and Indignation (n = 1, 2 %). This taxonomy—which emerged without prede-

7

Score

1

0

−1 0

5

10

15

20

25 30 Segment index

Gemini Valence

e2v Arousal

35

40

45

50

TRUST Pathos

Figure 1: Temporal profiles of Gemini Valence, emotion2vec (e2v) Arousal, and TRUST-Pathos across 51 segments of the Banaszak speech [10]. Gemini Valence tracks the rhetorical arc; e2v Arousal reflects acoustic energy independently of TRUST-Pathos. e2v Valence and Gemini Arousal are omitted for clarity; both show no significant correlation with TRUST-Pathos (Table 4). Abbreviations: e2v = emotion2vec; TRUST = Transparent Rhetorical Understanding and Scoring Tool. fined categories—captures the rhetorical profile of an opposition speech targeting the governing coalition, and corresponds to the negative Pathos cluster observed in TRUST scoring.

5

Discussion

Our results clarify why acoustic SER is insufficient as a Pathos proxy in political communication analysis, and suggest a path forward.

5.1

Two Different Constructs

As shown in Table 4, the near-zero correlation between emotion2vec Valence and TRUST-Pathos (ρ = +0.097, p = 0.499) is not a failure of emotion2vec—it reflects the fact that the two measures capture different constructs. emotion2vec captures acoustic Valence: the emotional colouring inferable from voice quality, fundamental frequency (F0) contour, and spectral features. TRUSTPathos captures political-rhetorical Valence: the societal impact of emotional language, including irony, sarcasm, and rhetorical strategy accessible only through semantic understanding. The segment “Das ist wirklich peinlich” (“That is truly embarrassing”) illustrates this gap: emotion2vec assigns high positive Valence (+0.74) and classifies it as happy based on acoustic energy. Gemini correctly identifies the utterance as strong disapproval (Valence −0.90). TRUST assigns Pathos = 0, reflecting that the statement is evaluatively charged but directed at a specific target rather than broadly divisive. More broadly, this points to a rhetorical strategy of decoupling: a speaker may deploy high prosodic activation to frame a statement emotionally while delivering the propositional core in a factually neutral register—a pattern that acoustic models cannot detect but LLMs can identify through semantic understanding.

5.2

The Post-Hoc Projection Problem

Equations (1) and (2) operationalise a theoretically motivated but empirically unvalidated mapping. The Warriner et al. [15] norms were derived from English word ratings, not spoken German

8

political discourse. The Russell Circumplex [14] is a model of subjective affect, not of acoustic signal properties. Future work should empirically validate or replace this projection using manual Arousal/Valence annotations of the target corpus.

5.3

EMO-DB Limitations

Our evaluation reveals three concerns beyond those previously documented. First, Disgust is acoustically unidentifiable without visual context. Second, Boredom is systematically misclassified by Gemini despite being readily identifiable to human listeners. Third, emotion-specific prosodic conventions create text-specific cues that inflate text-dependent accuracy estimates.

5.4

Limitations

This study has several limitations. First, n = 51 segments from a single speaker limits statistical power; we plan to extend the corpus to additional speakers and speech types. Second, Gemini Arousal and Valence are self-estimated scalar values derived from the same model that produces rhetorical labels, creating potential internal consistency effects. Third, TRUST-Pathos scores reflect a specific operationalisation of political emotion developed and evaluated on German political statements [1]. Whether this operationalisation generalises to other political systems, languages, or speech genres remains an open empirical question. Fourth, the training data composition of emotion2vec is not publicly documented.

6

Conclusion

We have shown that LLM-based multimodal emotion analysis substantially outperforms acoustic SER as a proxy for TRUST-Pathos in naturalistic political speech. The key mechanism is semantic-pragmatic understanding: Gemini integrates lexical content, rhetorical structure, and political context, whereas emotion2vec responds primarily to acoustic signal properties. Both modalities capture real information—but about different dimensions of political communication. These results have practical implications for the TRUST pipeline and for computational political communication analysis more broadly. Rather than replacing acoustic analysis, LLM-based emotion scoring should be used in conjunction with acoustic features, creating a complementary multimodal representation of political affect. Future work will extend this analysis to a multispeaker corpus, evaluate fine-tuned SER models (wav2vec 2.0 trained on quality-filtered EMO-DB and PAVOQUE), investigate whether such a model better approximates TRUST-Pathos than the base acoustic model, and compare multiple LLM providers (e.g., Gemini, GPT, Claude) as emotion annotators to assess whether the observed correlation with TRUST-Pathos is model-specific or reflects a more general property of semantically informed emotion analysis. A natural further extension is the incorporation of video analysis—combining facial expression recognition (Action Units via OpenFace), gaze estimation (L2CS-Net), and body posture tracking (MediaPipe)—to capture the full multimodal dimension of political communication.

Acknowledgements The author thanks Demian Frister (Democracy Intelligence gGmbH) for his critical review and valuable comments on an earlier version of this manuscript.

9

References [1] Jürgen Dietrich. From safety risk to design principle: Peer identity bias in multi-agent LLM systems for political statement analysis. arXiv preprint, 2026. arXiv:2604.08465. [2] Jürgen Dietrich. When roles fail: Epistemic constraints on advocate role fidelity in LLMbased political statement analysis. arXiv preprint, 2026. arXiv:2604.27228. [3] Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems, volume 33, pages 12449–12460, 2020. [4] Ziyang Ma, Mingjie Zheng, Jiaxin Yin, Sirui Li, Xie Li, and Xie Chen. emotion2vec: Self-supervised pre-training for speech emotion representation. arXiv preprint, 2023. arXiv:2312.15185. [5] Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan. IEMOCAP: Interactive emotional dyadic motion capture database. Language Resources and Evaluation, 42(4):335– 359, 2008. [6] Felix Burkhardt, Astrid Paeschke, Miriam Rolfes, Walter Sendlmeier, and Benjamin Weiss. A database of German emotional speech. In Proceedings of Interspeech, pages 1517–1520, 2005. [7] Ziyang Ma, Mingjie Zheng, Xie Chen, et al. EmoBox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark. In Proceedings of Interspeech 2024, 2024. [8] Bagus Tris Atmaja and Akira Sasou. Towards text-independent emotion recognition. Sensors, 22(17):6682, 2022. [9] Md Hamjajul Amin et al. Will affective computing emerge from foundation models and multimodal learning? A first evaluation on ChatGPT. arXiv preprint, 2023. arXiv:2307.14555. [10] Deutscher Bundestag. Plenarprotokoll 21/62, videomitschnitt. Video-ID 7649676, 2026.

Bundestag-Mediathek,

[11] Max Bain, Jaesung Huh, Tengda Han, and Andrew Zisserman. WhisperX: Time-accurate speech transcription of long-form audio. arXiv preprint, 2023. arXiv:2303.00747. [12] Alexis Plaquet and Hervé Bredin. Powerset multi-class cross entropy loss for neural speaker diarization. In Proceedings of Interspeech 2023, 2023. [13] Zhifu Gao et al. FunASR: A fundamental end-to-end speech recognition toolkit. arXiv preprint, 2023. arXiv:2305.11013. [14] James A. Russell. A circumplex model of affect. Journal of Personality and Social Psychology, 39(6):1161–1178, 1980. [15] Amy Beth Warriner, Victor Kuperman, and Marc Brysbaert. Norms of valence, arousal, and dominance for 13,915 English lemmas. Behavior Research Methods, 45(4):1191–1207, 2013.

10

A

Appendix A: EMO-DB Speaker×Emotion Matrix

Table 5 shows the complete distribution of utterances across speakers and emotion categories in EMO-DB. The matrix reveals systematic gaps—most notably, speaker 08 produced no Disgust utterances—that complicate speaker-independent cross-validation. A manual review of 140 of the 535 files revealed discrepancies between the published EMO-DB documentation [6] and the actual corpus files, including differences in gender coding and sentence transcription. A complete verification against the original recording notes is recommended for future users of the corpus. Table 5: EMO-DB utterance counts per speaker and emotion category. Speaker and emotion distribution as derived from filename metadata. Speaker

G

Ang

Bor

Dis

Fea

Hap

Neu

Sad

Total

03 08 09 10 11 12 13 14 15 16

F F F F F F M M M M

14 12 13 10 11 12 12 16 13 14

5 10 4 8 8 5 10 8 9 14

1 0 8 1 2 2 8 8 5 11

4 6 1 8 10 6 7 12 8 7

7 11 4 4 8 2 10 8 6 11

11 10 9 4 9 4 9 7 11 5

7 9 4 3 7 4 5 10 4 9

49 58 43 38 55 35 61 69 56 71

127

81

46

69

71

79

62

535

Total

Abbreviations: G = gender (F = female, M = male); Ang = Anger; Bor = Boredom; Dis = Disgust; Fea = Fear; Hap = Happiness; Neu = Neutral; Sad = Sadness.

11

B

Appendix B: Banaszak Speech – Segment-Level Scores

Table 6 lists the 41 segments retained for analysis after applying the TRUST relevance filter. Ten segments were excluded as they constitute non-evaluable utterances: procedural address to the chair, closing statement, and moderator announcements. The complete speech is publicly available [10].

12

13

Text (abridged)

Die Klima- und Umweltbewegung. . . Über 200 000 Unterschriften In dieser Debatte. . . 200 000 Unterschriften. . . Lassen Sie mich. . . 200 000 Menschen. . . Die drei von der Tankstelle. . . Aber dass Jens Spahn. . . Aber das System. . . Das Habäcksche. . . Und Mirsch erzählt. . . Nur weil man. . . Heute wäre die Chance. . . Heute wäre die Chance. . . Solaranlage. . . Jetzt habt ihr. . . Also gibt es keine. . . Heute wäre die Chance. . . Sie lassen diese Chance. . . Katharina Reich hat. . . Ja, das würde ich. . . Stattdessen. . . Traumabew. . . Alles, was nach Habeck. . . Das ist doch nicht Freiheit Sie bringen Energiearmut. . . Seit 2022 wissen wir. . . Sie haben nichts gelernt Es gibt kein Verbot. . . Mit Verboten kenne ich. . . Es gibt doch kein Verbot Es gibt kein Gesetz. . . Und da ist die Wand. . . Jetzt stehen Sie vor. . . Das ganze Land steht. . . Es gibt einen Morgen. . . Robin Mesarosch. . . Als hätten Sie. . . Das ist Ihr Fraktions. . . Ich sage Ihnen. . . Zeigen Sie Größe. . . auf unsere Sicherheit. . .

ID

s0003 s0004 s0005 s0006 s0007 s0008 s0009 s0013 s0014 s0015 s0016 s0017 s0018 s0019 s0020 s0022 s0023 s0024 s0025 s0027 s0028 s0029 s0030 s0031 s0032 s0033 s0034 s0035 s0036 s0037 s0038 s0039 s0040 s0041 s0042 s0043 s0044 s0045 s0046 s0047 s0048

0.10 0.16 0.27 0.18 0.13 0.09 0.31 0.46 0.20 0.23 0.55 0.49 0.65 0.60 0.69 0.75 0.44 0.41 0.33 0.04 0.10 0.44 0.48 0.42 0.62 0.37 0.23 0.35 0.50 0.58 0.54 0.74 0.46 0.13 0.45 0.53 0.56 0.45 0.13 0.05 0.27

e2v-A

Gem-A 0.50 0.40 0.50 0.60 0.40 0.60 0.50 0.70 0.50 0.50 0.70 0.60 0.80 0.50 0.80 0.80 0.60 0.80 0.70 0.30 0.40 0.70 0.80 0.70 0.80 0.50 0.60 0.50 0.50 0.50 0.70 0.60 0.60 0.70 0.50 0.50 0.70 0.50 0.70 0.90 0.40

e2v-V −0.07 0.21 0.36 0.22 0.12 0.00 0.33 0.06 0.27 0.27 −0.04 −0.36 0.28 0.72 −0.29 −0.74 −0.38 −0.32 −0.03 0.02 0.05 0.30 −0.21 0.01 0.45 0.02 −0.07 −0.03 0.01 −0.02 0.14 −0.66 −0.25 −0.04 −0.30 0.35 −0.15 −0.23 0.13 0.02 0.33 0.20 0.10 0.00 −0.70 −0.20 −0.70 −0.20 −0.80 −0.30 −0.40 −0.80 −0.60 −0.90 −0.20 −0.90 −0.90 −0.50 −0.90 −0.70 −0.10 −0.20 −0.80 −0.90 −0.70 −0.90 −0.30 −0.60 −0.30 −0.40 −0.30 −0.60 −0.40 −0.50 −0.70 0.40 −0.20 −0.80 −0.30 −0.80 −0.80 0.10

Gem-V 0 0 0 −1 0 −1 0 −1 0 0 −1 0 −1 0 −1 −1 0 −1 −1 0 0 −1 −2 −1 −1 0 −1 0 0 0 −1 0 0 −1 1 0 −1 0 −1 −1 0

Pathos Entschlossenheit Sachlichkeit Begeisterung Verachtung Entschlossenheit Spott Sarkasmus Empörung Kritik Sarkasmus Kritik Frustration Anklage Vorwurf Empörung Sarkasmus Kritik Anklage Vorwurf Sachlichkeit Sarkasmus Verachtung Empörung Empörung Empörung Sachlichkeit Vorwurf Sarkasmus Sarkasmus Kritik Kritik Metapher Kritik Metapher Zuversicht Sarkasmus Sarkasmus Kritik Empörung Appell Appell

Gem-Emotion Appell keines Appell Kritik keines Sarkasmus Sarkasmus Empörung Kritik Sarkasmus Sarkasmus Kritik Kritik Kritik Kritik Sarkasmus Kritik Appell Kritik keines Sarkasmus Kritik Sarkasmus Rhet. Question Anklage Appell Kritik Sarkasmus Sarkasmus Kritik Kritik Metapher Kritik Metapher Appell Sarkasmus Sarkasmus Kritik Appell Appell Appell

Gem-Rhetoric

Table 6: Segment-level scores for the Banaszak speech [10], 41 analysed segments. Pathos scores are integers on {−2, −1, 0, +1, +2}. Abbreviations: e2v-A = emotion2vec Arousal; e2v-V = emotion2vec Valence; Gem-A = Gemini Arousal; Gem-V = Gemini Valence; Pathos = TRUST-Pathos score; Gem-Emotion = Gemini primary emotion (German); Gem-Rhetoric = Gemini rhetorical function (English).

Record · ID 216862 · SHA-256 87531481c76b444e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.