Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy Serli Kopar ID 1,2 , Roshan Prakash Rane ID 1,2,3 , Christian Mychajliw ID 4,5 , Lydia Federmann ID 1,2 , Gerhard Eschweiler ID 4,5,6 , Daniela Berg ID 7,8 , Sam Gijsen ID 1,2,11 , Paula Andrea Perez-Toro ID 9 , Kerstin Ritter ID 1,2,10
arXiv:2605.27189v1 [cs.CL] 26 May 2026
1
Hertie Institute for AI in Brain Health, University of Tübingen, Tübingen, Germany 2 Tübingen AI Center, University of Tübingen, Tübingen, Germany 3 Department of Psychology, Humboldt-Universität zu Berlin 4 Geriatric Center, Tübingen University Hospital, Tübingen, Germany 5 Tübingen Center for Mental Health (TüCMH), Department of Psychiatry and Psychotherapy, Tübingen University Hospital, Tübingen, Germany 6 German Center for Mental Health (DZPG), Partner Site Tübingen, Tübingen, Germany 7 Department of Neurology, University Medical Center Schleswig-Holstein and Kiel University, Kiel, Germany 8 Center for Neurology, University Hospital Tübingen and Hertie Institute for Clinical Brain Research, Tübingen, Germany 9 Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany 10 Charité–Universitätsmedizin, Department of Psychiatry and Psychotherapy, Berlin, Germany [email protected], [email protected], [email protected]
Abstract This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment. Utilizing 5,754 German neuropsychological assessment recordings, we evaluate six cognitive tasks across three score levels: task, domain, and global levels. We compare hand-crafted acoustic features with self-supervised learning (SSL) embeddings. Results show that although SSL representations generally outperform hand-crafted features at lower levels, this trend reverses for MCI classification. Furthermore, task-specific constraints influence performance: tasks with greater response freedom exhibit performance dilution as hierarchical levels increase, suggesting “specialist” representations, whereas the performance of highly structured tasks increases toward higher levels, suggesting “generalist” representations. These findings show links between task constraints and assessment hierarchy in automated clinical speech analysis. Index Terms: hierarchical cognitive assessment, mild cognitive impairment, neuropsychological test battery, clinical speech analysis
izability [8]; and (iii) modeling clinical scores as independent, flat targets, thereby ignoring the hierarchical structure inherent to standardized cognitive assessment. In this paper, we address these gaps using 5,754 recordings collected during five CERAD+ tasks and one MMSE screening task from an elderly German cohort. Beyond binary classification, we model the hierarchical organization of clinical scores, linking acoustic features to task-, domain-, and globallevel scores. This enables a fine-grained analysis of the relationship between speech and cognitive decline across multiple diagnostic and screening tasks. Our analysis reveals a taskdependent pattern: for tasks with more open-ended responses, the predictive power of acoustic features decreases from tasklevel to domain-level and global scores, whereas for more constrained tasks, it increases across these levels. To the best of our knowledge, this is the first study to examine how acoustic feature predictiveness varies across hierarchical levels of the German CERAD+ battery in the context of MCI. To support future research, we publicly release our code.1
2. Methods 1. Introduction Mild cognitive impairment (MCI) is a clinical syndrome characterized by cognitive decline exceeding normal aging [1, 2]. As a prodromal stage of dementia, most commonly Alzheimer’s disease, it represents a critical window for early intervention [3]. Despite its clinical relevance, MCI remains substantially underdiagnosed [4]. Clinical diagnosis typically relies on standardized neuropsychological assessments. The Consortium to Establish a Registry for Alzheimer’s disease (CERAD+) is a wellvalidated, multi-domain neuropsychological battery assessing language, memory, executive function, and visuospatial abilities [5], generating structured task-, domain-, and global-level scores. In routine practice, shorter instruments such as the MiniMental State Examination (MMSE) are frequently used for efficient screening. Together, these instruments form the backbone of clinical cognitive assessment. Automated speech analysis has emerged as a promising complement to traditional testing [6]. However, current approaches face three methodological bottlenecks: (i) focusing on binary classification (Alzheimer’s Disease vs. healthy controls (HC)), which lacks sensitivity to the subtle cognitive change characteristics of MCI [7]; (ii) reliance on English-centric, single-task datasets, limiting general-
2.1. Dataset and Quality Control We used speech recordings from the TREND study2 [12], comprising one MMSE screening task and five CERAD+ diagnostic tasks: Word List Recognition (RW), Boston Naming Test (BNT), Word List Recall (RL), Verbal Fluency (VF), and Phonemic Fluency (PF). Our analysis follows the inherent three-level hierarchy of clinical assessment, as summarized in Fig. 1: Level 1 (Individual Tests) comprises raw tasklevel scores (e.g., PF: number of valid words beginning with “S”). Level 2 (Cognitive Domains) aggregates these task-level scores into domain-level composite measures [9], including Language (LAN), Memory (MEM), Executive Function (EXE), and Visuospatial Ability (VIS). These domains differ in their reliance on verbal versus drawing-based tasks (LAN: 1:0; MEM: 1:1; EXE and VIS: 0:1). Although EXE and VIS are traditionally assessed through drawing-based tasks, we evaluate the cross-domain predictive performance of speech by using acoustic features from verbal tasks to predict these non-verbal tar1 https://github.com/anon-interspeech/anon-interspeech-2026.git 2 All participants provided written informed consent. The study was approved by the Ethics Committee of (anonymized)
Figure 1: Our workflow for predicting hierarchical cognitive scores from speech. Using speech-derived acoustic features, independent models predict targets at three levels: (1) Level 1 (Individual Tests): task-level scores (e.g., Phonemic Fluency (PF), the number of valid words beginning with “S”, with scores ranging from [0, ∞)); (2) Level 2 (Cognitive Domains): domain-level composite scores with speech-to-drawing task ratios of 1:0 (Language, LAN), 1:1 (Memory, MEM), and 0:1 (Executive, EXE; Visuospatial, VIS) [9]; and (3) Level 3 (Global Status): global-level scores including the CERAD+ total score (modeled as continuous and binary with a threshold of 85) [10] and binary MCI status (defined as more than 1.5 standard deviations below the normative cohort) [11]. Arrows denote workflow from shared acoustic feature representations to independent models predicting hierarchical targets.
gets. This tests whether speech contains information that generalizes beyond verbal cognitive domains (MEM, LAN). Level 3 (Global Status) includes the global-level scores: CERAD+ total score (continuous and thresholded at 85) [10] and clinical MCI status (> 1.5 SD below demographically adjusted normative cohort) reported by Berres et al. [11]. To ensure diagnostic and acoustic integrity, we excluded nonnative speakers, incomplete profiles, and MCI-to-HC reverters. Additionally, we performed acoustic quality assessment, enforcing constraints on duration (minimum 15 s), energy (RMS > −55 dBFS), digital clipping (< 1.5%), and signal-to-noise ratio (SNR > 10 dB), estimated using reference-free, quantilebased methods [13, 14, 15]. Conditional inconsistencies between metrics (e.g., high speech activity ratio with low SNR) were manually reviewed. These filtering steps yielded 959 sessions (698 HC, 261 MCI) from 593 participants. 2.2. Preprocessing and Diarization Optimization To find the optimal preprocessing hyperparameters, we used a manually transcribed ground-truth subset (N = 89; ≈ 9% of the corpus, available only in two fluency tests: Phonemic (PF) and Verbal (VF)). We conducted a grid search of > 2,500 combinations using participant-disjoint tuning and validation splits. Hyperparameter tuning was based on diarization and Jaccard error rates (DER, JER), purity (PUR), and coverage (COV), with a 250 ms collar [16]. The resulting pipeline applied a 6th-order Butterworth high-pass filter (fc = 100 Hz) [17], spectral-gating noise suppression (α = 0.3) [18], and loudness normalization (−23 LUFS) [19]. The final configuration achieved 0.20 DER, 0.33 JER, 94% PUR and 97% COV on the participant-disjoint validation split. To enable high-density voice-quality features, we generated two audio streams: Prosody-Preserved (examiner masked, temporal structure retained) and Concatenated (participant segments merged using 10 ms linear cross-fades). Both streams were manually audited to ensure transition integrity.
puted prosodic features (EG Prosody) from the ProsodyPreserved stream to retain conversational timing and voicequality features (EG V-Qual) from the Concatenated stream to enable high-density voice-quality features. We combined them into a third, unified EG All feature set. In addition, following recent work on dementia detection with semantic and phonemic fluency tasks [21, 22, 23], we extract latent representations from the frozen final hidden layers of wav2vec 2.0 (W2V2; facebook/wav2vec2-base-960h) and HuBERT (facebook/hubert-large-ls960-ft) using global mean pooling from Prosody-Preserved stream. To validate our findings in an independent participant sample, we split the dataset into subject-disjoint development and holdout sets. We then confirmed that these sets were comparable (Tab. 1) within both HC and MCI by testing differences in hierarchical scores, age and sex (chi-square (χ2 ) tests for categorical variables and t-tests for continuous variables). No significant differences were observed (p > 0.05). Table 1: Demographic statistics and cognitive assessment scores for the development and hold-out sets. Values are presented as Mean (± standard deviation (SD)) or Count (%). Development (N=772)1 Feature
HC
Subjects (N) Age (years) Sex (% Female) Phonemic Fluency (PF) MMSE Language Domain (LAN) CERAD+ (Total) Avg. Rec.2 1
Hold-out (N=187)
MCI
359 115 73.1 ± 6.1 74.9 ± 5.8 53.2% 36.5% 14.8 ± 4.6 13.1 ± 4.8 28.3 ± 1.5 27.6 ± 2.0 0.78 ± 0.1 0.74 ± 0.1 85.7 ± 8.2 79.5 ± 10.6 1.7 ± 0.6 1.4 ± 0.5
HC
MCI
88 31 73.0 ± 6.0 74.9 ± 6.8 55.7% 35.5% 15.0 ± 4.2 13.3 ± 4.6 28.3 ± 1.4 27.6 ± 2.0 0.78 ± 0.1 0.72 ± 0.1 85.3 ± 8.1 77.5 ± 11.6 1.6 ± 0.6 1.5 ± 0.6
Number of recordings. 2 Average recordings per participant.
2.4. Prediction and Validation Framework 2.3. Feature Extraction From these streams, we extracted the extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) [20]. We com-
To evaluate prediction performance on the development set, we employed 5 × 3 nested cross-validation (NCV). For each hierarchical target, task, and feature set, models were trained from
scratch with strictly subject-disjoint folds, ensuring that no participant appeared in multiple folds or across training and validation/test sets. The pipeline incorporated z-score normalization and PCA variance thresholding (including a passthrough option) within the inner-loop grid search. We evaluated Ridge regression, support vector machines (SVM for classification; SVR for regression), and extreme gradient boosting (XGBoost). Inner-loop optimization targeted balanced accuracy for classification and R2 for regression. After NCV, the best-performing model architecture for each target was selected based on mean outer-fold performance, and only these results are reported for the development set. To determine the final configuration of model hyperparameters to be tested on a disjoint hold-out set, we used a majority vote across NCV folds. This model was then retrained from scratch on the full development set and evaluated on the hold-out set to verify generalization to unseen participants. All models were implemented in Python using scikit-learn (v1.8.0) [24] and xgboost (v3.1.2) [25].
3. Results 3.1. Level 1: Task-Level Score Prediction
MMSE
0.11
0.15
0.16
0.19
0.33
0.8
RW
0.07
0.11
0.14
0.11
0.26
0.7
BNT
0.15
0.22
0.24
0.25
0.46 ± 0.08
0.5
RL
0.30
0.45
0.50
0.52
0.67
0.4
VF
0.26
0.54
0.56
0.58
0.72
0.3
0.40
0.68
0.70
0.75
0.85
0.2
PF
ll GA
2V2
T BER
EG
± 0.07 ± 0.12 ± 0.12 ± 0.07 ± 0.06 ± 0.03
V-Q
ual EG
± 0.07 ± 0.05 ± 0.07 ± 0.08 ± 0.06 ± 0.02
y sod Pro
± 0.04 ± 0.06 ± 0.07 ± 0.06 ± 0.07 ± 0.02
E
± 0.06
± 0.04
± 0.10
± 0.10
± 0.07 ± 0.09
± 0.07
± 0.06
± 0.03
± 0.03
W
Level – Target
Input Test
Level 3: MCI (Binary) Level 3: CERAD+ (Binary)
MMSE RL
Level 3: CERAD+ (Total) Level 2: LAN Level 1: PF
RL PF PF
DEV Set
HO Set
eGeMAPS All 0.62 ± 0.07 HuBERT 0.70 ± 0.01
Feature
0.63 0.65
0.58 ± 0.07 0.70 ± 0.03 0.85 ± 0.02
0.49 0.68 0.80
HuBERT HuBERT HuBERT
tive scores (LAN, MEM, EXE, VIS). Performance is analyzed across feature sets (upper panel) and input tasks (lower panel). In the upper panel, HuBERT consistently achieves the strongest performance, while eGeMAPS All remains competitive and often matches W2V2. Performance drops in the drawing-based EXE and VIS domains. In the lower panel, task-level analysis shows that PF and VF are the strongest predictors within LAN. Within MEM, RL emerges as the dominant task. Notably, MMSE performs equally well for EXE and LAN (r = 0.38), followed by MEM and VIS. 3.3. Level 3: Global Status Score Predictions Extending the domain-level analysis to continous global cognition score (CERAD+ total), task grouping reveals distinct aggregation dynamics (Fig. 4). Open-ended tasks (PF, VF) exhibit dilution, with predictive performance decreasing from task (L1) to global-level (L3) scores. RL shows relatively stable performance across aggregation levels. In contrast, constrained tasks (MMSE, RW) show inverse dilution, where aggregation improves prediction. 3.4. Generalizability to Hold-out & Feature Importance
Pearson R (mean ± std)
± 0.02
0.6
Pearson R (mean)
CERAD+ Tests
Level 1 results (Fig. 2) report mean (± SD) Pearson correlations on the development set for predicting individual task-level scores using features extracted from single neuropsychological tasks. Two consistent trends emerge. First, performance improves from hand-crafted eGeMAPS features (EG V-Qual) to self-supervised learning (SSL) representations across all assessments, with HuBERT yielding the highest prediction performance. Second, performance increases as tasks allow more open-ended responses: constrained tasks (MMSE, RW, BNT) show weaker performance, whereas open-ended tasks (VF, PF) show stronger performance.
Table 2: Top-performing models across the scoring hierarchy, reporting balanced accuracy (Binary) and Pearson correlation on the Development and disjoint Hold-out sets.
0.1
Hu
Features
Figure 2: Level 1: Individual test score prediction. Pearson correlation (r) between predicted and ground-truth scores. The x-axis shows feature sets; the y-axis lists MMSE and CERAD+ subtests ordered by increasing response freedom (from constrained to open-ended). Values denote mean ± SD across cross-validation folds.
3.2. Level 2: Cognitive Domain Score Predictions Moving from individual tests to domains, Level 2 results (Fig. 3) evaluate prediction of composite domain-level cogni-
Following hierarchical analyses, we evaluated model generalizability on an independent hold-out (HO) set and examined feature importance for binary MCI classification. Tab. 2 shows that HO performance closely matches the mean performance on the development (DEV) set across hierarchical levels, indicating robust generalization. HuBERT achieves the strongest performance for continuous targets and binarized CERAD+ scores. In contrast, MCI classification performs best using MMSE recordings with eGeMAPS (DEV: 0.62±0.07; HO: 0.63). Feature importance derived from SVM weights for this model on the HO set highlights interpretable acoustic correlates of impairment (Fig. 5). Positive coefficients (associated with MCI) include increased low-frequency spectral slope variability (+0.22) and elevated F0 instability (+0.18). Moreover, HCs show wider F1 /F2 bandwidths. A polarity shift is observed for spectral slope: steeper slopes in voiced segments are associated with HC, whereas steeper slopes in unvoiced segments correlate with MCI.
4. Discussion and Future Work In this study, we demonstrated that the predictive performance of speech features depends heavily on both the hierarchical level of the cognitive target and the nature of the task itself. For open-ended tasks such as phonemic fluency, we observed a clear dilution effect, with predictive performance declining at higher aggregation levels. These tasks can be viewed as “specialist” tasks: they are optimized to capture specific cognitive processes. However, because global cognition is multi-domain
Pearson r (mean ± std)
Figure 3: Level 2: Cognitive domain score prediction (LAN: Language, MEM: Memory, EXE: Executive Function, VIS: Visuospatial Ability). The upper panel shows mean (± SD) Pearson correlations by feature set across tasks, and the lower panel shows results by input task across feature sets. Tasks are ordered within domains by increasing degrees of freedom, from constrained to open-ended.
Model Performance across Hierarchical Scoring Levels 1.0 0.8 0.6 0.4 0.2 0.0 L1 L2 L3 L1 L2 L3 L1 L2 L3 L1 L2 L3 L1 L2 L3 L1 L2 L3
MMSE
RW
BNT
RL
VF
PF
Figure 4: Hierarchical prediction patterns across aggregation levels. Lines depict mean Pearson correlation r across 5-fold cross-validation for each task at Level 1 (Individual Tests), Level 2 (Cognitive Domains), and Level 3 (CERAD+ total).
and the signal from a specialist task captures only a subset of the construct, predictive performance is diluted. In contrast, more constrained screening tasks (such as the MMSE and RW) appear to function as “generalists”. These tasks exhibited an inverse dilution effect, with predictive performance improving at higher levels of the hierarchy. Individual items often show ceiling effects, with both MCI and HC groups achieving near-perfect scores, but aggregating items across the full assessment increases predictive performance. This generalist profile is further supported by the MMSE-based models’ ability to predict executive function scores derived entirely from non-speech drawing tasks, as well as language domain scores derived solely from speech. This suggests that speech features from structured screenings capture a cross-modal sig-
Slope [V]0-500 (std) Slope [UV]500-1500 (mean) Pitch (F0) (std) MFCC 4V (std) Pitch (F0) (med) loudness_stddevRisingSlope F2bandwidth (mean) Slope [V]500-1500 (mean) HNR (std) F1bandwidth (mean)
0.2
0.1
0.0
0.1
HC (SVM Weight) MCI
0.2
Figure 5: Feature Importance of best Level 3: MCI Binary Prediction Model (SVM Weight-Based)
nature of cognitive health. Consistent with this, our MMSEbased MCI model achieves the strongest binary classification performance across tasks and relies on interpretable eGeMAPS features, specifically increased F0 and spectral slope instability in the MCI group. These markers point to reduced speech motor control and phonatory instability, consistent with evidence that MCI is associated with greater acoustic instability and altered voice quality [26]. Despite promising results, our evaluation has limitations. It is limited to a single German-speaking cohort and omits sociodemographic and lifestyle covariates. Future work should test whether the proposed “specialist” and “generalist” profiles generalize across languages and cultural contexts. Joint hierarchical modeling may further capture dependencies between individual tests and global cognitive levels, improving the robustness and interpretability of speech-based cognitive monitoring.
5. Generative AI Use Disclosure Generative AI tools were used only for minor language editing and to improve readability. All research ideas, study design, experiments, analyses, and interpretations were conceived and carried out by the authors. The authors take full responsibility for the originality, validity, and integrity of the work.
6. Acknowledgements This research was funded by Gemeinnützigen Hertie-Stiftung and the Deutsche Forschungsgemein- schaft (DFG) through RU 5187 (project number 442075332)and RU 5363 (project number 459422098). Additional support was provided by the Machine Excellence Cluster and DFG through the Germany’s Excellence Strategy (EXC 2064 - project number 390727645) and the following projects: CRC 1404 (project number 414984028) and TRR 265 (project number 402170461). The authors gratefully acknowledge Dr. Ulrike Sünkel and Dr. Anna-Katharina von Thaler for their valuable assistance with data collection and annotation. The authors thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Serli Kopar.
7. References [1] F. Portet, P. J. Ousset, P. J. Visser et al., “Mild Cognitive Impairment (MCI) in Medical Practice: a Critical Review of the Concept and New Diagnostic Procedure. Report of the MCI Working Group of the European Consortium on Alzheimer’s Disease,” Journal of Neurology, Neurosurgery & Psychiatry, vol. 77, pp. 714–718, 2006. [2] J. Smid, A. Studart-Neto, K. G. César-Freitas et al., “Subjective Cognitive Decline, Mild Cognitive Impairment, and Dementia – Syndromic Approach: Recommendations of the Scientific Department of Cognitive Neurology and Aging of the Brazilian Academy of Neurology,” Dementia & Neuropsychologia, vol. 16, pp. 1–24, 2022. [3] N. D. Anderson, “State of the Science on Mild Cognitive Impairment (MCI),” CNS Spectrums, vol. 24, pp. 78–87, 2019. [4] J. Bohlken and K. Kostev, “Coded Prevalence of Mild Cognitive Impairment in General and Neuropsychiatrists Practices in Germany Between 2007 and 2017,” Journal of Alzheimer’s Disease, vol. 67, pp. 1313–1318, 2019. [5] J. C. Morris, A. Heyman, R. C. Mohs et al., “The Consortium to Establish a Registry for Alzheimer’s Disease (CERAD). Part I. Clinical and Neuropsychological Assessment of Alzheimer’s Disease,” Neurology, vol. 39, no. 9, pp. 1159–1165, 1989. [6] A. König, A. Satt, A. Sorin et al., “Automatic Speech Analysis for the Assessment of Patients with Predementia and Alzheimer’s Disease,” Alzheimer’s & Dementia: Diagnosis, Assessment & Disease Monitoring, vol. 1, pp. 112–124, 2015. [7] K. Mekulu, F. Aqlan, and H. Yang, “The Mild Cognitive Impairment Window for Optimal Alzheimer’s Disease Intervention,” J Alzheimers Dis Rep, vol. 9, p. 25424823251370768, 2025. [8] K. Ding, M. Chetty, A. Noori Hoshyar et al., “Speech Based Detection of Alzheimer’s Disease: a Survey of AI Techniques, Datasets and Challenges,” Artificial Intelligence Review, vol. 57, p. 325, 2024. [9] R. O. Roberts, Y. E. Geda, D. S. Knopman et al., “The Mayo Clinic Study of Aging: Design and Sampling, Participation, Baseline Measures and Sample Characteristics,” Neuroepidemiology, vol. 30, pp. 58–69, 2008. [10] M. J. Chandler, L. H. Lacritz, L. S. Hynan et al., “A Total Score for the CERAD Neuropsychological Battery,” Neurology, vol. 65, no. 1, pp. 102–106, 2005.
[11] M. Berres, A. U. Monsch, F. Bernasconi et al., “Normal Ranges of Neuropsychological Tests for The Diagnosis of Alzheimer’s Disease,” Studies in Health Technology and Informatics, vol. 77, pp. 195–199, 2000. [12] TREND Study Group. (2026) Tübinger Erhebung von Risikofaktoren zur Erkennung von Neurodegeneration (TREND). University Hospital Tübingen. Accessed: 2026-01-03. [Online]. Available: https://www.trend-studie.de/ [13] C. Kim and R. Stern, “Robust Signal-To-Noise Ratio Estimation Based on Waveform Amplitude Distribution Analysis,” 09 2008, pp. 2598–2601. [14] C. K. A. Reddy, V. Gopal, and R. Cutler, “DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,” CoRR, vol. abs/2010.15258, 2020. [Online]. Available: https://arxiv.org/abs/2010.15258 [15] C. Oh, R. Morris, X. Wang et al., “Analysis of Emotional Prosody as a Tool for Differential Diagnosis of Cognitive Impairments: a Pilot Research,” Frontiers in Psychology, vol. Volume 14 - 2023, 2023. [Online]. Available: https://www.frontiersin.org/journals/ psychology/articles/10.3389/fpsyg.2023.1129406 [16] H. Bredin, “Pyannote.metrics: A Toolkit for Reproducible Evaluation, Diagnostic, and Error Analysis of Speaker Diarization Systems,” in Interspeech 2017, 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, August 2017. [Online]. Available: http://pyannote. github.io/pyannote-metrics [17] S. Butterworth, “On the Theory of Filter Amplifiers,” Experimental Wireless & the Wireless Engineer, vol. 7, pp. 536–541, 1930. [18] S. Boll, “Suppression of Acoustic Noise in Speech Using Spectral Subtraction,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 27, no. 2, pp. 113–120, 1979. [19] C. J. Steinmetz, “pyloudnorm: A simple Python Implementation of ITU-R BS.1770 Loudness,” 2020, gitHub repository. [Online]. Available: https://github.com/csteinmetz1/pyloudnorm [20] F. Eyben, K. R. Scherer, B. W. Schuller et al., “The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing,” IEEE Transactions on Affective Computing, vol. 7, no. 2, pp. 190–202, 2016. [21] T. Kuroda, K. Ono, M. Onishi et al., “Utility of Artificial Intelligence-based Conversation Voice Analysis for Detecting Cognitive Decline,” PLOS ONE, vol. 20, no. 6, pp. 1–12, 06 2025. [Online]. Available: https://doi.org/10.1371/journal.pone. 0325177 [22] P. Sapkota, H. Srivastava, H. K. Kathania et al., “Do All Features Matter? Layer-wise Feature Probing of Self-supervised Speech Models for Dysarthria Severity Classification,” Speech Communication, vol. 175, p. 103326, 2025. [Online]. Available: https:// www.sciencedirect.com/science/article/pii/S0167639325001414 [23] K. Chlasta, P. Struzik, and G. M. Wójcik, “Enhancing Dementia and Cognitive Decline Detection with Large Language Models and Speech Representation Learning,” Frontiers in Neuroinformatics, vol. Volume 19 - 2025, 2025. [Online]. Available: https://www.frontiersin.org/journals/neuroinformatics/ articles/10.3389/fninf.2025.1679664 [24] F. Pedregosa, G. Varoquaux, A. Gramfort et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, no. 85, pp. 2825–2830, 2011. [Online]. Available: http://jmlr.org/papers/v12/pedregosa11a.html [25] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 785–794. [26] C. Themistocleous, M. Eckerström, and D. Kokkinakis, “Voice Quality and Speech Fluency Distinguish Individuals with Mild Cognitive Impairment from Healthy Controls,” PLOS ONE, vol. 15, no. 7, p. e0236009, 2020.