IMPROVING CLINICAL INTERPRETABILITY OF LINEAR NEUROIMAGING MODELS THROUGH FEATURE WHITENING Sara Petiton1 , Antoine Grigis1 , Raphaël Vock1 , Edouard Duchesnay1 1
University Paris-Saclay, CEA, CNRS, NeuroSpin, Baobab UMR 9027, Saclay, France
arXiv:2604.20675v1 [cs.LG] 22 Apr 2026
ABSTRACT
can be challenging when features are correlated, a common scenario in neuroimaging, where brain regions or imagingLinear models are widely used in computational neuroimagderived measures often covary [3]. For example, model ing to identify biomarkers associated with brain pathologies. weights assigned to left and right hemispheric measurements However, interpreting the learned weights remains chalof symmetrical regions, or to gray matter (GM) and cerelenging, as they do not always yield clinically meaningful brospinal fluid (CSF) volumes within the same region, often insights. This difficulty arises in part from the inherent corcapture shared variance rather than individual contributions, relation between brain regions, which causes linear weights thereby limiting clinical interpretability. While “forward to reflect shared rather than region-specific contributions. models” [3] address collinearity post hoc, they ignore known In particular, some groups of regions, including homolobrain structure, such as inter-hemispheric correlations, which gous structures in the left and right hemispheres, are known we explicitly incorporate during model fitting in the present to exhibit strong anatomical correlations. In this work, we work. leverage this prior neuroanatomical knowledge to introduce a Whitening, or sphering, is an orthogonalization techwhitening approach applied to groups of regions with known nique originally introduced in signal processing to remove shared variance, designed to disentangle overlapping inforlinear dependencies between features [4]. It has since been mation across correlated brain measures. We additionally adopted across various domains to improve model stability propose a regularized variant that allows controlled tuning and interpretability. In machine learning, whitening has reof the degree of decorrelation. We evaluate this method cently gained attention in self-supervised learning (SSL), as using region-of-interest features in two psychiatric classifia strategy to prevent feature collapse and enhance latent repcation tasks, distinguishing individuals with bipolar disorder resentations [5]. Emerging studies also suggest its potential or schizophrenia from healthy controls. Importantly, unlike to enhance interpretability of convolutional neural networks PCA or ICA which use whitening as a dimensionality reduc[6]. In neuroimaging, whitening has primarily been applied tion step, our approach decorrelates anatomically informed to EEG or fMRI data for noise reduction [7, 8]. Recent work pairs of neuroanatomical regions while retaining the full input has begun evaluating whitening as a preprocessing step to signal, making it specifically suited for feature interpretation improve the performance of post-hoc XAI methods in genrather than feature selection. Our findings demonstrate that eral image classification settings [9], finding benefits that whitening improves the interpretability of model weights vary by method and architecture, but has no applications in while preserving predictive performance, providing a robust neuroimaging contexts yet. Overall, whitening’s potential for framework for linking linear model outputs to neurobiologienhancing the interpretability of brain MRI models remains cal mechanisms. largely unexplored. Index Terms— structural MRI, Machine Learning, WhitenIn this work, we leverage correlation-based zero-phase ing, Interpretability, Neuroimaging, Psychiatry, Zero-Phase component analysis (ZCA-cor) whitening to enhance the inComponent Analysis terpretability of linear models trained on region-of-interest (ROI)-based brain features. ZCA-cor whitening has been identified as an optimal procedure for generating sphered 1. INTRODUCTION variables that closely preserve the structure of the original Linear models are widely used in computational neuroimagdata [10]. We implemented both ZCA-cor, and an original ing due to their simplicity, interpretability, and ability to regularized ZCA-cor as custom transformers compatible with uncover relationships between brain measurements and clinthe scikit-learn API. Whitening is applied in two steps: (i) ical outcomes. They are applied in a variety of contexts, between left and right hemisphere measures of the same reincluding classification of psychiatric disorders, prediction gion, to disentangle hemisphere-specific contributions; and of cognitive scores, and identification of biomarkers from (ii) between gray matter (GM) and cerebrospinal fluid (CSF) brain imaging [1, 2]. However, interpreting linear weights volumes within each brain region, given their typically in-
verse relationship. We evaluate the approach on two classification tasks, distinguishing individuals with bipolar disorder (BD) or schizophrenia (SCZ) from healthy controls (HC). Our results demonstrate that whitening and regularized whitening improve the interpretability of model weights without degrading predictive performance, offering a robust framework for linking linear model estimates to neurobiological mechanisms.
left amygdala GM, left amygdala CSF). We selected such pairs based on prior evidence that corresponding regions in the left and right hemispheres exhibit strong correlations, and on the inverse relationship between GM and CSF volumes, such that higher GM volume corresponds to lower CSF volume within the same brain region. Although we demonstrate the approach on ROI-level features for simplicity, the framework naturally extends to voxel-wise data by whitening pairs of spatially symmetric voxels across hemispheres, or pairs of GM and CSF voxel intensities within the same region.
2. MATERIALS AND METHODS 2.1. Datasets Classification tasks for BD and SCZ were conducted using aggregated multi-site neuroimaging datasets. For BD, we used the BIOBD and BSNIP datasets [11, 12], comprising 861 participants across 12 acquisition sites. Among them, 56.4% are female, 44.1% were diagnosed with BD, and the mean age is 37.69 ± 11.75 years. For SCZ, classification was performed using the SCHIZCONNECT-VIP dataset [13], totaling 604 participants across 4 acquisition sites. 38.2% of its participants are female, 45.5% were diagnosed with SCZ, and the mean age is 33.33 ± 12.34 years. Voxel-based morphometry (VBM) measures were extracted from T1-weighted brain MRI data using the CAT12 toolbox [14], (v 12.7). The brain was parcellated according to the Neuromorphometrics atlas [15], which defines 140 ROIs across both hemispheres. Both GM and CSF volumes were included, resulting in a total of 280 ROIs. 2.2. Machine Learning setup For both classification tasks (i.e., BD vs. HC, SCZ vs. HC), train and test sets were constructed using a ten-fold crossvalidation (CV). Classification performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC) and balanced accuracy (BAcc) metrics. ROIbased measures were stratified by age, sex, acquisition site, and diagnosis using the MULM [16] package. For whitening experiments, features were first residualized, then whitened, and finally scaled. We chose logistic regression as our classification model for its interpretability and linearity, allowing for straightforward computation of weights in the whitened space and their projection back to the original feature space. The regression was performed with scikit-learn (v. 1.3.2) [17], with hyperparameter tuning via a grid search, to identify the best L2 regularization parameter C, evaluated over the set {0.001, 0.01, 0, 1, 10}. 2.3. Identification of highly covarying features We whitened features pairwise to corresponding left and right hemispheric regions (e.g., left amygdala, right amygdala), and to GM and CSF volumes of the same brain region (e.g.,
2.4. Choice of whitening method Several whitening formulations exist, including principal component analysis (PCA) whitening, zero-phase component analysis (ZCA) whitening, and Cholesky whitening. These methods are reviewed in [10], where ZCA-cor whitening is recommended as the ideal whitening procedure to generate sphered variables that remain as close as possible to the original features. ZCA-cor whitening differs from standard ZCA in that it uses the correlation matrix instead of the covariance matrix, operating on standardized features to produce decorrelated variables that remain maximally aligned with the original data and improving numerical stability. 2.5. ZCA-cor whitening Given p a pair of correlated features (e.g., left/right amygdala GM volume, or GM/CSF of left hippocampus), let Xp ∈ Rn×2 denote the sequence of observations of those features over n subjects. Each feature is standardized to have zero mean and unit variance. We then compute the singular value decomposition (SVD) of the 2×2 correlation matrix Rp of the (standardized) Xp : Rp = Up Λp Up⊤ , where Up is the matrix of singular vectors of Rp and Λp is the corresponding diagonal matrix of singular values. We opt for SVD over eigenvalue decomposition for numerical stability. We then define the whitening matrix Wp by Wp = Up Λ−1/2 Up⊤ , p
(1)
and the whitened features by Zp = Xp Wp ,
with
Corr(Zp ) = I2 ,
(2)
A linear classifier is then trained on the whitened data: ŷ = Zβ + b, where Z represents the whitened input features (i.e., the stacked matrices Zp ), β denotes the classifier’s weights in the whitened space, b the bias term, and ŷ the predicted scores. Let WZCA-cor denote the block-diagonal whitening matrix formed by combining the pairwise whitening matrices Wp . An interpretable classifier ŷ = Xθ + b (whose inputs are the original features) is obtained by projecting the coefficients
from whitened space to feature space, for each pair p, via the transformation: ⊤ θ = WZCA-cor β.
(3)
2.6. Mapping classifier weights back to feature space For each pair of whitened features (z1 , z2 ) (each given as vector of length n), let β1 and β2 be the classifier weights associated with z1 and z2 in the whitened space, and θ1 and θ2 the corresponding weights in the original feature space. These weights are related by the 2 × 2 whitening matrix w11 w12 Wp = , w21 w22 where w11 = w22 > w12 = w21 and w11 > 0, since a 2 × 2 ZCA-cor whitening matrix is symmetric, with positive and equal diagonal terms larger than the off-diagonal entries. The original-space weights can be expressed as θ1 = w11 β1 +w12 β2 ,
θ2 = w21 β1 +w22 β2 = w12 β1 +w11 β2 .
It follows that β1 > β2 ⇒ (w11 − w12 )(β1 − β2 ) > 0 ⇒ θ1 > θ2 . Thus, if z1 has a higher weight than z2 in the whitened space, the same order is preserved in the original feature space, preserving the relative contributions of input features (i.e., left/right or GM/CSF features).
Fig. 1. Correlation matrices in feature (top) and whitened (bottom) space, computed on the BD dataset. Only amygdala, hippocampus, putamen, and anterior cingulate gyrus correlations were plotted for clarity and due to their known relationships to BD and SCZ.
2.7. Regularized ZCA-cor whitening Since our primary objective is interpretability, we propose a novel whitening procedure to decorrelate features while preserving class-relevant correlation structure. Partial whitening applied to pairs allows the classifier to retain inter-hemispheric relationships (e.g. some regions may exhibit stronger left-right correlations in patients than controls), which may reflect disease-related coupling patterns. Let Wα be the weighted whitening matrix: Wα = αW + (1 − α)I
(4)
where α ∈ [0, 1] controls the degree of decorrelation. When α = 1, the transformation reduces to the previously described ZCA-cor whitening. 3. RESULTS 3.1. Feature disentanglement via whitening First, we applied a regularized ZCA-cor whitening to leftright pairs of homologous ROIs, setting α = 0.3 to soften the whitening transformation and preserve class-specific covariance patterns. We then applied full (α = 1) pairwise whitening to GM-CSF volumes within each region. These settings
reflect that whitening matrices were estimated on the training set of each CV fold, including both patients and HC; while GM-CSF correlations are expected to be stable across individuals, left-right hemisphere correlations may vary across diagnostic classes. The ZCA-cor whitening matrices were computed using the training set and subsequently applied to the testing set for each of the 10 CV folds, thereby preventing data leakage. Figures 1 and 2 were obtained using the training set of the first fold of the BD dataset, using standardized features to accurately reflect the transformations applied before classification. In Figure 1, after whitening, the correlations between the homologous regions in the left and right hemispheres are markedly reduced, most clearly illustrated in the putamen (the four red squares at the center of the plot). Inverse correlations between GM and CSF are also attenuated, particularly visible in the bilateral amygdala (top-left of the plot). Figure 2 highlights the disentanglement of left/right and GM/CSF measures. After whitening, feature dispersion increases, with variance spread more evenly across hemispheres (bottom plot, hippocampus) and taking a more spherical distribution (top plot, anterior cingulate gyrus). Similar results were found with other folds and the SCZ dataset.
Original Feature R Pallidum (GM) R Cerebrum and Motor (GM) L Pallidum (GM) R Superior Occipital Gyrus (GM) L Parahippocampus Gyrus (CSF) L Anterior Insula (GM) L Medial Frontal Cerebrum (GM) Whitened Feature R Pallidum (GM) R Fourth Ventricle (CSF) L Third Ventricle (CSF) L Middle Cingulate Gyrus (CSF) R Cerebrum and Motor (GM) L Cerebral White Matter (CSF) L Parahippocampus Gyrus (CSF)
Weight (± std) 0.196 ± 0.010 −0.138 ± 0.007 0.136 ± 0.014 −0.120 ± 0.017 −0.115 ± 0.007 −0.111 ± 0.006 0.111 ± 0.011 Weight (± std) 0.249 ± 0.013 0.194 ± 0.022 0.176 ± 0.060 0.164 ± 0.011 −0.162 ± 0.010 0.140 ± 0.012 −0.140 ± 0.008
Table 2. Top seven contributing brain regions and corresponding model weights (mean ± std) for BD vs. HC classification. The mean is computed over the 10 CV folds regionwise. Fig. 2. Sphering of GM/CSF and left/right ROI pairs from the BD dataset. Top (blue): left anterior cingulate gyrus GM vs. CSF before (left) and after (right) whitening. Bottom (green): left vs. right hippocampus GM before and after whitening. Arrows show eigenvectors scaled up by 1.5 for clarity. BD vs. HC
Original Whitened
SCZ vs. HC
ROC-AUC
BAcc
ROC-AUC
BAcc
76.39 ± 3.88 76.24 ± 4.37
69.58 ± 4.15 69.65 ± 3.49
81.68 ± 2.89 81.0 ± 3.12
73.76 ± 4.27 72.41 ± 3.58
Table 1. Classification test set results for BD vs. HC and SCZ vs. HC with and without whitening. Values are reported as percentages of mean ± standard deviation across CV folds. 3.2. Whitening preserves classification performance Applying whitening did not affect classification performance (see Table 1). A Student’s t-test comparing ROC-AUC and BAcc across cross-validation folds revealed no significant differences between classifiers trained on whitened versus unwhitened features (p > 0.05). In all experiments, the L2 regularization parameter C was fixed at 0.01, as determined via grid search. 3.3. Whitening enhances interpretability Tables 2 and 3 report the mean coefficients estimated from logistic regression using whitened and unwhitened features for BD vs. HC and SCZ vs. HC classification, respectively. Weights computed on whitened features were mapped back to the original feature space as described in Equation 3. The
listed regions correspond to the seven features with the largest absolute coefficients, reflecting their importance in classification. Whitening improved the interpretability of regression coefficients by increasing their alignment with ENIGMA meta-analyses for both BD and SCZ (see [18, 19, 20, 21] for listings of ENIGMA-derived significant regions). In BD, whitening significantly strengthened correlations with cortical rankings and elevated subcortical regions (ventricles, hippocampus, thalamus, amygdala) closer to literature findings. In SCZ, whitening improved the ranking of key cortical and subcortical regions (hippocampus, inferior temporal gyrus, fusiform gyrus, lateral ventricles), even though overall correlations remained insignificant due to atlas differences. Across both disorders, whitening consistently enhanced the biological plausibility of classifier-derived ROI importance, underscoring its utility for more interpretable neuroimaging classification. 4. CONCLUSION In this study, we investigated the use of a novel regularized ZCA-cor whitening approach to enhance the interpretability of linear classifiers trained on ROI-based neuroimaging features. Pairwise whitening was performed between left and right hemispheres as well as between GM and CSF volumes, and classifier weights were subsequently projected back into the original feature space. Our results demonstrate that whitening preserves classification performance while improving the alignment of model-derived region importance as established by the ENIGMA meta-analytic results. While we apply our method to pairs of regions informed by prior
Original Feature Right Pallidum (GM) Left Pallidum (GM) Right Putamen (GM) Left Putamen (GM) Left Accumbens (CSF) Left Frontal Pole (GM) Right Angular Gyrus (CSF) Whitened Feature Right Pallidum (GM) Right Putamen (GM) Left Exterior Cerebellum (CSF) Left Hippocampus (CSF) Right Central Operculum (CSF) Right Cerebrum and Motor (CSF) Left Anterior Cingulate Gyrus (CSF)
6. COMPLIANCE WITH ETHICAL STANDARDS
Weight (± std) 0.261 ± 0.012 0.174 ± 0.008 0.152 ± 0.014 0.128 ± 0.014 −0.116 ± 0.007 0.111 ± 0.008 −0.108 ± 0.008 Weight (± std) 0.311 ± 0.015 0.169 ± 0.019 −0.164 ± 0.012 −0.159 ± 0.013 0.145 ± 0.017 −0.137 ± 0.022 −0.136 ± 0.018
Table 3. Top seven contributing brain regions and corresponding model weights (mean ± std) for SCZ vs. HC classification. The mean is computed over the 10 CV folds regionwise.
knowledge of brain structure, applying it to the full set of features would yield results closely related to Haufe’s forward models [3]. While whitening has been applied in neuroimaging primarily as a dimensionality reduction technique (e.g., via PCA or ICA) or for noise reduction in EEG and fMRI signals [7, 8], it had not previously been used to improve the interpretability of neuroimaging models. Looking ahead, while whitening extensions have been explored in deep learning for feature interpretation [6], to our knowledge, this approach has not yet been adapted to neuroimaging applications. Our work addresses this gap and provides a foundation for extending the proposed methodology beyond linear models, integrating it into more complex architectures such as deep neural networks.
5. ACKNOWLEDGMENTS This research was generously supported by The Robert Debré Child Brain Institute (Paris) under grant ANR-23IAHU-0010 and the research program in precision psychiatry (PEPR PROPSY, ANR-22-EXPR-0001), both of which are funded by the France 2030 program and the French National Research Agency (ANR). Additional support was provided by two ”Investissements d’Avenir” Hospital-University Research in Health initiatives: RHU-PsyCARE (ANR-18RHUS 0014) and FAME (ANR-21-RHUS-0009).
This study was performed retrospectively using human participant data in accordance with local ethics guidelines. 7. REFERENCES [1] Abraham Nunes et al., “Using structural mri to identify bipolar disorders – 13 site machine learning study in 3020 individuals from the enigma bipolar disorders working group,” Molecular Psychiatry, vol. 25, no. 9, pp. 2130–2143, Sept. 2020. [2] Ashley N. Nielsen et al., “Machine Learning With Neuroimaging: Evaluating Its Applications in Psychiatry,” Biological Psychiatry: Cognitive Neuroscience and Neuroimaging, vol. 5, no. 8, pp. 791–798, Aug. 2020. [3] Stefan Haufe et al., “On the interpretation of weight vectors of linear models in multivariate neuroimaging,” NeuroImage, vol. 87, pp. 96–110, Feb. 2014. [4] Alain De Cheveigné and Lucas C. Parra, “Joint decorrelation, a versatile tool for multichannel data analysis,” NeuroImage, vol. 98, pp. 487–505, Sept. 2014. [5] Aleksandr Ermolov et al., “Whitening for selfsupervised representation learning,” 2020. [6] Zhi Chen, Yijie Bei, and Cynthia Rudin, “Concept whitening for interpretable image recognition,” Nature Machine Intelligence, vol. 2, no. 12, pp. 772–782, Dec. 2020. [7] Wiktor Olszowy, John Aston, Catarina Rua, and Guy B. Williams, “Accurate autocorrelation modeling substantially improves fmri reliability,” Nature Communications, vol. 10, no. 1, pp. 1220, Mar. 2019. [8] Denis A. Engemann and Alexandre Gramfort, “Automated model selection in covariance estimation and spatial whitening of meg and eeg signals,” NeuroImage, vol. 108, pp. 328–342, Mar. 2015. [9] Benedict Clark, Stoyan Karastoyanov, Rick Wilming, and Stefan Haufe, “The effect of whitening on explanation performance,” arXiv preprint arXiv:2602.09278, 2026. [10] Agnan Kessy, Alex Lewin, and Korbinian Strimmer, “Optimal whitening and decorrelation,” The American Statistician, vol. 72, no. 4, pp. 309–314, Oct. 2018. [11] S Sarrazin, A Cachia, F Hozer, et al., “Neurodevelopmental subtypes of bipolar disorder are related to cortical folding patterns: An international multicenter study,” Bipolar Disorders, vol. 20, pp. 721–732, 2018.
[12] Carol A. Tamminga et al., “Bipolar and schizophrenia network for intermediate phenotypes: Outcomes across the psychosis continuum,” Schizophrenia Bulletin, vol. 40, no. Suppl 2, pp. S131–S137, 2014. [13] SchizConnect, “Schizconnect: A multisite schizophrenia neuroimaging database,” https://www.schizconnect.org/. [14] Christian Gaser et al., “Cat: a computational anatomy toolbox for the analysis of structural mri data,” GigaScience, vol. 13, pp. giae049, Jan. 2024. [15] Neuromorphometrics, “Neuromorphometrics: Brain atlas and parcellation resources,” https://www. neuromorphometrics.com/, 2025, Accessed: Nov 14, 2025. [16] pylearn-mulm, “mulm: massive univariate linear model,” https://www.neurospin.fr/ pylearn-mulm/, 2021. [17] Fabian Pedregosa et al., “Scikit-learn: Machine learning in python,” Journal of Machine Learning Research, 2012. [18] D P Hibar et al., “Cortical abnormalities in bipolar disorder: an mri analysis of 6503 individuals from the enigma bipolar disorder working group,” Molecular Psychiatry, vol. 23, no. 4, pp. 932–942, Apr. 2018. [19] Hibar D. P. et al., “Subcortical volumetric abnormalities in bipolar disorder,” Molecular Psychiatry, vol. 21, no. 12, pp. 1710–1716, Dec. 2016. [20] TGM van Erp et al., “Cortical brain abnormalities in 4474 individuals with schizophrenia and 5098 control subjects via the enhancing neuro imaging genetics through meta analysis (enigma) consortium,” Biological Psychiatry, vol. 84, no. 9, pp. 644–654, 2018. [21] T G M Van Erp et al., “Subcortical brain volume abnormalities in 2028 individuals with schizophrenia and 2540 healthy controls via the enigma consortium,” Molecular Psychiatry, vol. 21, no. 4, pp. 547–553, Apr. 2016.