Conceptio › Archive › arXiv CS
arXiv CSopen access

SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis Yuqing YANG1 , Alexander SCHMATZ2,6 , Zhaozhao MA3 , Changkyu CHOI4 , Robert JENSSEN5 , Shujian YU5,6,* 2

1 Independent Researcher Leiden Institute of Advanced Computer Science (LIACS), Leiden University, Leiden, The Netherlands 3 School of Applied and Creative Computing, Purdue University, West Lafayette, IN, USA 4 Department of Informatics, University of Oslo, Oslo, Norway 5 Machine Learning Group, UiT — The Arctic University of Norway, Tromsø, Norway 6 Quantitative Data Analytics Group, Vrije Universiteit Amsterdam, Amsterdam, The Netherlands

arXiv:2609.05174v1 [cs.CV] 4 Sep 2026

*

Corresponding author: [email protected].

Abstract Explainability is increasingly seen as a crucial requirement in AI-based medical diagnosis, particularly in safety-critical clinical decision-making. Most existing explainability methods in healthcare operate in a post-hoc manner and are predominantly designed for unimodal data, which limits their applicability in increasingly prevalent multimodal diagnostic settings. This paper addresses the problem of self-explainable multimodal diagnosis by formulating it within the information bottleneck (IB) framework. We propose a unified learning paradigm that jointly optimizes predictive performance and modality-specific explainability by identifying the most informative elements inside each modality that contribute to diagnostic decisions. To enable tractable and stable optimization, we employ a matrix-based Rényi’s α-order entropy functional under the assumption of sufficiently expressive encoders. Extensive experiments on representative medical datasets spanning heterogeneous modalities demonstrate that the proposed method consistently achieves strong diagnostic performance, including an absolute accuracy improvement of 9.1 percentage points on the iCTCF dataset. Moreover, the learned explanations provide transparent and modality-aware insights into feature relevance, thereby improving both the explainability and generalization.

1

Introduction

Deep learning has achieved substantial success in a wide range of application domains, including medical diagnosis, where data-driven models have demonstrated strong predictive performance across diverse clinical tasks [1, 2]. Despite these advances, the increasing complexity of modern deep models has rendered their decision-making processes largely opaque; their “black-box" nature becomes a significant barrier to clinical adoption. In safety-critical medical settings, such opacity undermines clinicians’ confidence in assessing the reliability of automated diagnostic decisions, thereby motivating a growing demand for explainable artificial intelligence (XAI) [3, 4]. Representative post-hoc explanation methods explain the predictions of an already-trained model without modifying its training process. Methods such as Grad-CAM [5], LIME [6], and SHAP [7], aim to identify salient input features or regions that influence model outputs, and have been applied to explain medical images [8] as well as non-image clinical data [9]. While effective in unimodal scenarios, these methods are predominantly designed for natural images or text, and their post-hoc nature may limit the faithfulness of the resulting explanations. As a result, their applicability to complex clinical decision-making pipelines remains limited. In real-world medical diagnosis, clinical decisions are rarely based on a single data source. Instead, heterogeneous modalities—such as medical imaging, genomic profiles, and structured 1

clinical records—are jointly considered, with each modality providing unique and complementary information about patient health. Multimodal learning frameworks that integrate such heterogeneous data have demonstrated notable improvements in diagnostic performance by capturing cross-modal dependencies and shared representations [10–12]. Despite their success, multimodal diagnostic models introduce additional challenges for explainability. The heterogeneity of data types and the complex interactions between modalities make it difficult to interpret how individual features and modalities contribute to the final prediction. Existing explainability approaches for multimodal models largely rely on post-hoc extensions of unimodal methods [9, 13], which may not faithfully reflect the model’s decision-making process. Consequently, there remains a lack of principled frameworks that jointly optimize multimodal prediction and explanation within a unified learning objective. To fill this gap, we study self-explainable multimodal medical diagnosis from an informationtheoretic perspective by formulating it as an information bottleneck (IB) problem [14], where modality-specific explainers are learned to identify compact subsets of input elements that retain diagnostic information while discarding redundant information. This formulation naturally couples explanation generation with representation learning, ensuring that the selected features are intrinsically aligned with the model’s decision process rather than obtained post hoc. Moreover, distinct explainers are designed for each modality to reflect their heterogeneous clinical data characteristics. To enable tractable and stable optimization of the IB objective, we assume sufficiently expressive modality-specific encoders [15] and employ a matrix-based Rényi’s α-order entropy functional [16, 17] for mutual information estimation, which avoids introducing auxiliary parametric models such as the mutual information neural estimator (MINE) [18]. The main contributions of this work include: • We propose Self-explainable Multimodal Information bottLEneck (SMILE), which, to the best of our knowledge, is the first IB-based self-explainable framework for multimodal medical diagnosis. Different from post-hoc approaches that analyze an already-trained predictor, SMILE learns the predictive representation and explanation jointly, making the selected features an intrinsic component of the decision process. • We develop a relaxed and tractable training objective for SMILE that enables efficient and stable learning. Furthermore, we provide theoretical insights into the generalization behavior of SMILE, showing that it does not suffer from an intrinsic trade-off between explainability and predictive accuracy. • We evaluate SMILE on multiple representative multimodal medical datasets, where it consistently achieves state-of-the-art diagnostic performance. In addition, SMILE produces instance-wise, modality-specific explanations that are quantitatively and qualitatively validated, demonstrating improved faithfulness and explainability. The remainder of this paper is structured as follows: Section 2 gives a brief review of related works, and Section 3 introduces each module of SMILE and gives a generalization analysis of the model. Section 4 presents both the quantitative and qualitative experimental results on diverse representative datasets. Finally, Section 5 draws the conclusion.

2

2

Related Work

2.1

Explainability for Medical Diagnosis

Explainable AI (XAI) has been extensively studied to improve the transparency of AI-based diagnostic systems. A large body of work focuses on post-hoc explanation, which can be broadly grouped into saliency/attribution-based and perturbation-based approaches. Saliency-based methods are widely applied in medical imaging data to visualize important areas of interest, such as tumors or lesions, using methods like Grad-CAM [5] and SmoothGrad [19]. Perturbation-based methods estimate feature importance by measuring changes in model outputs under controlled input modifications; classical formulations include influence functions [20] and additive feature attribution such as SHAP [7]. Beyond generic post-hoc explanations, recent studies have begun to emphasize whether explanations are meaningful and relevant to clinical decision-making processes. For instance, guidelineoriented frameworks incorporate medical guidelines into explanations to improve clinical relevance [21], and clinician-centered principles provide practical recommendations for clinical XAI design [22]. In parallel, existing work has explored explainability methods tailored to specific data modalities, including structured electronic health records and clinical time series (e.g., Dynamask [23]) and graph-structured biomedical data (e.g., GNNExplainer [24]). With the growing availability of heterogeneous clinical data, explainability has become an increasingly important consideration for multimodal diagnostic models. Existing explanation methods for multimodal medical AI are predominantly post-hoc, often rely on classic approaches, and tend to focus on specific disease types, limiting their broader applicability. For example, [25] extends the concept activation method [13] to generate human-understandable concepts for explaining prostate cancer detection using PET and CT scans. In another study [9], image data and metadata are combined for skin lesion diagnosis, where Grad-CAM is applied to interpret image features, and kernel SHAP [7] is used to explain metadata. More recently, clinical perspectives highlight that diagnostic decisions often rely on multiple data modalities, which makes it more challenging to present and integrate explanations across modalities [26]. A closely related line of work studies multimodal explainability at the modality or interaction level, quantifying which modalities (or modality interactions) drive a prediction. Representative examples include interpretable fusion mechanisms that assign relevance scores to modalities and their interactions (e.g., InTense [27]) and methods that explicitly quantify cross-modal interactions and modality contributions (e.g., InterSHAP [28]). These modality-level analyses provide coarse-grained attribution over modalities; clinical interpretation often requires identifying which input elements within each modality are diagnostically informative for a given patient. In contrast to most existing multimodal explanation methods in medical AI, our work targets feature-level self-explainability in a modality-aware manner, providing a generic framework that can, in principle, be applied to diverse data modalities and potentially adapted to multiple disease types.

2.2

Information Bottleneck (IB) in Deep Learning

The IB principle [14, 29] considers extracting relevant information from a random variable X to predict another variable Y . It operates by identifying a “bottleneck" variable X̃ that maximizes its predictive power to Y , as expressed by mutual information I(Y ; X̃), while restricting the amount of information it carries about X, formulated as I(X; X̃): min −I(Y ; X̃) + βI(X; X̃), p(X̃|X)

3

(1)

Explanation (importance score)

Selection Network

Share

Modality-m

Explainer Conv Layers

Upsampling

....

number of selected patches

....

Modality-1 Linear Layers

Share

Explanation (importance score)

...

Image-based

....

Gumbel-Softmax

Prediction

Gumbel-Softmax

Metadata-based

Explanation (importance score)

Share

Modality-m

Figure 1: Overview of the proposed SMILE framework. We design a modality-specific explainer to identify the top-k most informative features within each modality that influence diagnostic decisions. All selected features {E m }M m=1 serve as the final explanations and are directly used for prediction, ensuring consistency between the explanations and the model decisions. where β ≥ 0 is a Lagrange multiplier. IB-inspired objectives have been studied to encourage compact, generalizable representations in deep learning [30], in which X̃ usually refers to the latent layer representations. Recently, IB has been extended to multimodal representation learning to address redundancy and modality-specific noise [12, 31, 32]. For example, multimodal information bottleneck formulations aim to learn minimal and sufficient unimodal and joint representations under information-theoretic regularization [32]. More recent work further investigates how to set regularization weights or handle imbalanced task-relevant information across modalities (e.g., OMIB) [33]. While these works have demonstrated strong empirical generalization performance, they lack theoretical justification and the ability to explain feature contributions to the decision. On the other hand, the IB principle has also been explored for single-modality explainability. Its central idea is to learn a differentiable mask M over the input X, such that the masked input X ⊙ M is maximally informative for the decision. Owing to its flexibility, this formulation has been adopted for both post-hoc explainability [34] and self-explainability [35, 36]. However, extending these approaches to heterogeneous multimodal clinical data remains an open challenge. Closer to our goal, recent studies explicitly leverage IB to explain multimodal representations (e.g., IB-based attribution for image–text representations [37], and cross-attention-guided multimodal IB attribution [38]). However, these methods are designed as post-hoc attribution tools for specific (mostly vision-language) architectures, and do not address self-explainable diagnosis for clinical data.

3

Methodology

3.1

Objective of SMILE

Given N i.i.d. multi-modal observations from M modalities with their corresponding labels M N m K {{xm i }m=1 , yi }i=1 , let xi denote the i-th sample in the m-th modality. Here, yi ∈ R , where K m represents the number of classes. Each xi can be a vector, such as genomic data; a three-dimensional (3D) volume, such as chest computed tomography (CT) scans; or a graph structure, such as brain 4

networks constructed from functional magnetic resonance imaging (fMRI) signals. Our goal is to learn a predictive model f : {X m }M m=1 → Y , which also integrates self-explainability in the sense that f is able to identify the most informative input elements from each modality. We therefore introduce a set of M built-in explainers to perform a deterministic mapping X m 7→ E m , where N m has the same size as X m and acts as an explanation to X m . Each element in {em i }i=1 ∈ E m E is a binary variable, with a value of 1 indicating an informative feature and 0 indicating a non-informative feature. Typically, E m is both sample-dependent and modality-dependent. In order to learn a modality-specific explanation E m , we employ the IB principle. Specifically, m m let x̃m i = ei ⊙ xi denote the set of all selected features in the m-th modality, where ⊙ represents element-wise multiplication, the objective for single modality explanation can be expressed as: min −I(Y ; X̃ m ) + βI(X̃ m ; X m ) + λ∥E m ∥0 , m E

(2)

where ∥.∥0 is the ℓ0 norm that encourages the sparsity of E m , only a few elements can be identified as informative, β and λ are non-negative regularization coefficients. A closely related objective to our Eq. (2) is the INstance-wise VAriable SElection (INVASE) [39], which formulates the learning of E in a single modality as: 



min DKL p(Y |X)∥p(Y |X̃) + λ∥E∥0 , E

(3)

where DKL refers to the Kullback-Leibler (KL) divergence. Proposition 1. The IB objective in (2) encompasses that of INVASE as a special case when β = 0. Proof. All proofs are provided in the supplementary material. One could argue that minimizing I(X; X̃) plays a similar role to the sparsity constraint ∥E∥0 . In fact, minimizing I(X; X̃) can be viewed as an implicit form of sparsity regularization; however, the two terms play complementary roles. Specifically, the ℓ0 constraint explicitly controls the number of selected features, whereas the IB compression term regulates the amount of information preserved in the selected representation. Since the explainer performs deterministic masking, we have H(X̃ | X) = 0, and therefore I(X; X̃) = H(X̃). Hence, minimizing the IB compression term encourages compact, low-entropy representations rather than merely reducing the number of selected features. Generally, higher-dimensional variables have greater entropy because additional dimensions introduce more uncertainty into their joint distribution. To illustrate the role of both regularization terms, let us consider a toy example in the 2D space, as shown in Fig. 2. Among the four features {x1 , · · · , x4 }, both pairs (x1 , x2 ) and (x3 , x4 ) satisfy ∥E∥0 = 2 and effectively distinguish between the two classes. Consequently, both pairs have high and similar mutual information values I(Y ; X̃), where X̃ represents (x1 , x2 ) or (x3 , x4 ). Hence, a regularization term ∥E∥0 alone is insufficient to select between (x1 , x2 ) and (x3 , x4 ). However, an entropy minimization penalty further drives our objective to favor (x3 , x4 ) over (x1 , x2 )1 . Note that a compressed representation is always a desirable property in downstream applications. Our analysis in Section 3.4 will further justify the rationale behind I(X; X̃). Therefore, we argue that a mutual information regularization term I(X; X̃) is essential for effective instance-wise feature selection. Our objective in Eq. (2) incorporates both terms, achieving effective and efficient feature selection. 1

Entropy is minimized if all points converge to a single point.

5

subject to

subject to

min

Figure 2: A 2D toy example illustrating how minimizing mutual information I(X; X̃) and entropy H(X̃) aids in feature selection.

...

...

...

Having illustrated the rationale behind each term in Eq. (2), it is natural to extend it to the multimodal scenario, in which we aim to learn {E m }M m=1 simultaneously: min

{E m }M m=1

− I(Y ; X) + β

M X

I(X̃ m ; X m ) + λ

m=1 1

s.t., X = gω (X̃ , · · · , X̃ M ),

M X

∥E m ∥0 ,

m=1 m m x̃i = em i ⊙ xi ,

(4)

where gω refers to a fusion network with parameter ω. However, in practical applications, especially in medical diagnosis, where data is often noisy, high-dimensional, and structured, the direct estimation of I(X̃ m ; X m ) is infeasible. Inspired by the invariance property of mutual information under reparameterization [40, 41]: if f1 : X 7→ X ′ and f2 : Y 7→ Y ′ are homeomorphisms (i.e., smooth invertible maps), then it holds that I(X ; Y) = I(X ′ ; Y ′ ), we introduce an encoder hϕm : X m 7→ Z m for each modality, and compute mutual information terms in the transformed, low-dimensional latent space. The lossless transformation assumption is also known as the sufficient encoder assumption in contrastive learning [15]. In addition, we implement the ℓ0 sparsity regularizer in Eq. (4) as an explicit top-k cardinality constraint. Thus, Eq. (5) should be understood as a constrained implementation of Eq. (4), where the sparsity penalty controlled by λ is replaced by the budget k: min

{E m }M m=1

−I(Y ; Z) + β

M X

I(hϕm (X̃ m ); hϕm (X m )),

m=1 1

s.t., Z = gω (hϕ1 (X̃ ), · · · , hϕM (X̃

M

(5)

)), ∥em i ∥0 ≤ k,

where hϕm denotes a modality-specific encoder with parameters ϕm . The latent variables Z m = hϕm (X m ) and Z̃ m = hϕm (X̃ m ) facilitate a tractable estimation of mutual information terms. In principle, this strategy can be applied to any modality or data type, provided there is an appropriate encoder to transform complex data into a vector representation. ∥em i ∥0 ≤ k refers to as the k-sparse norm [42]. In SMILE, we simply take gω with a concatenation operator, that is Z = [hϕ1 (X̃ 1 ), · · · , hϕM (X̃ M )]. Hence, our explainable predictive model f consists of three modules, namely the explainer for generating explanations, the encoder for transforming data to latent spaces, and the classifier for predictive tasks, as shown in Fig. 1.

3.2

Modality-specific, Built-in Explainer

Here, we introduce the design of our modality-specific built-in explainers. Our explainers operate directly on the input space of each modality and identify the most informative input elements 6

contributing to the diagnostic decision. Specifically, the definition of an input element is modalitydependent: for image data, the elements can correspond to pixels or image patches; for graphstructured data, they correspond to nodes or edges that form an informative subgraph; and for 3D volumetric data, they correspond to voxels or subvolumes. Therefore, we assume that each modality input consists of d elements, where d is modality-dependent. d m d m For an instance xm i ∈ R , given a fixed size k (0 ≤ k ≤ d), we can define ℘k = {si ⊂ 2 , |si | = k}, which represents the collection of all possible subsets of elements of size k drawn from a set of d elements. Our aim is to find a small subset sm i ∈ ℘k with k ≪ d that contains the most informative elements for predicting the model output. We thereby introduce a function S, mapping each m element in xm i to an importance score that indicates its likelihood of being part of the subset si . This importance score is learned through a neural network (i.e., selection network Sψm in Fig. 1) parameterized by ψm , and the subset selection is based on these scores. The architecture of the selection network is modality-dependent and can be adapted to the characteristics of different data types, such as images, graphs, and structured clinical features (see details in Section 4). d However, a direct estimation requires summing over k combinations of feature subsets, which is intractable. Therefore, we employ the Gumbel-Softmax [34, 43] to overcome this problem in a differentiable manner. Suppose we aim to approximate a categorical random variable represented as a one-hot vector in Rd with category probabilities p1 , p2 , . . . , pd , we start by adding a random perturbation to the log probability of each category log pi : Gi = − log(− log ui ) where ui ∼ Uniform(0,1), exp{(Gi + log pi )/τ } Ci = P d , i=1 exp{(Gi + log pi )/τ }

(6)

where τ is a tuning parameter for the temperature of the Gumbel-Softmax distribution. Then, we can define a Concrete random vector C = [C1 , · · · , Cd ], which serves as a continuous, differentiable approximation of a categorical random variable represented as a one-hot vector in Rd . We further define a continuous-relaxed random variable C ∗ = [C1∗ , · · · , Cd∗ ] as the element-wise maximum of the independently sampled Concrete vectors C (j) where j = 1, · · · , k: (j)

Ci∗ = max Ci j

i.i.d for j = 1, ..., k.

(7)

d We denote this sampling result C ∗ as the generated explanation em i ∈ R , where top-k elements are most informative for the prediction of the model.

3.3

Mutual Information Estimation

We now describe the estimation of mutual information. For the estimation of I(Y ; Z), we can use a standard cross entropy loss [30]. This is because, letting Q(Y |Z) be a variational approximation to P (Y |Z), we have: I(Y ; Z) ≥ H(Y ) + EP (Y,Z) [log Q(Y |Z)], (8) where H(Y ) is independent of the optimization and thereby can be ignored. EP (Y,Z) [log Q(Y |Z)] is equivalent to minimizing the usual cross-entropy loss [44]. For the term I(Z̃ m ; Z m ), which can be decomposed as: I(Z̃ m ; Z m ) = H(Z̃ m ) + H(Z m ) − H(Z̃ m , Z m ).

Maximizing

(9)

We employ the matrix-based Rényi’s α-order entropy functional [16, 17] to quantify each entropy term in (9) in a non-parametric manner. Unlike existing parametric estimators such as MINE [18] 7

...

...

...

Figure 3: Flow of information in SMILE. or InfoNCE [45], this approach directly estimates entropy from data using positive-definite matrices and avoids introducing an auxiliary parametric model, which complicates training. For the m-th modality, we denote the set of latent feature vectors from N instances as {zim }N i=1 , m . The Gram matrix K ∈ RN ×N can be computed as K = κ(z m , z m ) by where zim = hϕm (xm ) ∈ Z ij i i j using a positive definite kernel κ. Here, we use the radial basis function (RBF) kernel κ(zim , zjm ) = ∥z m −z m ∥2

exp(− i 2σ2j ). The normalized Gram matrix can be defined as A = K/tr(K), a matrix-based analogue to Rényi’s α-order entropy is subsequently defined as: N X 1 Hα (A) = log2 λi (A)α , 1−α i=1

!

α ∈ (0, 1) ∪ (1, ∞)

(10)

which estimates H(Z m ), and λi (A) denotes the i-th eigenvalue of A. With the explainable representation z̃ m , we can apply the same process and obtain Hα (Ã), where à is the normalized Gram m matrix derived from {z̃im }N i=1 ∈ Z̃ . Additionally, the joint entropy can be defined similarly using the Hadamard product ⊙: ! A ⊙ à , (11) Hα (A, Ã) = Hα tr(A ⊙ Ã) which estimates H(Z̃ m , Z m ). Given Eqs. (10) and (11), we can estimate I(Z̃ m ; Z m ) with Eq. (9).

3.4

Generalization Analysis

We finally present a generalization analysis of our SMILE. The information flow within SMILE is illustrated in Fig. 3. Specifically, SMILE consists of M channels, where each explanation X̃ m and its latent representation Z̃ m depend solely on X m . SMILE then fuses information from the set m M {Z̃ m }M i=1 to form Z̄, which is then used for prediction. For simplicity, we can also treat {X }i=1 as a joint information source, denoted by X, which transmits information through a parameterized channel to Z̄ and subsequently to Y . Extending the theoretical result in [46], we show that the P m m compression term M m=1 I(X̃ ; X ) introduced in (4) further reduces the generalization error, as demonstrated in Proposition 2. Proposition 2. With probability at least 1 − δ over the training data t = {xi , yi }N a i=1 drawn from  M , the generalization error ∆(t) = E t (x, y)) − data distribution p(x, y), where xi = {xm } ℓ(f i m=1 p(x,y) 1 PN t i=1 ℓ(f (xi , yi )) roughly obeys the following form: N

s

∆(t) = Õ 

PM m

I(X̃ m ; X m ) + 1 N

  as N → ∞,

where ℓ is a bounded per-sample loss, f t is the full model obtained by training over t. 8

(12)

It is widely believed that as model complexity increases to improve performance, interpretability often declines, leading to a trade-off where more interpretable models may sacrifice some predictive accuracy [47]. However, our theoretical analysis suggests that the IB-based compression mechanism does not necessarily introduce an intrinsic accuracy penalty, which is consistent with recent perspectives [48, 49]. We provide empirical evidence to support this argument in Section 4.

4

Experiments and Results

4.1

Experimental Setup

Datasets. To demonstrate the effectiveness of SMILE, we evaluate it on five representative datasets that cover commonly used medical multimodal settings: i) BRCA [10] (breast carcinoma PAM50 subtypes) and ROSMAP [10] (Alzheimer’s Disease), where each subject is characterized by three structured-data modalities, i.e., mRNA expression, DNA methylation, and miRNA expression; ii) iCTCF [50] (COVID-19 morbidity), which contains paired HRCT scans and clinical features; following [12], each 3D HRCT volume is preprocessed into a 700 × 700 2D montage for model input; iii) Glaucoma Grading [51], a dual-image-modality dataset consisting of fundus images (resized to 256 × 256) and OCT images (resized to 512 × 512) for glaucoma severity grading; iv) REST-meta-MDD [52], a large-scale multi-site resting-state consortium for major depressive disorder (MDD), aggregated from 25 cohorts. For sMRI, we use z-normalized gray matter volume (GMV) maps of size (121 × 145 × 121). For rs-fMRI, we construct functional connectivity matrices by computing pairwise Pearson correlations between the time series of 116 brain regions defined by the Automated Anatomical Labeling (AAL) atlas [53], followed by Fisher-z transformation. Each subject is therefore represented by a 116 × 116 functional connectivity matrix. After quality control (QC), 1,604 subjects are retained. See Table 1 for characteristics of datasets. Implementation Details. In the explainer module, we employ modality-specific selection networks to accommodate the heterogeneous characteristics of different data modalities. The detailed architectures and implementation settings of these selection networks are provided in the supplementary material. For image data, the explainer operates at the patch level rather than individual pixels and employs convolutional layers to capture local spatial patterns. For structured data, such as clinical variables and genomic features, fully connected networks are used to estimate feature importance. For graph-structured data, the explainer evaluates the importance of graph elements (i.e., edges) using modality-specific networks; in our implementation, the edge importance is estimated by an MLP based on the features of the corresponding node pairs. In the fusion stage, we use simple concatenation to fuse latent representations from each modality. For training the proposed method on BRCA and ROSMAP, we follow the experimental settings in [11] and set k = 30. For iCTCF, we adopt the settings from [12] and use k = 30 patches of size 70 × 70 for the image modality and k = 20 for clinical features. For Glaucoma Grading, we use the experimental setup in [54] and set k = 100, using 8 × 8 patches for fundus images and 4 × 4 slices for OCT images. For REST-meta-MDD, we use soft top-k selection with k = 25 edges (rs-fMRI; ∼1% of admissible pairs) and k = 25 subvolumes (sMRI; ∼5% of a 7 × 9 × 7 grid); in all cases, k is chosen empirically as the smallest value before performance drops. All experiments are performed on a single H100 GPU (80GB, cuDNN 9.3), and each experiment is repeated 10 times to report the mean and standard deviation.

9

Table 1: Dataset characteristics. Note that we use a prognosis task for iCTCF rather than a diagnostic task to identify infection, to better showcase explanations. Dataset

Modality

BRCA ROSMAP iCTCF Glaucoma Grading REST-meta-MDD

mRNA, DNA methylation, miRNA Diagnosis for breast carcinoma PAM50 subtype Normal: 115 / Basal: 131 / Her2: 46 / LumA: 436 / LumB: 147 mRNA, DNA methylation, miRNA Diagnosis for Alzheimer’s Disease Normal: 169 / AD: 182 HRCT scans, 81 clinical features Prognosis for COVID-19 morbidity outcome Severe symptoms: 202 / Mild symptoms: 549 2D fundus images, 3D OCT scans Diagnosis for Glaucoma Grad Non: 50 / Early: 26 / Mid advanced: 24 rs-fMRI graphs, 3D sMRI graphs Diagnosis for Major Depressive Disorder Major Depressive Disorder: 1300 / Healthy Controls: 1128 (830 / 771 after QC)

4.2

Task

Number of labels and patients

Quantitative Analysis

BRCA & ROSMAP. For these multi-omics datasets, we compare against five representative integration methods covering the main families: classical statistical integration (DIABLO [55]), graph-based deep integration (MOGONET [10]), uncertainty-aware fusion (TMC [56], Dynamic [11], and DMIB [12], the closest competitor as it is also IB-based but lacks any explanation mechanism. As shown in Table 2, SMILE achieves the best ACC, WeightedF1, and MacroF1 on BRCA (87.3%, 87.8%, 84.6%); the largest margin appears on MacroF1, indicating that the selected features remain discriminative for minority PAM50 subtypes. On the binary ROSMAP task, SMILE leads in ACC (85.1%) and F1 (85.6%), while DMIB retains the best AUC (91.6% vs. 90.3%). SMILE thus matches or exceeds the strongest IB-based competitor while additionally providing instance-wise explanations. Table 2: Performance of various multimodal methods on BRCA and ROSMAP datasets. The best results are in bold, and the second-best results are underlined. BRCA

ROSMAP

Method

ACC

WeightedF1

MacroF1

ACC

F1

AUC

DIABLO [55] MOGONET [10] TMC [56] Dynamic [11] DMIB [12] SMILE (Ours)

64.2±0.9 82.9±1.8 84.2±0.5 87.1±0.5 86.0±0.7 87.3±0.2

53.4±1.7 82.5±1.7 84.4±0.9 87.4±0.6 86.0±0.8 87.8±0.3

36.9±1.7 77.4±1.7 80.6±0.9 83.5±0.5 81.6±0.9 84.6±0.3

74.2±2.4 81.5±2.3 82.5±0.9 81.7±1.5 84.9±1.8 85.1±1.1

75.5±2.5 82.1±1.2 82.3±0.6 82.3±1.5 85.3±1.7 85.6±1.3

83.0±2.5 87.4±1.2 88.5±0.6 90.0±1.2 91.6±0.7 90.3±0.4

Table 3: Performance of various multimodal methods on iCTCF dataset. SMILE results are reported as the mean and standard deviation across 10 runs of 5-fold cross-validation.

De-COVID19-Net [57] HoFN+SCResNet [58] HUST-19 [50] DMIB [12] SMILE (ours)

ACC

AUC

JW0.5

JW0.6

72.4 68.4 83.3 73.2 92.4±0.5

77.3 80.2 92.1 82.3 95.6±0.6

69.8 70.1 – 72.1 84.3±0.5

69.2 71.6 – 71.7 82.2±0.3

iCTCF. As this dataset pairs imaging with clinical features and has dedicated models built on it, we compare against three COVID-19-specific prognosis models (De-COVID19 Net [57], HoFN+SCResNet [58], and HUST-19 [50], released with the dataset) plus the general-purpose DMIB [12], and follow [12] in adopting ACC, AUC and the weighted Youden indices JW 0.5 , JW 0.6 , which reflect the asymmetric clinical cost of missing severe cases. Table 3 shows that SMILE raises ACC from 83.3% to 92.4% (9.1 percentage points over HUST-19) and attains the highest AUC 10

Table 4: Performance comparison of various multimodal methods on the REST-meta-MDD dataset (%). Best results are highlighted in bold. Our proposed SMILE achieves compelling performance in terms of all four metrics.

DER [59] MoNIG [60] MIB [32] MEIB [61] SMILE (ours)

ACC

F1-score

AUC

MCC

56.3 63.6 55.9 52.2 67.4±2.3

41.3 64.8 70.1 68.5 69.7±2.6

66.8 67.9 77.7 82.2 73.5±3.6

17.4 27.0 21.3 6.7 35.0±4.6

Table 5: Performance of different strategies on Glaucoma Grading dataset. To compare directly with pre-trained models used for the classification task [51], we use them as encoders.

Kappa

DuelRes

DuelRes (SMILE)

Res-DEN

Res-DEN (SMILE)

65.4±1.2

69.2±0.8

70.1±0.5

72.4±0.5

(95.6%); among methods reporting the Youden indices, it also obtains the best JW 0.5 (84.3%) and JW 0.6 (82.2%), confirming that the gain is not driven by the majority class. Glaucoma Grading. To compare with baseline models [51], we replace encoders with pre-trained models (excluding the classifier layer) and keep Cohen’s Kappa as the evaluation metric. DuelRes with SMILE achieves a notable improvement with a Kappa score of 69.2% compared to 65.4% and presents better stability. Similarly, Res-DEN with SMILE reaches a Kappa score of 72.4% compared to 70.1% (see Table 5). This demonstrates that involving the pre-trained models in our framework consistently improves performance. REST-meta-MDD. Table 4 compares SMILE against four multimodal baselines on the RESTmeta-MDD dataset. SMILE achieves the highest accuracy (67.4%), exceeding the strongest baseline in accuracy, MoNIG [60], by 3.8 percentage points, and the largest margin appears on MCC (35.0% vs. 27.0%), the metric least affected by majority-class bias. Its F1-score (69.7%) is on par with the best baseline (MIB, 70.1%). Two baselines attain higher AUC, yet this does not translate into usable decisions: MEIB [61] reaches the highest AUC (82.2%) but collapses to 52.2% accuracy with near-chance MCC (6.7%), and MIB [32] likewise pairs a high AUC (77.7%) with 55.9% accuracy and an MCC of 21.3%. Similarly, DER [59] shows a wide gap between accuracy and F1-score (56.3% vs. 41.3%). These results suggest that SMILE provides a more balanced operating point under the default decision threshold.

4.3

Qualitative Analysis

The most important innovation of SMILE is its ability to generate modality-specific explanations. We present qualitative examples to show its effectiveness. For genomic data, we visualize the top-5 biomarkers identified by SMILE across three modalities of BRCA, which is shown at the top of Fig. 4. Meanwhile, we reference some representative studies to illustrate the relevance of these biomarkers to breast cancer progression. Specifically, for mRNA expression in Fig. 4a, in HER2-positive breast cancers, one of the most aggressive subtypes of breast cancer, KCTD10 can induce RhoB degradation and activation of Rac1 [62]. TEX10 is illustrated in [63], which finds that its interaction with the NF-κB pathway promotes tumor progression by 11

(a)

(b)

(c)

Figure 4: The top-5 biomarkers among the k = 30 features selected by SMILE are shown for two datasets (top row: BRCA; bottom row: ROSMAP) across three modalities: (a) mRNA expression, (b) DNA methylation, and (c) miRNA expression. enhancing the expression of key genes involved in inflammation and survival. [64] claims that the degradation of PHGDH by ring finger protein 5 inhibited breast cancer cell growth. TP53INP1 is an independent risk factor for overall survival in breast cancer patients [65]. For DNA methylation expression in Fig. 4b, ZCCHC11 [66], also known as TUT4, has been verified to be associated with breast cancer. TEX19 [67], SNORD82 [68] and FZD7 [69] are relevant to breast cancer and its diagnosis. Similar to miRNA expression in Fig. 4c, miR-520b [70], miR-874 [71] and miR-665 [72], all of these identified gene sequences are associated with breast cancer. In the ROSMAP dataset, the top-5 biomarkers identified by SMILE across three modalities are shown at the bottom of Fig. 4. These biomarkers are supported by leading medical journals linking them to Alzheimer’s Disease (AD). For mRNA in Fig. 4a, QDPR has been observed to be defective in the AD brain because it catalyzes the conversion of BH4 away from these neurotransmitters [73, 74]. [75, 76] indicate that DDIT4 is differentially expressed in the prefrontal cortex of AD patients. SLC5A11 is included in the top 5 cluster-enriched transcripts per astrocytes and oligodendrocytes which change in AD [77], while MPV17L2 is also recognized as a high-confidence AD-associated gene [78]. For DNA methylation in Fig. 4b, the dataset contains Cytosine-phosphate-Guanine (GpG) sites (e.g., cg22472290 ). We are able to find the relevance of their associated genes to the task, such as cg22472290 in 2NF577 [79], cg17253459 in OLFML3 [80]. The identified miRNAs in Fig. 4c, including miR-199a-5p [81], miR-346 [82], miR-484 [83], miR-129-3p [84] and miR-133a [85], have all been associated with breast cancer. Fig. 5 presents qualitative examples of iCTCF, which contains two modalities. We find that the generated visual explanations (see Fig. 5a) successfully localize COVID-19-related lung features, such as Consolidation, Ground-Glass Opacities (GGOs), and Crazy-paving patterns, for 2D montage images. For selected clinical information in Fig. 5b, ALG, LYP, CRP, GGT, LDH, NEP, and AST are task-relevant references [50]. Fig. 6 presents visual explanations for two modalities in Glaucoma Grading, demonstrating that SMILE effectively identifies key features in both multi-channel and single-channel images. Fig. 7 demonstrates the extracted explanations from both fMRI and sMRI modalities on the

12

Crazy-paving GGO

Consolidations

(a)

(b)

Figure 5: iCTCF visualizations on two modalities. (a) Qualitative examples of CT scan in the 2D montage: each explanation contains k = 30 patches. The left rows present severe patients’ samples; the right rows present mild patients’ samples. (b) Qualitative examples of clinical information: normalized results from 150 test samples, k = 20 features selected. RNFL RNFL

GCIPL GCIPL

Optical cup Optical cup

Macular fovea Macular fovea

Low Contribution Low Contribution

High Contribution High Contribution

(a)

(b)

Figure 6: Glaucoma Grading visualizations on two modalities. (a) Qualitative examples of color fundus (k = 100); (b) Qualitative examples of OCT slices (k = 100) REST-meta-MDD dataset. To examine what SMILE attends to at the population level, we aggregated per-subject selections into frequencies within each group. The structural selector returned sparse, anatomically coherent subvolumes rather than scattered voxels, and the functional selector a compact rather than diffuse edge set. In six of ten runs, structural selection converged on a posterior occipital–parietal territory, indicating that the pattern reflects the learned model rather than a single initialization. This territory coincides with regions of reduced cortical thickness identified in a vertex-wise meta-analysis of 64 cohorts from the ENIGMA MDD and DIRECT consortia—the latter supplied our data—including the lateral occipital and parietal cortices [86], and with morphometric findings on REST-meta-MDD itself [87]. The selected edges center on occipito-temporal pathways, complementing the default-mode alterations reported previously in this cohort [52]. One caveat: the HC and MDD maps overlap substantially, so they evidence anatomically plausible and reproducible attention, not group-discriminative biomarkers. The co-localization of sMRI subvolumes with fMRI edge endpoints nonetheless indicates that the model draws on convergent structural and functional evidence.

13

sMRI Patches

sMRI Patches

fMRI circular connectome

fMRI circular connectome fMRI Brain Plot

fMRI Brain Plot

(a)

(b)

Figure 7: REST-meta-MDD biomarker visualizations on two modalities. (a) HC (k = 25); (b) MDD (k = 25). Top: sMRI subvolumes on axial slices. Bottom left: selected rs-fMRI edges on the cortical surface, nodes colored by AAL-116 ROI. Bottom right: circular connectome of the same edges with ROI labels. Color encodes selection frequency. Descriptive aggregates, not statistical comparisons; one run, whose posterior structural pattern recurred in 6 of 10 runs.

Figure 8: Qualitative examples of explanations generated by LIME, SHAP, and SMILE on different datasets. For CT scans in iCTCF dataset, we select k = 30 with the patch size of 70 × 70; for OCT slices in Glaucoma Grading dataset, k = 100 with the patch size of 4 × 4.

4.4

Comparison to Post-hoc Explanations

A central claim of self-explainable modelling is that explanations optimized jointly with the predictor are more faithful than those attached post hoc. To ensure a fair comparison, we apply LIME and SHAP to Multimodal IB [32] with late fusion (L-MIB), which uses the same encoders as SMILE; the two models differ only in the built-in explainer. We evaluate the image modality with two metrics, both lower-is-better. Fidelity is the drop in predicted-class probability when only the regions marked as important are retained. Consistency is maximum sensitivity [88]: the largest change in the explanation under a small input perturbation. SMILE attains the best value on both (Table 6), with visual examples in Fig. 8. 14

Table 6: Comparison between built-in (ours) and post-hoc explanations. A lower value indicates better performance for metrics.

Dataset iCTCF Glaucoma

4.5

Fidelity (average confidence drop[%] ↓ )

Consistency (maximum sensitivity↓)

LIME 36.4±2.1 12.0±0.8

LIME 0.48 0.39

SHAP 48.1±1.8 15.1±1.4

SMILE 30.2±0.9 9.2±0.6

SHAP 0.37 0.41

SMILE 0.22 0.27

Ablation Study

In this section, we analyze the contributions of the explainer module to quantify the effectiveness of the generated explanations. We consider two variants of SMILE: (i) without specific modality X m , which serves as a modality-ablation baseline. (ii) with a random mask Er , which modifies (5) s.t. we use a random mask rim instead of explanation em i from the explainer. We present their performance with ACC scores (Kappa scores for the Glaucoma grading dataset) in Table 7. The results reveal two key insights: (i) Modalities interact synergistically to achieve high performance, highlighting the need to assess each modality’s contribution in multimodal contexts. (ii) While random masks in information bottleneck training improve performance over the baseline, they introduce higher variance. In contrast, our explainer captures more decision-relevant information and exhibits greater stability than the random mask approach.

5

Conclusion and Future Work

We proposed SMILE, a self-explainable multimodal framework grounded in the IB principle for AI-based medical diagnosis. SMILE comprises three modules: modality-specific explainers, encoders mapping medical data to low-dimensional latent representations, and a fused classifier. Theoretically, under the single-modality setting and with β = 0, the SMILE objective reduces to the widely used INVASE objective as a special case. Moreover, our analysis shows that introducing the IB-based compression term does not necessarily impose an intrinsic accuracy penalty; instead, it can improve the generalization bound by controlling the amount of information retained in the selected representations. Across five representative multimodal medical datasets spanning heterogeneous modalities, SMILE exceeds state-of-the-art diagnostic performance while producing faithful modality-specific explanations. Our future work is twofold. First, the sufficient encoder assumption [15] in (5) simplifies optimization but remains over-optimistic; estimating or approximating I(X; X̃) without it is an open problem, even from an information theory perspective. Second, SMILE assumes that all modalities may contribute and focuses on modality-specific informative features, leaving cross-modal structure implicit. Moving forward, we aim to explicitly model high-order modality interaction terms such as synergy.

15

Table 7: Ablation study results on different datasets. The modality mentioned refers to the object we operate on. Dataset

BRCA/ ROSMAP

Modality

w/o X m

w/ Er

mRNA DNA meth miRNA

80.4±1.1/78.6±1.8 84.2±0.7/81.2±1.4 83.9±0.9/79.0±1.5

84.7±1.9/82.3±2.0 85.0±1.5/81.8±1.8 84.3±1.5/82.1±1.7

mRNA+DNA meth mRNA+miRNA miRNA+DNA meth

68.4±1.4/63.7±1.8 66.7±1.8/61.5±2.0 71.4±1.1/65.8±1.3

77.8±4.2/75.4±5.2 75.9±4.7/71.1±5.8 79.6±3.9/78.3±4.9

All

–

70.7±8.3/67.1±9.7

Ours: 87.3±0.2/85.1±1.1

iCTCF

2D Montage Clinical features

66.8±2.3 68.9±2.7

82.2±4.9 83.6± 3.6

Both

–

80.3± 7.2

Ours: 92.4±0.5 Glaucoma Grading

Color fundus OCT slices

64.3±2.1 56.5±3.4

68.3±3.9 65.9±4.8

Both

–

67.1±6.2

Ours: 72.4±0.5

REST-meta-MDD

rs-fMRI sMRI

61.0±2.2 52.5±1.1

61.9 ±1.0 56.6±6.6

Both

–

63.9±0.9

Ours: 67.4±2.3

16

References [1] Andreas Holzinger, Georg Langs, Helmut Denk, Kurt Zatloukal, and Heimo Müller. Causability and explainability of artificial intelligence in medicine. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 9(4):e1312, 2019. [2] Zichang He, Bo Zhao, and Zheng Zhang. Active sampling for accelerated mri with low-rank tensors. In 2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), pages 3024–3028. IEEE, 2022. [3] Xun Jia, Lei Ren, and Jing Cai. Clinical implementation of ai technologies will require interpretable ai models. Medical physics, (1):1–4, 2020. [4] Wojciech Samek, Grégoire Montavon, Andrea Vedaldi, Lars Kai Hansen, and Klaus-Robert Müller. Explainable AI: interpreting, explaining and visualizing deep learning, volume 11700. Springer Nature, 2019. [5] Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradientbased localization. In Proceedings of the IEEE international conference on computer vision, pages 618–626, 2017. [6] Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. " why should i trust you?" explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, pages 1135–1144, 2016. [7] Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, page 4768–4777, 2017. [8] Selene Gallo et al. Functional connectivity signatures of major depressive disorder: machine learning analysis of two multicenter neuroimaging studies. Molecular Psychiatry, 28(7):3013– 3022, 2023. [9] Sutong Wang, Yunqiang Yin, Dujuan Wang, Yanzhang Wang, and Yaochu Jin. Interpretabilitybased multimodal convolutional neural networks for skin lesion diagnosis. IEEE Transactions on Cybernetics, 52(12):12623–12637, 2021. [10] Tongxin Wang, Wei Shao, Zhi Huang, Haixu Tang, Jie Zhang, Zhengming Ding, and Kun Huang. Mogonet integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nature communications, 12(1):3445, 2021. [11] Zongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang, and Jianhua Yao. Multimodal dynamics: Dynamical fusion for trustworthy multimodal classification. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 20707–20717, 2022. [12] Yingying Fang, Shuang Wu, Sheng Zhang, Chaoyan Huang, et al. Dynamic multimodal information bottleneck for multimodality classification. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 7696–7706, 2024. [13] Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning, pages 2668–2677. PMLR, 2018. 17

[14] Naftali Tishby, Fernando C. Pereira, and William Bialek. The information bottleneck method. In Proc. of the 37-th Annual Allerton Conference on Communication, Control and Computing, pages 368–377, 1999. [15] Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. What makes for good views for contrastive learning? Advances in neural information processing systems, 33:6827–6839, 2020. [16] Luis Gonzalo Sanchez Giraldo, Murali Rao, and Jose C Principe. Measures of entropy from data using infinitely divisible kernels. IEEE Transactions on Information Theory, 61(1):535–548, 2014. [17] Shujian Yu, Luis Gonzalo Sanchez Giraldo, Robert Jenssen, and Jose C Principe. Multivariate extension of matrix-based rényi’s α-order entropy functional. IEEE transactions on pattern analysis and machine intelligence, 42(11):2960–2966, 2019. [18] Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and Devon Hjelm. Mutual information neural estimation. In International conference on machine learning, pages 531–540. PMLR, 2018. [19] Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. Smoothgrad: removing noise by adding noise. arXiv preprint arXiv:1706.03825, 2017. [20] Pang Wei Koh and Percy Liang. Understanding black-box predictions via influence functions. In International conference on machine learning, pages 1885–1894. PMLR, 2017. [21] Peifei Zhu and Masahiro Ogino. Guideline-based additive explanation for computer-aided diagnosis of lung nodules. In International Workshop on Multimodal Learning for Clinical Decision Support, pages 39–47, 2019. [22] Weina Jin, Xiaoxiao Li, Mostafa Fatehi, and Ghassan Hamarneh. Guidelines and evaluation of clinical explainable ai in medical image analysis. Medical image analysis, 84:102684, 2023. [23] Jonathan Crabbé and Mihaela Van Der Schaar. Explaining time series predictions with dynamic masks. In International Conference on Machine Learning, pages 2166–2177. PMLR, 2021. [24] Zhitao Ying, Dylan Bourgeois, Jiaxuan You, Marinka Zitnik, and Jure Leskovec. Gnnexplainer: Generating explanations for graph neural networks. Advances in neural information processing systems, 32, 2019. [25] Rosa CJ Kraaijveld, Marielle EP Philippens, Wietse SC Eppinga, Ina M Jürgenliemk-Schulz, Kenneth GA Gilhuijs, Petra S Kroon, and Bas HM van der Velden. Multi-modal volumetric concept activation to explain detection and classification of metastatic prostate cancer on psma-pet/ct. In International Workshop on Interpretability of Machine Intelligence in Medical Image Computing, pages 82–92, 2022. [26] Aurélie Pahud de Mortanges, Haozhe Luo, Shelley Zixin Shu, Amith Kamath, Yannick Suter, Mohamed Shelan, Alexander Pöllinger, and Mauricio Reyes. Orchestrating explainable artificial intelligence for multimodal and longitudinal data in medical imaging. NPJ digital medicine, 7 (1):195, 2024.

18

[27] Saurabh Varshneya, Antoine Ledent, Philipp Liznerski, Andriy Balinskyy, Purvanshi Mehta, Waleed Mustafa, and Marius Kloft. Interpretable tensor fusion. arXiv preprint arXiv:2405.04671, 2024. [28] Laura Wenderoth, Konstantin Hemker, Nikola Simidjievski, and Mateja Jamnik. Measuring cross-modal interactions in multimodal models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 21501–21509, 2025. [29] Ravid Shwartz-Ziv and Naftali Tishby. Opening the black box of deep neural networks via information. arXiv preprint arXiv:1703.00810, 2017. [30] Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy. Deep variational information bottleneck. arXiv preprint arXiv:1612.00410, 2016. [31] Jingqi Song, Yuanjie Zheng, Jing Wang, Muhammad Zakir Ullah, and Wanzhen Jiao. Multicolor image classification using the multimodal information bottleneck network (mmib-net) for detecting diabetic retinopathy. Optics Express, 29(14):22732–22748, 2021. [32] Sijie Mai, Ying Zeng, and Haifeng Hu. Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations. IEEE Transactions on Multimedia, 25: 4121–4134, 2022. [33] Qilong Wu, Yiyang Shao, Jun Wang, and Xiaobo Sun. Learning optimal multimodal information bottleneck representations. In International Conference on Machine Learning, pages 67584– 67622. PMLR, 2025. [34] Seojin Bang, Pengtao Xie, Heewook Lee, Wei Wu, and Eric Xing. Explaining a black-box by using a deep variational information bottleneck approach. In Proceedings of the AAAI conference on artificial intelligence, volume 35, pages 11396–11404, 2021. [35] Junchi Yu, Tingyang Xu, Yu Rong, Yatao Bian, Junzhou Huang, and Ran He. Graph information bottleneck for subgraph recognition. In International Conference on Learning Representations, 2021. [36] Kaizhong Zheng, Shujian Yu, Baojuan Li, Robert Jenssen, and Badong Chen. Brainib: Interpretable brain network-based psychiatric diagnosis with graph information bottleneck. IEEE Transactions on Neural Networks and Learning Systems, 36(7):13066–13079, 2024. [37] Ying Wang, Tim GJ Rudner, and Andrew G Wilson. Visual explanations of image-text representations via multi-modal information bottleneck attribution. Advances in Neural Information Processing Systems, 36:16009–16027, 2023. [38] Pauline Bourigault, Emmanuelle Bourigault, and Danilo P Mandic. Multi-modal information bottleneck attribution with cross-attention guidance. BMVC, 2024. [39] Jinsung Yoon, James Jordon, and Mihaela Van der Schaar. Invase: Instance-wise variable selection using neural networks. In International conference on learning representations, 2019. [40] Michael Tschannen, Josip Djolonga, Paul K Rubenstein, Sylvain Gelly, and Mario Lucic. On mutual information maximization for representation learning. arXiv preprint arXiv:1907.13625, 2019.

19

[41] Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. Invariant risk minimization. arXiv preprint arXiv:1907.02893, 2019. [42] Alireza Makhzani and Brendan Frey. K-sparse autoencoders. arXiv preprint arXiv:1312.5663, 2013. [43] Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016. [44] Artemy Kolchinsky, Brendan D Tracey, and David H Wolpert. Nonlinear information bottleneck. Entropy, 21(12):1181, 2019. [45] Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. [46] Kenji Kawaguchi, Zhun Deng, Xu Ji, and Jiaoyang Huang. How does information bottleneck help deep learning? In International Conference on Machine Learning, pages 16049–16096. PMLR, 2023. [47] Dang Minh, H Xiang Wang, Y Fen Li, and Tan N Nguyen. Explainable artificial intelligence: a comprehensive review. Artificial Intelligence Review, pages 1–66, 2022. [48] Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. Interpretable machine learning: Fundamental principles and 10 grand challenges. Statistic Surveys, 16:1–85, 2022. [49] Junlin Hou, Sicen Liu, Yequan Bie, Hongmei Wang, Andong Tan, Luyang Luo, and Hao Chen. Self-explainable ai for medical image analysis: A survey and new outlooks. arXiv preprint arXiv:2410.02331, 2024. [50] Wanshan Ning, Shijun Lei, Jingjing Yang, Yukun Cao, Peiran Jiang, Qianqian Yang, Jiao Zhang, Xiaobei Wang, Fenghua Chen, Zhi Geng, et al. Open resource of clinical data from patients with pneumonia for the prediction of covid-19 outcomes via deep learning. Nature biomedical engineering, 4(12):1197–1207, 2020. [51] Junde Wu, Huihui Fang, Fei Li, Huazhu Fu, Fengbin Lin, Jiongcheng Li, Yue Huang, Qinji Yu, Sifan Song, Xinxing Xu, et al. Gamma challenge: glaucoma grading from multi-modality images. Medical Image Analysis, 90:102938, 2023. [52] Chao-Gan Yan, Xiao Chen, Le Li, Francisco Xavier Castellanos, Tong-Jian Bai, Qi-Jing Bo, Jun Cao, Guan-Mao Chen, Ning-Xuan Chen, Wei Chen, et al. Reduced default mode network functional connectivity in patients with recurrent major depressive disorder. Proceedings of the National Academy of Sciences, 116(18):9078–9083, 2019. [53] Nathalie Tzourio-Mazoyer, Brigitte Landeau, Dimitri Papathanassiou, Fabrice Crivello, Octave Etard, Nicolas Delcroix, Bernard Mazoyer, and Marc Joliot. Automated anatomical labeling of activations in spm using a macroscopic anatomical parcellation of the mni mri single-subject brain. Neuroimage, 15(1):273–289, 2002. [54] Huihui Fang, Fangxin Shang, Huazhu Fu, Fei Li, Xiulan Zhang, and Yanwu Xu. Multi-modality images analysis: A baseline for glaucoma grading via deep learning. In Ophthalmic Medical Image Analysis: 8th International Workshop, OMIA 2021, Held in Conjunction with MICCAI 2021, Strasbourg, France, September 27, 2021, Proceedings 8, pages 139–147. Springer, 2021. 20

[55] Amrit Singh, Casey P Shannon, Benoît Gautier, Florian Rohart, Michaël Vacher, Scott J Tebbutt, and Kim-Anh Lê Cao. Diablo: an integrative approach for identifying key molecular drivers from multi-omics assays. Bioinformatics, 35(17):3055–3062, 2019. [56] Zongbo Han, Changqing Zhang, Huazhu Fu, and Joey Tianyi Zhou. Trusted multi-view classification with dynamic evidential fusion. IEEE transactions on pattern analysis and machine intelligence, 45(2):2551–2566, 2022. [57] Lingwei Meng, Di Dong, Liang Li, Meng Niu, Yan Bai, Meiyun Wang, Xiaoming Qiu, Yunfei Zha, and Jie Tian. A deep learning prognosis model help alert for covid-19 patients at high-risk of death: a multi-center study. IEEE journal of biomedical and health informatics, 24(12): 3576–3584, 2020. [58] Jinzhao Zhou, Xingming Zhang, Ziwei Zhu, Xiangyuan Lan, Lunkai Fu, Haoxiang Wang, and Hanchun Wen. Cohesive multi-modality feature learning and fusion for covid-19 patient severity prediction. IEEE Transactions on Circuits and Systems for Video Technology, 32(5):2535–2549, 2021. [59] Alexander Amini, Wilko Schwarting, Ava Soleimany, and Daniela Rus. Deep evidential regression. Advances in neural information processing systems, 33:14927–14937, 2020. [60] Huan Ma, Zongbo Han, Changqing Zhang, Huazhu Fu, Joey Tianyi Zhou, and Qinghua Hu. Trustworthy multimodal regression with mixture of normal-inverse gamma distributions. Advances in Neural Information Processing Systems, 34:6881–6893, 2021. [61] Qi Zhang, Shujian Yu, Jingmin Xin, and Badong Chen. Multi-view information bottleneck without variational approximation. In ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 4318–4322. IEEE, 2022. [62] Annapaola Angrisani, Annamaria Di Fiore, Enrico De Smaele, and Marta Moretti. The emerging role of the kctd proteins in cancer. Cell Communication and Signaling, 19(1):56, 2021. [63] Ziyang Wang, Chunjie Sheng, Guangyan Kan, et al. Rnai screening identifies that tex10 promotes the proliferation of colorectal cancer cells by increasing nf-κb activation. Advanced Science, 7(17):2000593, 2020. [64] Chae Min Lee, Yeseong Hwang, Minki Kim, et al. Phgdh: a novel therapeutic target in cancer. Experimental & Molecular Medicine, 56(7):1513–1522, 2024. [65] Mayumi Nishimoto, Sayaka Nishikawa, Naoto Kondo, et al. Prognostic impact of tp53inp1 gene expression in estrogen receptor α-positive breast cancer patients. Japanese journal of clinical oncology, 49(6):567–575, 2019. [66] Pengcheng Zhang, Mallory I Frederick, and Ilka U Heinemann. Terminal uridylyltransferases tut4/7 regulate microrna and mrna homeostasis. Cells, 11(23):3742, 2022. [67] Huantao Liu, He Wang, Hongyu Zhang, Miaomiao Yu, and Yu Tang. Tex19 increases the levels of cdk4 and promotes breast cancer by disrupting skp2-mediated cdk4 ubiquitination. Cancer Cell International, 24(1):207, 2024.

21

[68] Emmi Kärkkäinen, Sami Heikkinen, Maria Tengström, Veli-Matti Kosma, et al. Expression profiles of small non-coding rnas in breast cancer tumors characterize clinicopathological features and show prognostic and predictive potential. Scientific Reports, 12(1):22614, 2022. [69] L Yang, X Wu, Y Wang, K Zhang, J Wu, YC Yuan, X Deng, L Chen, CCH Kim, S Lau, et al. Fzd7 has a critical role in cell proliferation in triple negative breast cancer. Oncogene, 30(43): 4437–4446, 2011. [70] Ya-Ching Lu, Ann-Joy Cheng, Li-Yu Lee, et al. Mir-520b as a novel molecular target for suppressing stemness phenotype of head-neck cancer by inhibiting cd44. Scientific reports, 7 (1):2042, 2017. [71] Alimasi Aersilan, Naoko Hashimoto, Kazuyuki Yamagata, et al. Microrna-874 targets phosphomevalonate kinase and inhibits cancer cell growth via the mevalonate pathway. Scientific reports, 12(1):18443, 2022. [72] Xin-Ge Zhao, Jing-Ye Hu, Jun Tang, et al. mir-665 expression predicts poor survival and promotes tumor metastasis by targeting nr4a3 in breast cancer. Cell death & disease, 10(7): 479, 2019. [73] Jingshu Xu, Stefano Patassini, Nitin Rustogi, Isabel Riba-Garcia, Benjamin D Hale, Alexander M Phillips, Henry Waldvogel, Robert Haines, Phil Bradbury, Adam Stevens, et al. Regional protein expression in human alzheimer’s brain correlates with disease severity. Communications biology, 2(1):43, 2019. [74] Hansruedi Mathys, Jose Davila-Velderrain, Zhuyu Peng, Fan Gao, Shahin Mohammadi, Jennie Z Young, Madhvi Menon, Liang He, Fatema Abdurrob, Xueqiao Jiang, et al. Single-cell transcriptomic analysis of alzheimer’s disease. Nature, 570(7761):332–337, 2019. [75] Saranya Canchi, Balaji Raao, Deborah Masliah, Sara Brin Rosenthal, Roman Sasik, Kathleen M Fisch, Philip L De Jager, David A Bennett, and Robert A Rissman. Integrating gene and protein expression reveals perturbed functional networks in alzheimer’s disease. Cell reports, 28(4):1103–1116, 2019. [76] Leticia Pérez-Sisqués, Anna Sancho-Balsells, Júlia Solana-Balaguer, et al. Rtp801/redd1 contributes to neuroinflammation severity and memory impairments in alzheimer’s disease. Cell Death & Disease, 12(6):616, 2021. [77] Jessica S Sadick, Michael R O’Dea, Philip Hasel, Taitea Dykstra, Arline Faustin, and Shane A Liddelow. Astrocytes and oligodendrocytes undergo subtype-specific transcriptional changes in alzheimer’s disease. Neuron, 110(11):1788–1805, 2022. [78] Daniel Zhang, Timothy Penwell, Yan-Hua Chen, Addison Koehler, Rui Wu, Shayan Nik Akhtar, and Qun Lu. G-protein signaling in alzheimer’s disease: Spatial expression validation of semi-supervised deep learning based computational framework. Journal of Neuroscience, 2024. [79] Roy Lardenoije, Janou AY Roubroeks, Ehsan Pishva, Markus Leber, Holger Wagner, Artemis Iatrou, Adam R Smith, Rebecca G Smith, Lars MT Eijssen, Luca Kleineidam, et al. Alzheimer’s disease-associated (hydroxy) methylomic changes in the brain and blood. Clinical epigenetics, 11:1–15, 2019.

22

[80] Eleanor Drummond, Tomas Kavanagh, Geoffrey Pires, Mitchell Marta-Ariza, Evgeny Kanshin, Shruti Nayak, Arline Faustin, Valentin Berdah, Beatrix Ueberheide, and Thomas Wisniewski. The amyloid plaque proteome in early onset alzheimer’s disease and down syndrome. Acta neuropathologica communications, 10(1):53, 2022. [81] Dandan Song, Guoxiang Li, Yu Hong, Pan Zhang, Jingling Zhu, Lei Yang, and Jin Huang. mir-199a decreases neuritin expression involved in the development of alzheimer’s disease in app/ps1 mice. International Journal of Molecular Medicine, 46(1):384–396, 2020. [82] Justin M Long, Bryan Maloney, Jack T Rogers, and Debomoy K Lahiri. Novel upregulation of amyloid-β precursor protein (app) by microrna-346 via targeting of app mrna 5-untranslated region: Implications in alzheimer’s disease. Molecular psychiatry, 24(3):345–363, 2019. [83] Yin-zhao Jia, Jing Liu, Geng-qiao Wang, and Zi-fang Song. mir-484: a potential biomarker in health and disease. Frontiers in oncology, 12:830420, 2022. [84] Ellis Patrick, Sathyapriya Rajagopal, Hon-Kit Andus Wong, Cristin McCabe, Jishu Xu, Anna Tang, Selina H Imboywa, Julie A Schneider, Nathalie Pochet, Anna M Krichevsky, et al. Dissecting the role of non-coding rnas in the accumulation of amyloid and tau neuropathologies in alzheimer’s disease. Molecular neurodegeneration, 12:1–13, 2017. [85] Phoebe P Chum, Md A Hakim, and Erik J Behringer. Cerebrovascular microrna expression profile during early development of alzheimer’s disease in a mouse model. Journal of Alzheimer’s Disease, 85(1):91–113, 2022. [86] Chao-Gan Yan, Zi-Han Wang, Laura KM Han, Nina Alexander, Dag Alnæs, Zeynep Başgöze, Stéphanie EEC Bauduin, Jochen Bauer, Francesco Benedetti, Klaus Berger, et al. Vertex-wise cortical abnormalities in major depressive disorder from 64 cohorts from the direct and enigma mdd consortia. nature mental health, 4(7):1175–1186, 2026. [87] KangCheng Wang, YuFei Hu, ChaoGan Yan, MeiLing Li, YanJing Wu, Jiang Qiu, XingXing Zhu, REST meta MDD Consortium, et al. Brain structural abnormalities in adult major depressive disorder revealed by voxel-and source-based morphometry: evidence from the rest-meta-mdd consortium. Psychological medicine, 53(8):3672–3682, 2023. [88] Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I Inouye, and Pradeep K Ravikumar. On the (in) fidelity and sensitivity of explanations. Advances in neural information processing systems, 32, 2019. [89] Zoe Piran, Ravid Shwartz-Ziv, and Naftali Tishby. The dual information bottleneck. arXiv preprint arXiv:2006.04641, 2020. [90] Yury Polyanskiy and Yihong Wu. Information theory: From coding to learning. Cambridge university press, 2024. [91] Yihong Wu. Lecture notes on information-theoretic methods for high-dimensional statistics.

23

A

Proof to Proposition 1

Our proposed IB objective for single modality feature selection is expressed as: min −I(Y ; X̃) + βI(X̃; X) + λ∥E∥0 .

(13)

E

Proposition 3. The IB objective in Eq. (13) encompasses that of INVASE as a special case when β = 0, i.e., INVASE lacks a compression term. Proof. The INVASE formulates the learning of E in a single modality as: h



min EX DKL p(Y |X)∥p(Y |X̃)

i

E

+ λ∥E∥0 ,

(14)

where DKL refershto the Kullback-Leibleri(KL) divergence. The term EX DKL p(Y |X)∥p(Y |X̃) can be decomposed as follows: h



EX DKL p(Y |X)∥p(Y |X̃) Z

!

Z

p(Y |X) log

p(X)

= X

i

Y

"

p(Y |X) p(Y |X̃)

= Ep(X,Y ) log

p(Y |X) dY p(Y |X̃)

!

dX

!#

(15)

= −H(Y |X) + H(Y |X̃) = −H(Y |X) + H(Y ) − (−H(Y |X̃) + H(Y )) = I(X; Y ) − I(Y ; X̃), where we used the Markov chain property X̃ ← X → Y , i.e., P (Y |X) = Ph(Y |X, X̃). i From the above relationship, minimizing the expected KL divergence EX DKL p(Y |X)∥p(Y |X̃) is equivalent to maximizing the mutual information I(Y ; X̃), since I(X; Y ) is a fixed value which only depends on training data and is irrelevant to optimization [89]. Hence, the INVASE objective essentially amounts to: min −I(Y ; X̃) + λ∥E∥0 . E

B

(16)

Proof to Proposition 4

Proposition 4. With probability at least 1 − δ over the training data t = {xi , yi }N a i=1 drawn from  M , the generalization error ∆(t) = E t (x, y)) − data distribution p(x, y), where xi = {xm } ℓ(f i m=1 p(x,y) 1 PN t i=1 ℓ(f (xi , yi )) roughly obeys the following form: N

s

∆(t) = Õ 

PM

m m m I(X̃ ; X ) + 1  as N → ∞, N

where ℓ is a bounded per-sample loss, f t is the full model obtained by training over t. 24

(17)

Proof. Our proof is based on Theorem 1[46]. Theorem 1. With probability at least 1 − δ over the training data t = {xi , yi }N i=1 drawn from a  1 PN t data distribution p(x, y), the generalization error ∆(t) = Ep(x,y) ℓ(f (x, y)) − N i=1 ℓ(f t (xi , yi )) roughly obeys the following form: s

∆(t) = Õ 

I(X; Zlt ) + 1  as N → ∞, N

(18)

where ℓ is a bounded per-sample loss, f t is the full model obtained by training over t, Zlt is representation obtained after passing X through the first l layers of model f t . In our case, we can treat {X m }M i=1 as a joint information source, denoted by X, which transmits information through a parameterized channel to Z̄ and subsequently to Y (see also Fig. 3 in the main text). Then, according to Theorem 1, s

∆(t) = Õ 

I(X; Z̄) + 1  as N → ∞. N

(19)

Further, we can treat {Z m }M i=1 as a joint information source, denoted by Z. Due to the data processing inequality (Z̄ = gω (Z)), the generalization error can be upper bounded by: s

∆(t) = Õ 

I(X; Z) + 1  as N → ∞. N

(20)

Lastly, we note that each Z m only depends on the corresponding X m through the modalityspecific encoder, i.e., X 1 → Z 1, X 2 → Z 2,

(21)

··· XM → ZM . Then the conditional distribution of Z given X factorized as: p(Z|X) =

M Y

p(Z m |X m ).

(22)

m=1

So as long the channels are decoupled like this, we always have [90][91, Chapter 12]: I(X; Z) ≤

M X

I(X m ; Z m ),

(23)

m=1

which, due to the data processing inequality, can be further upper bounded by: I(X; Z) ≤

M X

I(X m ; X̃ m ).

(24)

m=1

Hence, the generalization error ∆(t) scales as Õ

25

rP

M I(X̃ m ;X m )+1 m

N

!

.

C

Implementation Details

C.0.1

BRCA

mRNA encoder. The mRNA encoder takes the selected 1000-dimensional expression vector as input and maps it to a 500-dimensional latent representation via a fully connected projection layer, followed by ReLU activation and dropout with p = 0.5. The same encoder is applied to both the original and selected inputs, allowing the mutual-information regularizer to compare modality representations before and after selection. mRNA selector. The mRNA selector is an MLP with layer dimensions 1000 → 2000 → 2000 → 1000. GELU activations are applied after the first two linear layers, and the output layer produces one selection logit per mRNA feature. SMILE applies a differentiable top-k relaxation to these logits and obtains the hard explanation mask by thresholding at the k-th largest logit. We set k = 30 for this modality. DNA-methylation encoder. The DNA-methylation encoder follows the same fully connected design as the mRNA encoder. It maps the selected 1000-dimensional methylation vector to a 500-dimensional latent representation using a linear projection, ReLU activation, and dropout with p = 0.5. DNA-methylation selector. The DNA-methylation selector uses layer dimensions 1000 → 2000 → 2000 → 1000 with GELU activations between hidden layers. The selector produces featurelevel logits over all methylation variables, and the SMILE top-k operation retains k = 30 selected methylation features. miRNA encoder. The miRNA encoder maps the selected 503-dimensional miRNA-expression vector to a 500-dimensional latent representation. As in the other omics branches, the encoder consists of a fully connected projection followed by ReLU activation and dropout with p = 0.5. miRNA selector. The miRNA selector uses layer dimensions 503 → 1000 → 1000 → 503. Its output logits are converted into a sparse differentiable mask by the SMILE selector, and the selected mask keeps k = 30 miRNA features. Classifier head. The three selected modality representations are concatenated into a 1500dimensional multimodal representation. A fully connected classifier maps this fused representation to the five PAM50 subtype labels. Training details. We train the BRCA model for 2000 epochs using Adam with an initial learning rate of 10−4 and weight decay 10−4 . The learning rate is multiplied by 0.2 every 500 epochs. The mutual-information regularization weight is set to β = 0.08. C.0.2

ROSMAP

mRNA encoder. The mRNA encoder takes the selected 200-dimensional expression vector and projects it to a 300-dimensional latent representation with a fully connected layer, ReLU activation, and dropout with p = 0.5. mRNA selector. The mRNA selector is an MLP with layer dimensions 200 → 400 → 400 → 200. It outputs one selection logit per mRNA feature, and SMILE retains the top k = 30 features through the differentiable top-k selector. DNA-methylation encoder. The DNA-methylation encoder maps the selected 200dimensional methylation vector to a 300-dimensional latent representation through the same fully connected encoder used for the mRNA branch. DNA-methylation selector. The methylation selector uses layer dimensions 200 → 400 → 400 → 200 with GELU activations. The selector produces methylation-feature logits and keeps k = 30 methylation features. 26

miRNA encoder. The miRNA encoder maps the selected 200-dimensional miRNA-expression vector to a 300-dimensional latent representation using a fully connected projection, ReLU activation, and dropout with p = 0.5. miRNA selector. The miRNA selector also uses layer dimensions 200 → 400 → 400 → 200. The selected miRNA mask is obtained with the same differentiable top-k procedure, with k = 30. Classifier head. The three selected 300-dimensional modality embeddings are concatenated into a 900-dimensional representation. A fully connected classifier maps the fused representation to the binary diagnostic label. Training details. We train the ROSMAP model for 1000 epochs using Adam with an initial learning rate of 10−4 and weight decay 10−4 . The learning rate is multiplied by 0.2 every 200 epochs. The mutual-information regularization weight is set to β = 0.1. C.0.3

iCTCF

Clinical-feature encoder. The clinical encoder maps the selected 81-dimensional clinical vector to a 1024-dimensional latent representation. It consists of two fully connected layers with dimensions 81 → 64 → 1024, with ReLU activations after both layers. Clinical-feature selector. The clinical selector follows the non-image selector used for omics data. It uses an MLP with layer dimensions 81 → 100 → 100 → 81 and GELU activations between hidden layers. The output logits define feature-level selection scores over all clinical variables, and SMILE keeps k = 20 selected clinical features. HRCT encoder. The HRCT encoder takes a ten-channel 700 × 700 image tensor as input. It uses three convolutional blocks. The first block maps the input channels to 32 channels with a 3 × 3 convolution of stride 2, followed by ReLU and 2 × 2 max pooling. The second and third blocks use the same design and map 32 → 64 and 64 → 128 channels, respectively. The spatial resolution is reduced as 700 × 700 → 175 × 175 → 44 × 44 → 11 × 11. The resulting feature map is flattened and projected to a 1024-dimensional HRCT representation, followed by ReLU activation. HRCT selector. The HRCT selector operates at the patch level. It applies a lightweight convolutional selector to produce logits on a 10 × 10 patch grid for each montage channel, corresponding to 70 × 70 image patches in the original 700 × 700 resolution. The selector logits are converted into a differentiable top-k mask, which is upsampled by nearest-neighbor interpolation before being multiplied by the HRCT input. We set the HRCT selection budget to k = 60 patches. Classifier head. The 1024-dimensional clinical representation and 1024-dimensional HRCT representation are concatenated into a 2048-dimensional multimodal vector. The classifier maps this vector through 2048 → 256 → 2 with ReLU activation before the output layer. Training details. We train five models from scratch, one for each validation fold, for 300 epochs using Adam with a learning rate of 10−3 , weight decay of 10−4 , and a batch size of 8. The model with the highest validation AUC is selected for final testing. The mutual-information regularization weights are set to 0.01 for the clinical branch and 0.005 for the HRCT branch. C.0.4

Glaucoma Grading

Fundus encoder. The fundus encoder uses an EfficientNet-B3 branch to map the selected color fundus image to a 1000-dimensional representation. This representation is used as the modality embedding for multimodal fusion. Fundus selector. The fundus selector operates on the resized 256 × 256 RGB image. It contains three convolutional blocks with channel progression 3 → 64 → 128 → 256. Each block uses a 3 × 3 convolution, batch normalization, ReLU activation, and 2 × 2 max pooling, reducing the spatial

27

grid to 32 × 32. A final 1 × 1 convolution produces patch-level logits. SMILE selects k = 100 fundus patches from this grid, upsamples the mask to the input resolution, and multiplies it with the fundus image before encoding. This corresponds to 8 × 8 patches in the original fundus resolution. OCT encoder. The OCT encoder uses a ResNet-18 branch adapted to ten-channel OCT input. Specifically, the first convolution is modified to accept 10 input channels, and the final fully connected classification layer is removed. The resulting OCT representation is 512-dimensional. OCT selector. The OCT selector contains four convolutional blocks with channel progression 10 → 64 → 128 → 256 → 512. Each block uses a 3 × 3 convolution, batch normalization, ReLU activation, and 2 × 2 max pooling. The selector therefore reduces the 512 × 512 OCT tensor to a 32 × 32 patch-logit grid. A final 1 × 1 convolution outputs OCT selection logits, and SMILE keeps k = 100 OCT patches. The selected mask is upsampled to the OCT input resolution and applied before the OCT encoder. Each selected location corresponds to a 16 × 16 OCT patch. Classifier head. The 1000-dimensional fundus representation and 512-dimensional OCT representation are concatenated into a 1512-dimensional multimodal vector. The classifier maps this vector through 1512 → 756 → 3 with ReLU activation before the final three-class output. Training details. We train with Adam using a learning rate of 10−3 , weight decay of 10−4 , batch size 8, and a maximum of 1000 epochs. A step scheduler multiplies the learning rate by 0.2 every 500 epochs. The mutual-information regularization weight is set to 0.001 for each image modality. C.0.5

REST-meta-MDD

Functional connectivity graphs. For rs-fMRI, we use ROI-averaged BOLD time series parcellated by AAL-116. The preprocessing follows the consortium pipeline: the first 10 volumes are discarded; slice-timing correction, head-motion realignment, MNI normalization, temporal band-pass filtering from 0.01 to 0.10 Hz, and nuisance regression are applied. The nuisance regressors include head motion, global brain signal, white matter, cerebrospinal fluid, and linear/quadratic drift terms. For each subject, we compute a 116 × 116 Pearson functional-connectivity matrix and apply Fisher’s z transform. Nodes correspond to AAL ROIs, and each node feature is its full Fisher-z correlation profile. To obtain a shared topology, subject-level connectivity matrices are averaged at the dataset level, and the top 20% positive correlations are retained as admissible edges. This gives a fixed binary adjacency shared by all subjects. Graph encoder. The rs-fMRI encoder is a 3-layer GCN with ReLU activations and dropout p = 0.15. It maps node features through 116 → 116 → 96 → 64. The resulting 64-dimensional node embeddings are ℓ2 -normalized and aggregated by second-order pooling. Specifically, a 3-layer MLP maps node embeddings through 64 → 32 → 32 → 32, and the bilinear outer product is flattened into a 1024-dimensional graph embedding. Subgraph selector. The subgraph selector first linearly projects ROI correlation profiles to 64-dimensional node embeddings. It then constructs pairwise edge embeddings by concatenating node pairs, maps each 128-dimensional edge embedding through an MLP 128 → 64 → 1, and reshapes the outputs into an edge-logit matrix. Non-admissible edges are masked out using the shared adjacency. A soft top-k operation converts the remaining logits into a sparse differentiable edge mask. The selection temperature is annealed from 2.0 to 0.05 during warm-up to avoid premature hard selection while preserving gradient flow. In our multimodal experiment, the selector retains k = 25 rs-fMRI edges, approximately 1% of admissible pairs. sMRI preprocessing. For sMRI, we use gray-matter-volume (GMV) maps generated by a voxel-based morphometry pipeline. The pipeline includes bias-field correction, GM/WM/CSF segmentation, MNI normalization, and cross-site intensity normalization. The input image is the 28

(a)

(b)

Figure 9: iCTCF visualizations on 2D HRCT montages. Panels (a) and (b) show representative severe and mild patients, respectively. The highlighted regions indicate the HRCT patches selected by SMILE as prognostic evidence. mwc1 GMV map with spatial size 121 × 145 × 121, and each subject’s GMV map is z-normalized before model input. sMRI encoder. The sMRI encoder is a five-layer 3D CNN adapted to volumetric inputs. The first four convolutional layers use kernel size 3, stride 1, padding 1, batch normalization, ReLU activation, and 2 × 2 × 2 max pooling, reducing the spatial size as 121 × 145 × 121 → 60 × 72 × 60 → 30 × 36 × 30 → 15 × 18 × 15 → 7 × 9 × 7. The fifth convolution maps the feature map to 64 channels without further pooling, preserving the 7 × 9 × 7 grid. The channel progression is 28 → 58 → 128 → 256 → 64. The final feature map is globally averaged and passed through a projection head with dropout p = 0.5 and two GELU-activated fully connected layers 64 → 64 → 64, producing a 64-dimensional subject embedding. Volume selector. The sMRI selector is a lightweight 3D convolutional network that reduces the volume to a 7 × 9 × 7 logit grid, corresponding to 441 candidate subvolumes. A final 1 × 1 × 1 convolution outputs one logit per subvolume. Soft top-k selection retains k = 25 subvolumes, approximately 5% of the grid. During the first 200 warm-up epochs, the selection budget is annealed from 441 to 25 and the temperature is annealed from 1.0 to 0.05. Logistic noise is added to the subvolume logits for robustness. The resulting soft mask is upsampled to the input resolution and multiplied with the GMV volume before encoding. Classifier head. The 1024-dimensional rs-fMRI graph embedding and 64-dimensional sMRI subject embedding are concatenated into a 1088-dimensional multimodal representation and passed to a classifier head for MDD diagnosis. Training details. The maximum training budget is 2000 epochs. The encoders are warmed up for 200 epochs before joint optimization, and early stopping is applied with a patience of 100 epochs according to validation performance. The batch size is 32. Experiments are repeated over 10 runs. The first random seed is 42, and each subsequent run increases the seed by 1024.

29

(a)

(b)

Figure 10: Glaucoma Grading visualizations. Panels (a) and (b) show additional representative cases under the same multimodal setting. SMILE highlights selected regions on color fundus images and OCT slices, providing modality-specific evidence for glaucoma-grade prediction.

D

Additional Visualization Results

iCTCF visualizations. For iCTCF, we visualize both modalities (see Fig. 9). The HRCT explanation highlights the selected patches on the 2D montage, enabling spatial inspection of imaging evidence. The clinical-feature explanation reports the normalized importance of the selected clinical variables over the test set. We separately show representative severe and mild cases so that the reader can assess whether the model attends to different image regions across prognostic outcomes. Glaucoma Grading visualizations. For Glaucoma Grading, we visualize the color fundus and OCT modalities separately. The selected regions are displayed as heatmaps, enabling examination of whether the model uses anatomically meaningful evidence in each image stream (see Fig. 10).

30

Record · ID 660816 · SHA-256 778abbf1fc5c984e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.