Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation
arXiv:2606.30374v1 [cs.CV] 29 Jun 2026
Seunghun Baek∗ , Jihwan Park∗ , Jaeyoon Sim, Hoseok Lee, Seungjoo Lee, and Won Hwa Kim Pohang University of Science and Technology, Pohang, South Korea {habaek4, pjh58110, simjy98, hslee0608, seungjoo612, wonhwa}@postech.ac.kr
Abstract. Multimodal MRI is essential for accurate brain tumor segmentation. However, acquiring all modalities at inference is often challenging in practice, which causes intrinsic uncertainty due to unavoidable information loss. Without modeling this uncertainty, existing methods encode incomplete evidence into deterministic representations that appear plausible but lack reliability. In this regime, we propose a probabilistic representation framework that models representations as Gaussian distributions, where their mean captures task information and their variance measures uncertainty from missing evidence. To make variance reflect information deficiency, we regularize the mean from each partial configuration toward its full-modality counterpart, while scaling the variance with the discrepancy between their aligned means. We further introduce a set-inclusive strategy that exploits the hierarchical structure of modality subsets and enforces an ordering constraint to maintain their consistent uncertainty relationships. Extensive experiments on BraTS 2018 and 2020 demonstrate that our approach offers superior performance over baselines across diverse missing-modality scenarios. Code and model checkpoint are available at https://github.com/atlas-sky/SIUM. Keywords: Brain Tumor Segmentation · Probabilistic Representation
1
Introduction
Brain tumor segmentation is critical for assessing disease progression and treatment planning [3,5,17]. In clinical practice, four standard magnetic resonance imaging (MRI) modalities are utilized: T1, T1c, T2, and FLAIR [10,11], which provide complementary information for delineating tumor subregions, including whole tumor (WT), tumor core (TC), and enhancing tumor (ET) [4,6,11,17]. Since no single modality fully captures every tumor subregion, accurate delineation requires the joint integration of all modalities [12,18]. However, acquiring all of them is often impractical due to cost and patient burden [1,2]. As brain tumor segmentation relies on the complementary evidence from multiple modalities, when a modality goes missing, its unique information becomes inaccessible. *
S. Baek and J. Park contributed equally to this paper.
2
S. Baek & J. Park et al.
To address potential incompleteness, previous works [8,14,15,21] train their frameworks to be robust across missing-modality scenarios. They encode each modality configuration into a deterministic embedding to represent task-relevant information. While the absence introduces unavoidable information loss in the resulting embedding, deterministic approaches still update their prediction models without considering informational incompleteness (i.e., uncertainty). Consequently, the models may rely on hallucinated features that lack evidence. In this regime, we model each configuration as a Gaussian distribution, where the mean encodes task-relevant information and the variance captures uncertainty. This design acknowledges that missing modalities cause intrinsic information deficiency, which should be reflected as uncertainty rather than concealed within a deterministic embedding. To make the variance encode uncertainty, we exploit the hierarchical structure of modality subsets, where smaller sets are nested within larger ones. Within this set-inclusive hierarchy, the variance is encouraged to increase when the mean of a subset deviates from its full-modality counterpart, since deviation indicates missing evidence. It is further regularized to remain higher for subsets than for their supersets, which reflects their reduced informational completeness. Through uncertainty guidance, the model enables reliable predictions that prioritize well-supported evidence across configurations. Our main contributions are summarized as follows: 1) We propose a probabilistic framework that integrates set-inclusive hierarchy with uncertainty guidance to model intrinsic uncertainty and emphasize reliable cues. 2) We provide a theoretical analysis of how uncertainty influences task-representation optimization and empirical validation of the learned uncertainty to support the need for probabilistic embeddings. 3) Extensive experiments on BraTS 2018 and 2020 demonstrate superior performance under diverse missing-modality scenarios. Related works on Brain Tumor Segmentation. Early approaches adopt imputation methods that synthesize absent modalities [7,9,19]. However, these methods require additional generative models, which are computationally heavy and time-consuming. Recent works therefore have shifted toward imputation-free strategies. In particular, RFNet [8] introduces a region-aware fusion module that assigns adaptive modality importance across tumor regions. mmFormer [21] improves it using inter- and intra-modal Transformer [20] to model local and global context. M3 AE [15] mitigates missing-modality effects by modality- and patchlevel masking and self-distillation. DC-Seg [14] disentangles modality-invariant and -specific representations via contrastive learning for anatomical consistency.
2
Methods
General problem setting. Given a sample xi (i.e., subject) with M modali(m) ties, we denote its m-th modality input (e.g., T1 scan) as xi for m =1, . . . , M . (m) Each modality input xi is encoded by a modality-specific encoder E (m) to (m) (m) produce an embedding ei = E (m) (xi ). To simulate missing-modality scenar(m) ios during training, a Bernoulli indicator δi ∈ {0, 1} is adopted to mask the
Set-Inclusive Uncertainty Modeling
3
(a) Main Pipeline Set-Inclusive Modality Masking 𝒇𝒖𝒍𝒍 𝒔𝒖𝒑 𝓢𝒔𝒖𝒃 ⊂ 𝓢𝒊 ⊂ 𝓢𝒊 𝒊
(1)
(𝑚)
𝑥𝑖
ℎ𝑖𝑠𝑢𝑏
𝒇𝜽
Uncertainty Guidance 𝝈 Head
𝜇𝑖𝑠𝑢𝑏
𝒇𝝓
𝜎𝑖𝑠𝑢𝑏
(𝑀) 𝑥𝑖 𝑠𝑢𝑝
ℎ𝑖
(𝑀)
𝑥𝑖
𝑓𝑢𝑙𝑙
𝑥𝑖
ℎ𝑖
𝒩(𝜇𝑖𝑠𝑢𝑏 , (𝜎𝑖𝑠𝑢𝑏 )2 )
𝑠𝑢𝑝 𝑠𝑢𝑝
𝜎𝑖
𝑓𝑢𝑙𝑙
Goal: ∀𝑘 ∈ 𝑠𝑢𝑏, 𝑠𝑢𝑝 , 𝑓𝑢𝑙𝑙
𝜎𝑖𝑘 → 𝜇𝑖
− 𝜇𝑖𝑘
2
𝑓𝑢𝑙𝑙
& 𝜇𝑖𝑘 → 𝜇𝑖
𝑓𝑢𝑙𝑙
𝜇𝑖
𝜇𝑖𝑘
𝒉
𝑦ෝ𝑖 𝑠𝑢𝑏
𝑠𝑢𝑝
𝒉
𝑦ෝ𝑖 𝑠𝑢𝑝
𝒉
𝑦ෝ𝑖 𝑓𝑢𝑙𝑙
𝑟𝑖
𝜇𝑖
𝒇𝜽
𝜎𝑖𝑘
𝑟𝑖𝑠𝑢𝑏
𝑠𝑢𝑝 𝑠𝑢𝑝 𝒩(𝜇𝑖 , (𝜎𝑖 )2 )
(b) Uncertainty Guidance 1. Uncertainty-aware Alignment 𝓛𝑼𝑨
Task Head
𝜇𝑖
𝒇𝜽 𝒇𝝓
(𝑀)
⋮
⋮
𝑥𝑖
(𝑚)
𝑥𝑖
⋮
⋮
𝒇𝒖𝒍𝒍
𝓢𝒊
(1)
𝑥𝑖
(𝑚) 𝑥𝑖
⋮
𝒔𝒖𝒑
𝓢𝒊
(1) 𝑥𝑖
⋮
𝓢𝒔𝒖𝒃 𝒊
𝝁 Head
Encoder & Decoder
Notations 2. Uncertainty Ordering 𝓛𝑼𝑶 𝑠𝑢𝑝 2
Goal: (𝜎𝑖𝑠𝑢𝑏 )2 ≥ (𝜎𝑖
)
(𝒎) 𝒙𝒊 : 𝑖𝑡ℎ sample’s 𝑚𝑡ℎ modality 𝒇𝜽 , 𝒇𝝓 , 𝒉 : learnable network
: weight shared
Fig. 1. Overall framework. (a) For each sample xi , three modality sets Sisub ⊂ Sisup ⊂ Sifull are constructed and then processed independently. Subset configurations are modeled probabilistically as rik ∼N (µki , (σik )2 ), whereas the full configuration serves as a dek terministic anchor as µfull i . (b) LUA aligns each subset mean µi with the full-modality anchor while enabling σik to reflect their discrepancy. LUO further enforces an ordering constraint such that smaller subsets maintain higher uncertainty than their supersets. (m) (m)
modality embedding, resulting in δi ei . A shared decoder D then aggregates (m) (m) the set of masked embeddings {δi ei }M m=1 to obtain a task-specific represen(m) (m) M tation zi = D({δi ei }m=1 ). While incomplete modalities inherently imply a high degree of uncertainty, existing methods [8,14,15,21] typically maintain zi as a deterministic embedding. By directly passing zi to a task head h to obtain the prediction ŷi = h(zi ), such approaches do not explicitly account for the uncertainty induced by unobserved modalities. As a consequence, the model may be trained with excessive confidence in the presence of missing modalities. 2.1
Theoretical Analysis of Probabilistic Representation
To explicitly model the uncertainty, we represent the task embedding zi in a probabilistic manner. The task representation zi is transformed into a Gaussian embedding composed of a mean µi = fθ (zi ) and a standard deviation σi = fϕ (zi ), where fθ and fϕ are learnable networks parametrized by θ and ϕ. The role of µi is to encode the task-specific embedding, while σi models the uncertainty inherent in µi . We sample ri from N (µi , σi2 ) to obtain a probabilistic representation that captures both the embedding and its uncertainty. To enable gradient optimization, we employ the reparameterization trick [13] on ri as ri = µi + ϵ ⊙ σi ,
ϵ ∼ N (0, I).
(1)
During training, task prediction is obtained as ŷi = h(ri ) to incorporate uncertainty. At inference, only µi is used as ŷi = h(µi ) for the consistent prediction. Our training introduces stochastic perturbations around µi , whose magnitude is governed by σi , thereby influencing the loss landscape. To clarify this effect, we
4
S. Baek & J. Park et al.
analyze how σi affects the optimization of θ, where µi =fθ (zi ), by examining the gradient of the loss L under sampling noise ϵ∼N (0, I) (i.e., ∇θ L(θ; ϵ)). Applying i the chain rule, gradient with respect to θ decomposes into ∇ri L(θ; ϵ) and ∂r ∂θ as ∇θ L(θ; ϵ) = (
∂µi ⊤ ∂ri ∂ri ∂µi ⊤ ) ∇ri L(θ; ϵ) = ∇ri L(θ; ϵ) (∵ = I). ∂µi ∂θ ∂θ ∂µi
(2)
Since the noise ϵ⊙σi is added to µi in ri , the gradient ∇ri L(θ; ϵ) can be linearized around the reference point µi (i.e., ϵ = 0). A first-order Taylor expansion yields ∇ri L(θ; ϵ) ≈ ∇ri L(θ; 0) + ∇2ri L(θ; 0)(ϵ ⊙ σi ).
(3)
To characterize the influence of ϵ on ∇θ L(θ; ϵ), we first analyze the covariance of ∇ri L(θ; ϵ) with respect to the sampling noise ϵ (i.e., Covϵ [∇ri L(θ; ϵ)]). Using the approximation in Eq. (3) and the covariance identity Cov[Ax]=ACov[x]A⊤ , the resulting gradient covariance with respect to ϵ is proportional to σi2 as Covϵ [∇ri L(θ; ϵ)] ≈ Covϵ [∇ri L(θ; 0)] + Covϵ [∇2ri L(θ; 0)(ϵ ⊙ σi )] = ∇2ri L(θ; 0) diag σi2 ∇2ri L(θ; 0)⊤ (∵ Covϵ [ϵ ⊙ σi ] = diag σi2 ). (4) The chain rule (Eq. (2)) further links Covϵ [∇θ L(θ; ϵ)] to Covϵ [∇ri L(θ; ϵ)] as Covϵ [∇θ L(θ; ϵ)] =
∂µi ∂µi ⊤ Covϵ [∇ri L(θ; ϵ)] , ∂θ ∂θ
(5)
which also scales with σi2 according to Eq. (4). Larger σi induces higher gradient variance in the µi -determining parameters θ. The resulting stochastic updates fluctuate across iterations and partially cancel in expectation, which attenuate effective parameter updates and thereby diminish their impact on the task. 2.2
Set-Inclusive Modality Masking and Uncertainty Guidance
While σi modulates the updates of the model parameter θ, optimizing the task loss alone (i.e, L = LTask ) collapses σi toward zero. To be explicit, analogous to the local expansion in Eq. (3), the expected loss admits the approximation as Eϵ [L(µi + ϵ ⊙ σi )] ≈ L(µi ) + 21 Tr ∇2ri L(θ; 0) diag(σi2 ) . (6) Since the second term is non-negative and strictly increasing in σi2 , minimizing the task loss drives σi → 0, which necessitates additional regularization to preserve stochasticity in ri . In this regime, we posit that σi should increase when task-relevant information is insufficient, which naturally occurs as available modalities are progressively removed. This motivates our subsequent design, which incorporates set-inclusive training and uncertainty regularization. Set-Inclusive Modality Masking. Let M = {1, . . . , M } denote the index set of all modalities. For a given sample xi , each modality configuration is represented as a subset Si ⊆ M indicating the observed modalities. The masking
Set-Inclusive Uncertainty Modeling
5
(m)
indicator δi is set to 1 if and only if m ∈ Si , and to 0 otherwise. To enable setinclusive uncertainty modeling, we consider three modality index sets associated with the same sample xi . Let Sifull = M denote the full-modality configuration. We further define two partial modality sets Sik ∈ {Sisub , Sisup }, which satisfy the inclusion relation as Sisub ⊂ Sisup ⊂ Sifull . For each partial modality set Sik , the task representation zik is parameterized as a Gaussian embedding with mean µki and variance (σik )2 . We then perform stochastic sampling (Eq. (1)) as rik = µki + ϵ ⊙ σik to obtain the task prediction ŷik = h(rik ). For a full-modality set, which serves as an uncertainty-free anchor, the prediction ŷifull is obtained directly from µfull i . Accordingly, the task loss is defined over the modality sets as LTask = ℓ(ŷisub , yi ) + ℓ(ŷisup , yi ) + ℓ(ŷifull , yi ),
(7)
where ℓ is a task-specific loss (e.g., Dice loss) and yi is the ground-truth label. Set-Inclusive Uncertainty Guidance. As the full-modality configuration Sifull provides the most complete evidence, we treat its mean µfull as an i uncertainty-free anchor. For each incomplete configuration k ∈ {sub, sup}, we align its mean µki with this anchor, while allowing σik to quantify the discrepancy between µki and µfull i . We define the corresponding uncertainty-aware alignment loss LUA as a Gaussian negative log-likelihood with gradients detached from µfull i . For µki , σik ∈ RC×H×W×D with channel dimension C and spatial shape H×W ×D, LUA is computed element-wise and averaged over all elements via E as hP i k full 2 1 (µi −GradStop(µi )) 1 k 2 LUA = E + log(σ ) . (8) i k∈{sub,sup} 2 2 (σ k )2 i
k The former term penalizes deviations of µki from µfull i , while enlarging σi when k the discrepancy is substantial. The latter prevents σi from growing unbounded. These terms drive µki toward µfull while guiding σik to reflect their discrepancy. i
We further impose an ordering constraint consistent with the set-inclusion hierarchy. Since Sisub ⊂ Sisup , uncertainty from Sisub should be greater. To penalize violations of this hierarchy, we introduce an uncertainty ordering loss LUO as LUO = E max 0, (σisup )2 − (σisub )2 . (9) The overall objective L combines the task loss with our two uncertainty regularization terms, weighted by hyperparameters α and β, respectively, as L = LTask + αLUA + βLUO .
3
Experimental Settings
3.1
Datasets and Evaluation Metrics
(10)
We evaluate our framework on two brain tumor segmentation benchmarks, BraTS 2018 and 2020 [17], to assess performance across cohorts of different scales and compositions. Both datasets provide four MRI modalities (i.e., T1, T1c, T2,
6
S. Baek & J. Park et al.
Table 1. Test performance measured by the Dice similarity coefficient (DSC, %) on the BraTS 2020 dataset. For each modality, • and ◦ denote its presence and absence, respectively. The best and second-best results are highlighted in bold and underline. Modalities
Whole tumor (WT)
Tumor core (TC)
Enhancing tumor (ET)
F T1 T1c T2 RFNet mmFormer M3 AE DC-Seg Ours RFNet mmFormer M3 AE DC-Seg Ours RFNet mmFormer M3 AE DC-Seg Ours ◦ ◦ ◦ • ◦ ◦ • ◦ • • • • • ◦ •
◦ ◦ • ◦ ◦ • • • ◦ ◦ • • ◦ • •
◦ • ◦ ◦ • • ◦ ◦ ◦ • • ◦ • • •
• ◦ ◦ ◦ • ◦ ◦ • • ◦ ◦ • • • •
Average
86.05 76.77 77.16 87.32 87.74 81.12 89.73 87.73 89.87 89.89 90.69 90.60 90.68 88.25 91.11
85.51 78.04 76.24 86.54 87.52 80.70 88.76 86.94 89.49 89.31 89.79 89.83 90.49 87.64 90.54
86.10 78.90 79.00 88.00 87.10 80.10 89.60 87.30 90.10 89.50 89.60 90.20 90.50 87.40 90.40
86.72 79.54 78.47 87.80 88.17 82.22 90.01 88.09 90.32 89.99 90.65 90.77 90.62 88.73 90.95
71.02 81.51 66.02 69.19 83.45 83.40 73.07 73.13 74.14 84.65 85.07 75.19 84.97 83.47 85.21
63.36 81.51 63.23 64.60 82.69 82.81 71.76 67.76 70.34 83.79 84.44 72.42 83.94 83.66 84.61
71.80 83.60 69.40 68.70 85.60 83.80 72.80 72.90 74.30 85.50 85.60 74.40 85.80 85.80 86.20
70.88 73.23 84.62 86.15 66.63 64.88 71.27 72.07 86.34 86.89 85.18 87.02 74.50 74.71 73.09 73.53 75.11 76.34 85.90 86.44 86.29 86.40 75.53 75.98 86.21 86.39 86.49 86.96 86.46 86.38
46.29 74.85 37.30 38.15 75.93 78.01 40.98 45.65 49.32 76.67 76.81 49.92 77.12 76.99 78.00
49.09 78.30 37.62 36.68 77.20 81.71 42.98 49.12 49.06 79.44 80.65 50.08 78.73 77.34 79.92
47.10 73.60 40.40 40.20 76.00 75.30 43.70 48.70 47.10 75.90 76.30 48.20 77.40 78.00 77.50
47.76 49.55 78.90 81.93 42.19 41.63 41.66 40.56 80.43 83.31 79.25 83.11 46.90 46.81 50.19 48.86 51.32 50.82 80.28 82.25 81.41 83.20 52.05 50.80 79.42 81.54 81.66 83.55 81.52 83.35
86.98
86.49
86.90
87.54 88.07 78.23
76.06
79.10
79.63 80.22 61.47
63.19
61.70
65.00 66.02
87.24 80.11 80.19 87.85 88.43 83.41 90.20 88.67 90.42 90.44 91.09 91.00 91.27 89.14 91.58
Table 2. Average Dice similarity coeffi- Table 3. Ablation study on BraTS 2020. cient (DSC, %) on BraTS 2018 dataset. (Set &Prob: set-inclusive probabilistic design) Methods WT TC ET RFNet 85.49 76.00 58.44 ∗ mmFormer 86.20 73.65 56.20 85.82 77.37 59.85 M3 AE DC-Seg∗ 86.59 77.14 59.55 Ours 87.10 78.37 62.52
Set&Prob LUA LUO × × × ✓ × × ✓ ✓ × ✓ × ✓ ✓ ✓ ✓
WT TC ET 86.98 78.23 61.47 87.54 79.20 63.74 87.73 79.32 65.57 87.57 79.31 64.61 88.07 80.22 66.02
and FLAIR) and standardized tumor annotations, and performance is evaluated on WT, TC, and ET. All volumes are skull-stripped, aligned to a common anatomical template, and resampled to an isotropic resolution of 1mm3 , followed by zero-mean, unit-variance normalization within the brain region. BraTS 2018 contains 285 subjects, whereas BraTS 2020 extends the cohort to 369 subjects. We follow the same train, validation, and test splits as M3 AE [15] for BraTS 2018 (199:29:57) and RFNet [8] for BraTS 2020 (219:50:100). During training, volumetric patches of 112×112×112 are randomly extracted. At inference, we adopt the sliding-window strategy of RFNet [8] and average overlapping predictions. Performance is evaluated using the Dice similarity coefficient (DSC) [16]. 3.2
Baselines and Implementation Details
To validate our work, we compare our method with recent state-of-the-art approaches [8,14,15,21]. We report the results on each dataset from the original papers when available. Otherwise, we reproduce them using the official code and hyperparameters. These reproduced results are marked with * in Tab. 2. Following the baselines [14,15,21], we adopt the backbone architecture and training protocol of RFNet [8]. The output of the penultimate layer (i.e., the feature representation before the segmentation head) is replaced with a probabilistic embedding. Two separate 3D CNN layers are used for fθ and fϕ , respectively.
Set-Inclusive Uncertainty Modeling T1
T1c
T2
Ground truth
T2
T1c+T2
Flair+T1c
T1+T1c+T2
F+T1+T1c+T2
T2
T1c+T2
Flair+T1c
T1+T1c+T2
F+T1+T1c+T2
T2
T1c+T2
Flair+T1c
T1+T1c+T2
F+T1+T1c+T2
Ours
DC–Seg
RFNet
Modalities
Flair
7
Fig. 2. Qualitative results on BraTS 2020. Our model produces improved segmentations over RFNet [8] and DC-Seg [14] across various configurations. (Region: ET (blue) ⊂ TC (blue, red) ⊂ WT (blue, red, green), Sample: “HG_BraTS20_Training_357”) Flair
T1c
Ground truth
T1
T1+Flair +T1
T1+Flair+T1c
Modalities
Uncertainty
T1
Fig. 3. Visualization of input modalities and ground-truth mask, with trained uncertainty overlaid on T1. Uncertainty, initially higher within tumor regions, decreases as additional modalities are incorporated. (Sample: “HG_BraTS20_Training_045”)
Our framework was trained for 800 epochs using the Adam optimizer with a learning rate of 2×10−4 and a batch size of 2, where α=0.01 and β=5.
4
Results and Analysis
4.1
Performance on the BraTS 2018 and 2020 datasets
Quantitative Results. We first compare our method with recent baselines on BraTS 2020 dataset in Tab. 1. Our method outperformed the baselines by a clear margin in terms of the averaged Dice similarity score across the three tumor subregions. In particular, it consistently achieved the best or second-best performance under almost every modality configuration, which demonstrates its robustness against diverse missing-modality scenarios. A similar trend was observed on BraTS 2018 dataset, as shown in Tab. 2. These consistent improvements across the datasets indicate our superiority across underlying data cohorts. Qualitative Results. We visualize center-slice predictions on BraTS 2020 under modality configurations, together with RFNet [8] and DC-Seg [14] (Fig. 2). Our model yields more accurate and consistent segmentations across various configurations, notably preserving the background regions adjacent to the whole tumor. In contrast, the baselines overfill these regions (in blue) and overestimate the ET. These results indicate that our model relies on reliable cues, whereas deterministic embeddings in the baselines yield spurious hallucinations.
S. Baek & J. Park et al.
0.098 0.093
2.5 2.0
DSC
|| 2||
8
Gradient w.r.t ri
ri L( ; )
2 i:
0.02
0.04
ri L( ; 0)
0.06
Var [ riL( ; )] 0.08
0.10
0.12
0.14
0.16
0.18
0.20
Gradient Variance
Fig. 4. Variance magnitude ( ) and summed test DSC across tumor subregions ( ) for each modality configuration on BraTS 2020. • and ◦ indicate observed and missing modalities (Flair, T1, T1c, T2). Higher uncertainty correlates with lower performance. The shaded regions highlight that a superset exhibits lower uncertainty than its subset.
Fig. 5. Effect of variance magnitude on gradient behavior under controlled scaling.
4.2
Model Behavior Analyses
Ablation Study. We analyze the contribution of each component over the deterministic baseline (i.e., RFNet [8]) on BraTS 2020 (Tab. 3). While adopting probabilistic embeddings with set-inclusive modality masking improves segmentation, the variance collapse restricts its benefit (see Eq.(6)). LUA mitigates this collapse and further enhances performance, whereas LUO alone brings marginal gains as it constrains ordering rather than magnitude. Combining LUA and LUO yields the best results by providing complementary uncertainty supervision. Uncertainty Analysis. We first examine the spatial distribution of learned uncertainty by averaging σ across channels and overlaying it on the T1 image (Fig. 3). Elevated uncertainty is primarily localized within tumor regions, particularly the enhancing tumor (blue label). By incorporating Flair and T1c together, uncertainty in these regions progressively decreases, indicating that complementary evidence reduces ambiguity and stabilizes model updates. We further investigate the relationship between uncertainty and performance on BraTS 2020 across modality configurations (Fig. 4). Uncertainty is quantified by the mean squared magnitude of σ during training, denoted as ||σ 2 ||. Performance is measured by the summed Dice similarity coefficient over the three P tumor subregions, denoted as DSC. A clear inverse relationship is observed that configurations with larger variance magnitudes consistently correspond to lower test performance, and vice versa. This trend suggests that when missing modalities induce greater task-relevant information loss, the model responds by increasing uncertainty rather than committing to incomplete evidence. Moreover, as highlighted by the shaded samples, superset configurations consistently exhibit lower uncertainty and higher performance than their subsets. This ordering behavior aligns with our set-inclusive design, confirming that uncertainty systematically reflects modality-induced information incompleteness. Gradient Scaling Analysis. To empirically validate Eq. (4), which predicts that uncertainty σi modulates gradient covariance, we analyze the ∇ri L(θ; ϵ) at an arbitrary element of ri with fixed µi while varying σi (Fig. 5). For each σi , we
Set-Inclusive Uncertainty Modeling
9
sample ϵ 30 times to compute ∇ri L(θ; ϵ). Compared to deterministic gradient (ϵ = 0; dashed line), the sampled ones become increasingly dispersed as σi2 grows (i.e., enlarged variance). This confirms the theoretical relationship in Eq. (4).
5
Conclusion
In this work, we address the intrinsic uncertainty induced by missing modalities through a probabilistic framework for incomplete multimodal brain tumor segmentation. This design explicitly reduces reliance on spurious cues, thereby yielding more reliable predictions. Extensive experiments on BraTS 2018 and 2020 demonstrate robustness across missing-modality scenarios. Both theoretical and empirical analyses support the necessity of our uncertainty modeling. Acknowledgments. This research was supported by RS-2025-02216257 (60%), RS2026-25494850 (30%), RS-2019-II1091906 (AI Graduate Program at POSTECH, 5%), and RS-2022-II220290 (5%). Disclosure of Interests. The authors have no competing interests to declare that are relevant to the content of this article.
References 1. Baek, S., Sim, J., Dere, M., et al.: Modality-agnostic style transfer for holistic feature imputation. In: ISBI. pp. 1–5. IEEE (2024) 2. Baek, S., Sim, J., Wu, G., et al.: Ocl: Ordinal contrastive learning for imputating features with progressive labels. In: MICCAI. pp. 334–344. Springer (2024) 3. Baid, U., Ghodasara, S., Mohan, S., et al.: The rsna-asnr-miccai brats 2021 benchmark on brain tumor segmentation and radiogenomic classification. arXiv preprint arXiv:2107.02314 (2021) 4. Bakas, S., Sako, C., Akbari, H., et al.: The university of pennsylvania glioblastoma (upenn-gbm) cohort: advanced mri, clinical, genomics, & radiomics. Scientific data 9(1), 453 (2022) 5. Bakas, S., Shukla, G., Akbari, H., et al.: Overall survival prediction in glioblastoma patients using structural magnetic resonance imaging (mri): advanced radiomic features may compensate for lack of advanced mri modalities. JMI 7(3), 031505– 031505 (2020) 6. Blystad, I., Warntjes, J.M., Smedby, Ö., et al.: Quantitative mri for analysis of peritumoral edema in malignant gliomas. PLoS One 12(5), e0177135 (2017) 7. Conte, G.M., Weston, A.D., Vogelsang, D.C., et al.: Generative adversarial networks to synthesize missing t1 and flair mri sequences for use in a multisequence brain tumor segmentation model. Radiology 299(2), 313–323 (2021) 8. Ding, Y., Yu, X., Yang, Y.: Rfnet: Region-aware fusion network for incomplete multi-modal brain tumor segmentation. In: ICCV. pp. 3975–3984 (2021) 9. Dorent, R., Joutard, S., Modat, M., et al.: Hetero-modal variational encoderdecoder for joint modality completion and segmentation. In: MICCAI. pp. 74–82. Springer (2019)
10
S. Baek & J. Park et al.
10. Havaei, M., Davy, A., Warde-Farley, D., Biard, A., Courville, A., Bengio, Y., Pal, C., Jodoin, P.M., Larochelle, H.: Brain tumor segmentation with deep neural networks. Medical image analysis 35, 18–31 (2017) 11. Işın, A., Direkoğlu, C., Şah, M.: Review of mri-based brain tumor image segmentation using deep learning methods. Procedia computer science 102, 317–324 (2016) 12. Kamnitsas, K., Ledig, C., Newcombe, V.F., et al.: Efficient multi-scale 3d cnn with fully connected crf for accurate brain lesion segmentation. Medical image analysis 36, 61–78 (2017) 13. Kingma, D.P., Welling, M.: Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013) 14. Li, H., Li, Z., Mao, Y., et al.: Dc-seg: Disentangled contrastive learning for brain tumor segmentation with missing modalities. In: MICCAI. pp. 138–148. Springer (2025) 15. Liu, H., Wei, D., Lu, D., et al.: M3ae: multimodal representation learning for brain tumor segmentation with missing modalities. In: AAAI. vol. 37, pp. 1657–1665 (2023) 16. Lr, D.: Measures of the amount of ecologic association between species. Ecology 26, 297–302 (1945) 17. Menze, B.H., Jakab, d., Bauer, S., et al.: The multimodal brain tumor image segmentation benchmark (brats). TMI 34(10), 1993–2024 (2014) 18. Pereira, S., Pinto, A., Alves, V., et al.: Brain tumor segmentation using convolutional neural networks in mri images. TMI 35(5), 1240–1251 (2016) 19. Sharma, A., Hamarneh, G.: Missing mri pulse sequence synthesis using multi-modal generative adversarial network. TMI 39(4), 1170–1183 (2019) 20. Vaswani, A., Shazeer, N., Parmar, N., et al.: Attention is all you need. NeurIPS 30 (2017) 21. Zhang, Y., He, N., Yang, J., et al.: mmformer: Multimodal medical transformer for incomplete multimodal learning of brain tumor segmentation. In: MICCAI. pp. 107–117. Springer (2022)