Information-Theoretic Optimization for Task-Adapted Compressed Sensing Magnetic Resonance Imaging arXiv:2604.12709v1 [cs.LG] 14 Apr 2026
Xinyu Peng, Ziyang Zheng*, Member, IEEE, Wenrui Dai*, Member, IEEE, Duoduo Xue, Member, IEEE, Shaohui Li, Member, IEEE, Chenglin Li, Member, IEEE, Junni Zou, Member, IEEE, and Hongkai Xiong, Fellow, IEEE
✦
* Corresponding authors: Ziyang Zheng and Wenrui Dai. Xinyu Peng is with the Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai 200240, China. E-mail: [email protected]. Ziyang Zheng, Wenrui Dai, Chenglin Li, Junni Zou, and Hongkai Xiong are with the Department of Electronic Engineering, Shanghai Jiao Tong University, Shanghai 200240, China. E-mail: {zhengziyang, daiwenrui, lcl1985, zoujunni, xionghongkai}@sjtu.edu.cn. Duoduo Xue is with the Department of Computer Science, City University of Hong Kong, Hong Kong SAR, China. E-mail: [email protected]. Shaohui Li is with the College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310007, China. E-mail: [email protected]. This is a draft and the final version has been accepted by IEEE TPAMI (DOI: 10.1109/TPAMI.2026.3683201).
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
Abstract Task-adapted compressed sensing magnetic resonance imaging (CS-MRI) is emerging to address the specific demands of downstream clinical tasks with significantly fewer k-space measurements than required by Nyquist sampling. However, existing task-adapted CS-MRI methods suffer from the uncertainty problem for medical diagnosis and cannot achieve adaptive sampling in end-to-end optimization with reconstruction or clinical tasks. To address these limitations, we propose the first task-adapted CS-MRI from the information-theoretic perspective to simultaneously achieve probabilistic inference for uncertainty prediction and adapt to arbitrary sampling ratios and versatile clinical applications. Specifically, we formalize the task-adapted CS-MRI optimization problem by maximizing the mutual information between undersampled k-space measurements and clinical tasks to enable probabilistic inference for addressing the uncertainty problem. We leverage amortized optimization and construct tractable variational bounds for mutual information to jointly optimize sampling, reconstruction, and task-inference models, which enables flexible sampling ratio control using a single end-to-end trained model. Furthermore, the proposed framework addresses two kinds of distinct clinical scenarios within a unified approach, i.e., i) joint task and reconstruction, where reconstruction serves as an auxiliary process to enhance task performance; and ii) task implementation with suppressed reconstruction, applicable for privacy protection. Extensive experiments on large-scale MRI datasets demonstrate that the proposed framework achieves highly competitive performance on standard metrics like Dice compared to deterministic counterpart but provides better distribution matching to the ground-truth posterior distribution as measured by the generalized energy distance (GED).
Index Terms Magnetic resonance imaging, compressed sensing, variational inference, information theoretic metric.
1
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
2
C ONTENTS 1
Introduction
4
2
Related Works
6
2.1
Deep Learning for MR Sampling Patterns . . . . . . . . . . . . . . . .
6
2.2
Optimization of Mutual Information . . . . . . . . . . . . . . . . . . .
7
2.3
Task-adapted CS-MRI . . . . . . . . . . . . . . . . . . . . . . . . . . . .
7
2.4
Stochastic Medical Image Segmentation . . . . . . . . . . . . . . . . .
8
3
Proposed Method
9
3.1
Preliminary: Information Theoretic Metric . . . . . . . . . . . . . . . .
9
3.2
End-to-end Variational Information Optimization for Task-Adapted CS-MRI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 3.2.1
β < 0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
3.2.2
β ≥ 0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
3.3
Adaptive Sampling via Amortized Optimization . . . . . . . . . . . . 13
3.4
Tractable Optimization With Variational Bounds (β < 0) . . . . . . . . 16
3.5
Tractable Optimization With Marginal Entropy Minimization (β ≥ 0) 19
3.6
Relation to Existing Methods . . . . . . . . . . . . . . . . . . . . . . . . 24 3.6.1
Relation to LI-Net [15] . . . . . . . . . . . . . . . . . . . . . 24
3.6.2
Relation to Existing End-to-End CS-MRI Reconstruction Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
4
3.6.3
Relation to Existing Task-Adapted MRI Methods . . . . . . 28
Experiments
29
4.1
Comparisons to CS-MRI Reconstruction Methods . . . . . . . . . . . 29 4.1.1
Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . 29
4.1.2
Reconstruction Performance . . . . . . . . . . . . . . . . . . 32
4.1.3
Performance with Advanced Reconstruction Methods . . . 34
4.1.4
Effectiveness of Amortized Optimization . . . . . . . . . . 34
4.1.5
Reconstruction Uncertainty Quantification . . . . . . . . . . 37
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
4.2
4.3
4.4 5
3
Comparisons to Task-adapted CS-MRI Methods . . . . . . . . . . . . 38 4.2.1
Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . 38
4.2.2
Quantitative Results . . . . . . . . . . . . . . . . . . . . . . . 42
4.2.3
Segmentation Under Uncertainty . . . . . . . . . . . . . . . 42
Privacy-Preserving Task-Adapted CS-MRI . . . . . . . . . . . . . . . . 43 4.3.1
Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . 44
4.3.2
Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
Discussion and Limitation . . . . . . . . . . . . . . . . . . . . . . . . . 46 46
Conclusions
Appendix A: Additional Experimental Results
48
A.1
Additional Results on CS-MRI Reconstruction . . . . . . . . . . . . . 48
A.2
Additional Results on Task-adapted CS-MRI . . . . . . . . . . . . . . 48
A.3
A.4
A.2.1
More Results . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
A.2.2
Confidence Calibration . . . . . . . . . . . . . . . . . . . . . 49
A.2.3
Visualization of Privacy-Protected Compressed Learning . 53
Ablation Study . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 A.3.1
Sensitivity analysis for hyperparameters . . . . . . . . . . . 53
A.3.2
Analysis on Inference Complexity . . . . . . . . . . . . . . . 57
Training Stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57 A.4.1
Analysis on Performance vs. Acceleration . . . . . . . . . . 58
Appendix B: Additional Experimental Details
59
References
63
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
1
4
I NTRODUCTION
Magnetic Resonance Imaging (MRI) is a cornerstone for medical diagnostics, renowned for its high contrast and non-invasive nature in providing detailed images of soft tissues. However, the clinical utility of MRI is often limited by the lengthy scan time required to acquire necessary k-space measurement for high-quality imaging. The drive to expedite this process without compromising diagnostic accuracy has stimulated the development of compressed sensing MRI (CS-MRI). Contrary to traditional MRI methods, CS-MRI exploits the inherent sparsity in MR images to reconstruct images with clinically acceptable quality from significantly fewer samples. Existing studies on CS-MRI predominantly focus on optimizing image reconstruction using deep learning techniques. Early studies on deep learning-based CS-MRI primarily learn the mapping from the subsampled k-space measurements to the MR images [1], [2], [3]. Recently, joint optimization of MRI sampling patterns and reconstruction has garnered substantial interest [4], [5], [6], [7], [8], [9]. However, these methods are trained for a specified sampling ratio, and thus cannot support adaptive sampling via a single model. To achieve adaptive sampling and flexibly control sampling ratios, active acquisition methods [10], [11], [12], [13] formalize sampling optimization as a Markov decision process to iteratively generate sampling patterns based on preceding sampling patterns, and the acquired measurements and the corresponding reconstructed MR images. However, these methods heavily rely on a pre-trained reconstruction model that remains fixed without fine-tuning during sampling optimization. The reconstruction model could fail to align with the sampling patterns generated by the trained sampling strategy and yield potentially degraded performance [8]. Beyond reconstruction-centric CS-MRI methods that aim to enhance the reconstruction fidelity of MR images, recent Task-adapted CS-MRI methods [14], [15], [16] directly optimize MR pipeline for accommodating downstream clinical applications [17]. These methods perform downstream tasks directly from k-space measurements without realizing reconstruction or considering reconstruction as an auxiliary process. Un-
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
5
fortunately, existing task-adapted CS-MRI methods cannot fully address the inherent uncertainty problem that critically affects their reliability in clinical applications [14], [18]. The problem arises from the ambiguity in achieving accurate clinical diagnosis due to the factors such as the vague boundaries caused by low imaging quality and diverging expert opinions on the same MR images. The only one deterministic diagnosis could suffer from the uncertainty problem that leads to misdiagnosis or ineffective treatment by disregarding other potentially valid interpretations. The uncertainty becomes even more pronounced for CS-MRI following the Markov chain of “task T → ground truth image X → measurement Y ”, especially at low sampling ratios. Due to the data-processing inequality [19], the uncertainty of task T is inevitably greater given the measurement Y than given the ground truth image X . The challenge is also supported by the fact that the presumption of deterministic mapping from Y to T is invalid for highly undersampled Y [15]. LI-Net [15] partially addresses the uncertainty problem for segmentation by leveraging latent variables to ensure that predictions lie on the segmentation manifold. However, it remains deterministic and fails to capture all plausible solutions consistent with the measurements. To address these issues, in this paper, we propose an information-theoretic optimization framework for task-adapted CS-MRI that for the first time simultaneously achieves probabilistic inference for uncertainty prediction and adapts to arbitrary sampling ratios and versatile clinical applications. The proposed framework provides a unified method to optimize acquisition, reconstruction, and downstream clinical tasks for CS-MRI in an end-to-end fashion based on the information-theoretic metric (ITM) [20]. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance in task-adapted CS-MRI for various clinical tasks. The contributions of this paper are summarized as below. •
Probabilistic inference for uncertainty prediction: The proposed framework formalizes
task-adapted CS-MRI optimization by maximizing the mutual information between undersampled k-space measurements and clinical tasks to address the uncertainty problem of inferring task T from measurement Y . Unlike [15] that solely predicts segmentation in a deterministic manner, the proposed framework achieves probabilistic uncertainty
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
6
prediction for both reconstruction-centric and task-adapted CS-MRI. •
End-to-end optimized adaptive sampling: The proposed framework leverages amortized
optimization [21] and constructs tractable variational bounds for mutual information to support arbitrary sampling ratios for CS-MRI with a single end-to-end trained model. Adaptive sampling is jointly optimized with reconstruction and task-inference models to achieve efficient learning and significant performance gains. •
Versatile adaptation to clinical scenarios: The proposed framework provides a unified
approach for a wide range of clinical scenarios characterized by i) jointly optimizing downstream tasks and MR image reconstruction to enhance task performance with auxiliary reconstruction process, and ii) implementing downstream tasks while suppressing MR image reconstruction for scenarios involving privacy-preserving clinical diagnosis and compressed learning [22]. The rest of the paper is organized as follows. Section 2 reviews the related literature. Section 3 elaborates the proposed information-theoretic optimization framework, including the formulation of optimization problems, adaptive sampling via amortized optimization, and procedures addressing the two typical clinical scenarios. Section 4 presents experimental results and comparisons with existing methods. Finally, Section 5 concludes the paper. Throughout this paper, we use normal symbols for scalars and bold symbols for vectors. We use upper italic cases to represent random variables (e.g., X , Y , and T ) and lower bold cases for their realizations (e.g., x, y, and t).
2
R ELATED W ORKS
2.1
Deep Learning for MR Sampling Patterns
Existing approaches for optimizing MRI sampling patterns can be broadly categorized into two classes: end-to-end methods and active acquisition methods. End-to-end methods. These methods jointly optimize the sampling pattern along with the reconstruction model [4], [5], [6], [7], [8], [9]. To enable the use of stochastic gradientbased optimization for model parameters, these approaches typically require a continuous relaxation of the binary sampling patterns.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
7
Active acquisition methods. In these approaches, sampling patterns are generated iteratively based on previously acquired information [10], [11], [12], [13], [23]. A pioneering work [10] in this category selects the next measurement that minimizes the current uncertainty. More recently, reinforcement learning tools have been adopted, framing the problem as a partially observable Markov decision process [11], [12], [13]. Different from end-to-end optimized methods that are generally specialized for a fixed sampling ratio, active acquisition methods allow for continuous control of the sampling ratio due to their sequential nature. However, in terms of performance, endto-end methods significantly outperform active acquisition methods, as demonstrated in [8]. This performance gap is primarily attributed to training-test mismatches arising from the reliance on pre-trained reconstruction models in active acquisition methods. 2.2
Optimization of Mutual Information
In this paper, we aim to leverage mutual information for optimizing MR sampling patterns. The criterion for optimizing linear projection (sampling pattern) in CS based on mutual information has its roots in Bayesian CS [24], [25]. However, the application of this criterion presents a challenge due to the well-known difficulty in estimating or optimizing Mutual Information. Traditional methods typically rely on approximate Bayesian inference [24], [26] to handle the intractable posterior involved in mutual information. Recently, approaches utilizing amortized variational inference have been introduced to enable more efficient and accurate posterior inference [27], [28], as well as contrastive bounds to bypass explicit posterior computation [29]. Additionally, mutual information can be estimated using various techniques, including ratio estimation [30], neural estimation [31], and Laplace importance sampling [32], among others. 2.3
Task-adapted CS-MRI
Task-adapted CS-MRI focuses on optimizing the performance of medical tasks based on subsampled k-space measurements, where reconstruction either does not occur or serves as an auxiliary process. Here, we briefly review the two categories of application scenarios for task-adapted CS-MRI.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
8
Joint Task and Reconstruction. Simultaneously inferring T and reconstructing X could provide references or evaluate model predictions for clinical decision making. Pioneering work [17] achieved joint reconstruction and segmentation from undersampled k-space measurement based on patch-based sparse representation and Gaussian mixture models. In recent years, deep learning has significantly advanced task-adapted CS-MRI. Most deep learning-based task-adapted CS-MRI methods [7], [33], [34], [35], [36] adopt a two-stage model design: first, constructing a reconstruction model with k-space measurements as input, and then cascading a task-inference model for clinical tasks. Both models are jointly optimized using a multi-task loss. More recently, [16] introduced a similar model design but proposed a two-step training strategy. This approach trains the reconstruction model independently and fine-tunes both models for clinical tasks, optimizing solely the task-specific loss during the second step. Compressed Learning. Compressed learning focuses on inferring downstream tasks [4], [22] without requiring reconstruction results. This approach is particularly desirable for privacy-preserving medical applications, such as MRI [22], where sharing sensitive patient data is not feasible. In [15], a direct mapping from undersampled k-space measurement to segmentation maps was proposed. However, such methods fail to address the inevitable uncertainty caused by information loss in CS-MRI. LI-Net [15] partially addresses the uncertainty problem by generating segmentation maps based on latent variables to ensure that predictions lie on the segmentation manifold. Nonetheless, this deterministic approach does not fully resolve the uncertainty issue, as it cannot capture all plausible solutions consistent with the measurement. In Section 3.6, we make a comprehensive comparison between the proposed method and LI-Net. 2.4
Stochastic Medical Image Segmentation
The advent of deep learning has revolutionized the field of medical image segmentation [37]. A critical challenge in this domain is that inferring segmentations from medical images is inherently ambiguous, with annotations from domain experts often exhibiting significant variability. Conventional methods typically rely on pixel-wise categorical distributions to model the conditional distribution of segmentation T given an MRI
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
9
image X . However, this approach can be insufficient due to the multimodal nature of the distribution and the necessity for spatial coherence in segmentation [14]. To address these challenges, stochastic semantic segmentation methods have emerged, leveraging more expressive distributions to model segmentation uncertainty. These methods primarily integrate deep generative models to generate multiple plausible predictions for a single image, such as variational autoencoders [18], [38], [39], autoregressive models [40], normalizing flows [41], and diffusion models [42]. Additionally, several approaches avoid deep generative models, instead employing techniques such as lowrank multivariate normal distributions to model logit distributions [14] or mixtures of stochastic experts for mode estimation [43].
3
P ROPOSED M ETHOD
This section delves into the optimization of task-adapted CS-MRI from an informationtheoretic perspective, enabling probabilistic inference for uncertainty prediction and achieving an end-to-end optimized adaptive CS-MRI framework. Section 3.1 introduces the preliminaries on information-theoretic metrics (ITMs) and the linear feature design problem. Section 3.2 formulates a linear feature design problem based on ITMs for optimizing task-adapted CS-MRI given a fixed sampling ratio. Section 3.3 further solves multiple problem instances of task-adapted CS-MRI characterized by varying sampling ratios, and achieves adaptive sampling through a single end-to-end trained model via amortized optimization. Section 3.4 achieves joint optimization of downstream tasks and MRI signal reconstruction, and Section 3.5 realizes downstream tasks by removing irrelevant MRI reconstruction to enable privacy preserving medical diagnosis and compressed learning. Section 3.6 highlights the contributions of the proposed method and discusses its relations to existing methods. 3.1
Preliminary: Information Theoretic Metric
Consider the linear observation model y = Φx+n, where Φ ∈ CM ×N is the measurement matrix, n represents the noise, x and y are realizations of the signal X and observation Y ,
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
10
respectively. The Information Theoretic Metric (ITM) [20] provides a concise framework for optimizing Φ given the downstream task T that can be inferred from X , also known as the linear feature design problem. ITM builds a Markov sequence T → X → Y 1 to characterize X , Y , and T , and optimizes Φ by balancing the mutual information I(Y ; T ) for measuring the information preserved in the observation Y for the downstream task T and I(Y ; X) for assessing the quality of recovering the signal X from Y : max ITM(Φ, β) := max{I(Y ; T ) − βI(Y ; X)}. Φ
Φ
(1)
Here, the parameter β ∈ R controls the balance between I(Y ; T ) and I(Y ; X). When β < 0, signal recovery and downstream task inference are simultaneously considered. In
the extreme cases, the downstream task inference is prioritized in analogy to compressed learning [22] for β = 0, while signal recovery is solely considered as β → −∞. On the contrary, when β > 0, ITM behaves similarly to the information bottleneck (IB) [44], [45] that maximally preserves the information about T in Y , while discarding the remaining irrelevant information. Noting that Φ does not explicitly appear in ITM(Φ, β ) but influences Y . We omit Φ (or equivalent symbols) without introducing confusion.
3.2
End-to-end Variational Information Optimization for Task-Adapted CS-MRI
CS-MRI accelerates the MRI acquisition process with subsampling in the k-space. Given a sampling ratio r = M/N < 1, subsampled measurements y ∈ CM for a vectorized MR image x ∈ CN are acquired in the k-space using a linear projection model with Φ = Um F , where Um ∈ RM ×N is an undersampling matrix selecting M rows from the N ×N identity matrix according to the binary sampling pattern m ∈ {0, 1}N and F ∈ CN×N denotes the discrete Fourier transform. y = Um Fx + n.
(2)
Here, n ∈ CM ∼ Nc (0, σ 2 I) represents additive complex white Gaussian noise [46] with a standard deviation of σ . For instance, when N = 5, if m = [1, 1, 0, 1, 0], Um ∈ R3×5 is 1. Our framework’s validity hinges on the conditional independence T ⊥ Y | X , meaning the measurements Y are uninformative about the task T given the image X . This assumption is invariant to the causal direction, making graphical models like T → X → Y and T ← X → Y equivalent.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
11
formed by selecting the first, second, and fourth rows of the 5×5 identity matrix. Without loss of generality, consider x and y are realizations from X and Y that represent the underlying distribution of MR images and their measurements, respectively. According to (2), the conditional distribution of Y given X = x can be characterized as an isotropic Gaussian with mean Um Fx and variance σ 2 , i.e., p(y|x) = Nc (y|Um Fx, σ 2 I). Note that, in real-world clinical scenarios, noise may exhibit non-Gaussian characteristics, e.g., Rician, Poisson or mixed noise. We could adapt to other noise models by modifying the likelihood term p(y|x). For task-adapted CS-MRI, finding an optimal sampling pattern m can be formulated as a specific instance of the linear feature design problem, i.e., optimizing the linear projection Φ = Um F by maximizing the ITM objective in (1). Given that the discrete Fourier transform remains fixed, the focus shifts to optimizing m. In CS-MRI, m satisfies that ∥m∥0 = M and some other physical constraints C imposed by the real-world MR scanners. From (1), we formulate the ITM-based objective function for task-adapted CSMRI as max {I(Y ; T ) − βI(Y ; X)} , m
s.t. m ∈ M,
(3)
where M = m : m ∈ {0, 1}N , ∥m∥0 = M, M < N ∩ C represents the sampling pattern constraint. It is intractable to compute the mutual information between high-dimensional random variables for optimizing (3). As an alternative, we propose tractable proxy objectives based on variational inference for the case that β < 0 and marginal approximation for the case that β ≥ 0. 3.2.1
β<0
We first establish a tractable lower bound for (3) using variational inference in Proposition 1 for β < 0. Proposition 1. Let t and y be the realizations of T and Y , respectively. Given a variational approximation q(t|y) for the posterior p(t|y), the mutual information I(Y ; T ) satisfies
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
12
that I(Y ; T ) =Ep(y) [DKL (p(t|y)∥q(t|y))] + H(T ) + Ep(t,y) log q(t|y),
(4)
where DKL denotes the Kullback-Leibler (KL) divergence and H(T ) is the entropy of T . Furthermore, the Barber-Agakov lower bound [47] of I(Y ; T ) is I(Y ; T ) ≥ H(T ) + Ep(t,y) log q(t|y).
(5)
Let x be the realization of X . The variational lower bound of mutual information I(Y ; X) is similarly obtained by I(Y ; X) ≥ H(X) + Ep(x,y) log q(x|y).
(6)
Proposition 1 suggests that the Barber-Agakov lower bounds (5) and (6) are tight, only when q(t|y) = p(t|y) and q(x|y) = p(x|y) for all possible y. In this case, maximizing the ITM-based objective function in (3) is approximately equivalent to maximizing Ep(x,t,y) [log q(t|y)−β log q(x|y)]. This fact implies that q should approximate p to optimize m. Note that the data entropy terms H(T ) and H(X) are omitted in optimizing m, since
they do not depend on m. To this end, we develop the proxy objective for maximizing (3) as max Ep(x,t,y) [log q(t|y) − β log q(x|y)], m,q
s.t. m ∈ M.
(7)
According to (7), sampling pattern m, task inference q(t|y), and MR image reconstruction q(x|y) can be optimized in an end-to-end fashion. The sampling pattern is adapted to the downstream task and signal recovery for enhanced overall performance. In Section 3.4, we provide a detailed explanation of the procedure for solving (7). 3.2.2
β≥0
The objective function in (3) cannot be maximized using the lower bound of I(Y ; X) for β ≥ 0. Since I(Y ; X) = H(Y ) − H(Y |X) and H(Y |X) remains fixed for all m ∈ M, the
objective in (3) is reformulated as max {I(Y ; T ) − βH(Y )} , m
s.t. m ∈ M.
(8)
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
13
In (8), the regularization term −βH(Y ) constrains the marginal entropy of Y . We develop a tractable proxy objective for (8) by introducing variational approximation q(t|y) for p(t|y) and marginal entropy estimator Ĥ(Y ) for H(Y ). n o max Ep(x,t,y) [log q(t|y) − β Ĥ(Y )] , m,q
s.t. m ∈ M.
(9)
We calculate Ĥ(Y ) and solve (9) in Section 3.5. 3.3
Adaptive Sampling via Amortized Optimization
Both the optimization problems (7) for β < 0 and (9) for β ≥ 0 are formulated to obtain a single sampling pattern m for a specified sampling pattern constraint M, which corresponds to a fixed sampling ratio r. This necessitates training separate models for different sampling ratios. In this section, we further consider solving multiple problem instances of (3) characterized by varying sampling pattern constraint M to accommodate to arbitrary sampling ratios via a single end-to-end trained model, i.e., achieving adaptive sampling. Note that it differs from scalable sampling, which emphasizes progressive sampling and signal reconstruction. To address the computational complexity arising from potentially infinite sampling ratios, we employ amortized optimization [21] to efficiently produce a stochastic sampling strategy that maps r to a distribution of m, supported on the sampling pattern constraint M. The strategy is implemented by a pattern generation network (PGN) with learnable
parameters θ. Given a distribution p(r) over r, the objective for optimizing the PGN is max Er∼p(r) EM∼p(M|r) Eπθ (m|M) [I(Y;T )−βI(Y;X)], θ
(10)
where πθ (m|M) denotes the parametric conditional probability that generates m given M. We then substitute the intractable quantity I(Y ; T ) − βI(Y ; X) in (10) with tractable
variational bounds obtained in Section 3.2 to enable optimization. Define {r, M, m, x, t, y} ∼ p(r)p(M|r)πθ (m|M)pdata (x, t)p(y|x). When β < 0, we introduce (7) and obtain: max Er,M,m,x,t,y [log q(t|y) − β log q(x|y)], θ,q
(11)
and for β ≥ 0, we consider (9) such that: max Er,M,m,x,t,y [log q(t|y) − β Ĥ(Y )]. θ,q
(12)
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
𝒇𝜽
𝑝(𝑟) ~
𝒂
14
𝒃
𝒓
𝑟
Expand
Uniform distribution
Conv 1x1
Conv 1x1
Rescale
ReLU
Sigmoid
𝒪𝑅
Concat Positional embedding map
Hidden layers
𝝁𝜽 (𝒓)
Fig. 1: Illustration of the parametric mapping r 7→ µθ (r).
We then describe the implementation of p(r)p(M|r)πθ (m|M). Given that the constraint M on the sampling pattern m is considered for any sampling ratio r ∈ [0, 1], we define it as follows: M = {m ∈ {0, 1}N : |∥m∥0 − rN | < ϵ},
(13)
where ϵ is a tolerance hyperparameter to ensure that the sampling pattern is consistent with the sampling ratio. This implies that the unknown distribution p(M|r) can alternatively be achieved using a specified distribution p(r) over the sampling rate r according to (13). In this paper, we set p(r) to be a uniform distribution U[a, b] with 0 ≤ a ≤ b ≤ 12 . Once M is produced from p(r)p(M|r), we generate the sampling pattern m using the proposed PGN, i.e., m ∼ πθ (m|M), and then decide to accept m as a sample or reject it and return to the sampling step, depending on whether m belongs to M. This technique is known as rejection sampling [48]. If a sample of m is accepted, it is used to generate the corresponding measurements, obtain the final outputs, and optimize the whole MRI model for downstream tasks and signal recovery. We now return to the proposed PGN, which implements a stochastic sampling strategy to generate the sampling pattern m in a differentiable manner, using input r 2. The uniform distribution corresponds to a scenario with no preference for sampling rates within the specified range [a, b]. However, in cases where there is a preference for a specific range of sampling rates, alternative distributions may be considered.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
15
sampled from U[a, b]. We first define the distribution πθ (m|M) over m by treating each element mi of m as an independent Bernoulli random variable, as follows: (N ) Y πθ (m|M) ∈ B(mi |µi ) : µi ∈ [0, 1] for all i ,
(14)
i=1
where the subscript i indexes the k-space position, µi represents the probability of sampling the point at position i, and B(mi |µi ) denotes that p(mi = 1) = µi and p(mi = 0) = 1 − µi . To optimize the sampling strategy in a differentiable manner, we replace µi with a parametric mapping r 7→ µθ (r) that maps the sampling ratio r to a set of
valid probabilities µθ (r), where each element µθ (r)i lies in [0, 1]. We define a proposal distribution π̃θ (m|M) for πθ (m|M) given r as π̃θ (m|M) =
N Y
B(mi |µθ (r)i ).
(15)
i=1
Thus, πθ (m|M) is determined by accepting samples m from π̃θ (m|M) when m ∈ M, or rejecting them and re-implementing (15) otherwise. The parametric mapping r 7→ µθ (r) is defined as µθ (r) = fθ [Concat(r · 1N , Mpe )],
(16)
where Mpe ∈ RN is a learnable matrix corresponding to position embedding, 1N denotes an all-one vector of dimension N , and fθ is a simple neural network whose input is the concatenation of r · 1N and Mpe . In this paper, we set fθ as a multilayer perceptron (MLP) with one hidden layer, followed by a sigmoid activation and a rescale operator proposed in [49], as illustrated in Fig. 1. Note that the linear layers of the MLP are implemented using equivalent one-by-one convolutions. Given arbitrary b ∈ RN , the rescale operator OR is defined as OR (b) =
rN b,
if ∥b∥1 ≥ r
1 − N −rN (1 − b),
otherwise
∥b∥1
N −∥b∥1
.
(17)
From (17), OR (b) ∈ [0, 1]N and ∥OR (b)∥1 = rN = M , thereby guaranteeing that m ∈ M with high probability. We illustrate the training steps with respect to sampling in Algorithm 1. Once the model is well-trained, one can generate m according to steps 3-10 in Algorithm 1 for a specified r, achieving adaptive sampling via a single model.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
16
Algorithm 1 Adaptive Sampling via Amortized Optimization Input: Tolerance hyperparameter ϵ, MRI data dimension N , PGN with parameter θ, constants a, b with 0 ≤ a ≤ b ≤ 1 1: while training the CS-MRI model do 2:
Sample r from the uniform distribution r ∼ U[a, b]
3:
Determine M according to (13)
4:
Obtain µθ (r) through PGN according to (16)
5:
Generate m through (15)
6:
if m ̸∈ M then
7: 8: 9:
Reject m, re-implement (15), and return to step 5 else {m ∈ M} Accept m as a sample
10:
end if
11:
Use m to generate measurements, obtain final outputs, and optimize the whole MRI model (including PGN)
12: end while
3.4
Tractable Optimization With Variational Bounds (β < 0)
In this section, we solve (11) in the case that β < 0 for joint optimization of downstream tasks and MRI reconstruction. We optimize the tractable variational lower bounds for (11) given the sampling pattern m obtained through Algorithm 1. To solve (11), it is essential to construct tractable formulations of q(t|y) and q(x|y), balancing computational efficiency with fidelity to the true posteriors p(t|y) and p(x|y). Note that there exists statistical correlation between elements of t conditioned on y for several downstream tasks, such as segmentation. To capture this dependency, we introduce an auxiliary latent random variable Z with a prior distribution q(z|y) and a conditional distribution q(t|z, y), enabling the reformulation of q(t|y) through marginalization as follows: Z q(t|y) =
q(t|z, y)q(z|y) dz.
(18)
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
17
In this section, we use Segmentation-adapted CS-MRI as an example to elaborate on the optimization steps with an auxiliary latent variable for this task3 . While statistical correlations also exist between elements of the MRI reconstruction x conditioned on y, we model q(x|y) as a diagonal Gaussian distribution in this paper. This approach
is chosen to balance computational efficiency with accuracy, as maximizing log q(x|y) is equivalent to minimizing the mean square error (MSE) between the signal x and a transformation of y (see more details in Section 3.6), a typical objective in signal reconstruction tasks. Directly evaluating the log-likelihood of q(t|y) in (18) remains generally intractable. To address this, we construct the Evidence Lower Bound (ELBO, [50]) for log q(t|y) by introducing a variational approximation q̂(z|t, y) for the intractable latent posterior q(z|t, y) ∝ q(t|z, y)q(z|y):4 log q(t|y) ≥ ELBO := Eq̂(z|t,y) [log q(t|z, y) − DKL [q̂(z|t, y)||q(z|y)]].
(19)
By substituting log q(t|y) in (11) with its ELBO from (19), we obtain a tractable variational lower bound for (11): max Er,M,m,x,t,y Eq̂(z|t,y) [log q(t|z, y) − DKL [q̂(z|t, y)||q(z|y)] − β log q(x|y)], θ,ϕ
(20)
where DKL represents the Kullback–Leibler (KL) divergence. Notably, (20) offers intuitive interpretations. The first term, log q(t|z, y), aims to reconstruct the ground truth t sampled from training data given the measurement y and the latent variable z generated from q̂(z|t, y). Since q̂(z|t, y) is conditioned on
the ground truth t, an ideal reconstruction of t can be achieved. The second term, DKL [q̂(z|t, y)||q(z|y)], encourages q(z|y) to generate latent variables z that resemble those
produced by q̂(z|t, y). Here, q̂(z|t, y) acts as a “teacher” distribution that leverages the information of t, while q(z|y) serves as a “student” distribution, lacking information 3. We demonstrate in Section 4.2.2 that omitting the correlation between elements, in the absence of an auxiliary latent variable, significantly degrades segmentation performance. 4. We omit the derivation of (19) as it follows standard variational inference principles; see [50] for a detailed explanation.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
18
on t and attempting to approximate q̂(z|t, y) “in an average sense”. We refer to this approximation as the average sense because the optimal solution of q(z|y) that minimizes Er,M,m,x,t,y Eq̂(z|t,y) DKL [q̂(z|t, y)||q(z|y)] is Ep(t|y) [q̂(z|t, y)], referred to as the marginal or aggregated latent posterior [51]. The third term, log q(x|y), seeks to reconstruct the MRI signal x based on the measurement y, thereby ensuring data fidelity. Starting from here, we present the solving steps for (20). As mentioned in Section 3.1, the sampling pattern m influences the measurement y, implying that the posterior distributions q(t|y) and q(x|y) are dependent on m. Therefore, it is necessary to condition these distributions on m, meaning that the input for the task-inference network and signal-reconstruction network should contain information about both y and m. Due to the mismatched dimensionality and different data types of y (continuous) and m (discrete), it is undesirable to directly use the paired data (y, m) to optimize the MRI models. To address this, we use the zero-filling reconstruction x̂zf := F H Um T y ∈ CN as a representation of (y, m), as it contains all the information about (y, m). Specifically, y can be recovered from x̂zf by Um F x̂zf , and m can be recovered by identifying the zero elements in F x̂zf . To optimize (20), we implement probabilistic modeling for q(t|z, y), q(z|y), q̂(z|t, y), and q(x|y) with the latent variable model (LVM) [50], a powerful off-the-shelf probabilistic model that can be trained by maximum likelihood. Since the input for the taskinference network and signal-reconstruction network is x̂zf := F H Um T y, we first encode the information of x̂zf into the feature ŷ by a measurement encoder EϕY : ŷ = EϕY (F H Um T y).
(21)
ŷ is then fed into the latent decoder DϕZ , task-inference decoder DϕT , and signal-reconstruction
decoder DϕX to characterize q(z|y), q(t|z, y), and q(x|y), respectively. A typical choice of variational family to formulate these conditional probability distribution is the meanfield variational family, where the elements of a variable are assumed to be mutually independent [52]. q(x|y) is modeled as a diagonal Gaussian distribution whose mean and variance are produced by DϕX : q(x|y) = N (x|DϕX (ŷ)).
(22)
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
19
q(z|y) is also modeled as a diagonal Gaussian distribution with learnable mean and
variance that is commonly adopted for modeling the conditional distribution of latent variables. q(z|y) = N (z|DϕZ (ŷ)).
(23)
q(t|z, y) is formulated associated with the tasks. Taking MRI segmentation task as an
example, q(t|z, y) is modeled as an independent Categorical distribution for each pixel, whose class probability is produced by DϕT (z, ŷ). q(t|z, y) =
Y
Categorical(ti |DϕT (z, ŷ)i ).
(24)
i
As the KL divergence of q̂(z|t, y) aims to match the prior distribution q(z|y), we also formulate q̂(z|t, y) as a diagonal Gaussian distribution, using the same decoder DϕZ as q(z|y) and another measurement encoder Eϕ̃Y : q̂(z|t, y) = N (z|DϕZ (Eϕ̃Y (t, F H Um T y))),
(25)
where Eϕ̃Y shares the same architecture to EϕY except the input has an additional channel for the ground truth segmentation map t. Note that (23) and (24) can be performed multiple times to generate different conceivable segmentation maps that are consistent with the posterior samples from p(t|y), thereby naturally addressing the uncertainty problem imposed by the MR subsampling discussed in Section 2.3. The training steps for optimizing (20) are summarized in Algorithm 2 and illustrated in Fig. 2. The trained model is then used to obtain the MRI reconstruction and a series of segmentation maps according to Algorithm 3. 3.5
Tractable Optimization With Marginal Entropy Minimization (β ≥ 0)
In this section, we consider the case that β ≥ 0 in (12), i.e., implementing downstream tasks without reconstructing the original image, including privacy preserving clinical decision making and compressed learning. We solve (12) with β ≥ 0 by maximizing the log-likelihood of q(t|y) and minimizing the estimated marginal entropy Ĥ(Y ). For tasks that consider statistical correlations between elements of t conditioned on y, the log-likelihood of q(t|y) is maximized using the approach described in Section 3.4,
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
20
Algorithm 2 Optimizing (20) for Segmentation-Adapted CS-MRI Input: Initialized encoders and decoders: measurement encoders EϕY and Eϕ̃Y , latent decoder DϕZ , task-inference decoder DϕT , signal-reconstruction decoder DϕX . 1: while ϕ, θ not converge do 2:
Sample x, t from the training set: x, t ∼ pdata (x, t)
3:
Sample m according to Algorithm 1
4:
Obtain y and x̂zf = F H Um T y ∈ CN
5:
Encode x̂zf in the feature ŷ following (21)
6:
Sample z ∼ q̂(z|t, y) following (25)
7:
Evaluate log q(x|y) following (22)
8:
Evaluate log q(t|z, y) following (24)
9:
Evaluate DKL [q̂(z|t, y)∥q(z|y)] following (23), (25)
10:
Do back-propagation of (20) to compute gradient with respect to ϕ and θ
11: end while
Algorithm 3 Inference for Segmentation-Adapted CS-MRI Input: Measurement encoder EϕY , latent decoder DϕZ , segmentation decoder DϕT , image decoder DϕX , measurement y, sampling pattern m. 1: Obtain y and x̂zf = F H Um
T
y ∈ CN
2: Encode x̂zf in the feature ŷ following (21) 3: Obtain posterior approximation q(x|y) following (22) 4: while obtaining sufficient segmentation maps do 5:
Sample z ∼ q(z|y) following (23)
6:
Sample t ∼ q(t|z, y) following (24)
7: end while
Output: q(x|y) and a series of segmentation maps ts. The mean of q(x|y) can be taken as a specific MRI reconstruction image.
which involves introducing an auxiliary latent variable. However, tasks producing scalar outputs such as pathological classification do not require capturing statistical correla-
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
PGN
𝒙, 𝒕 ~ 𝑝𝑑𝑎𝑡𝑎 (𝑥, 𝑡)
𝒎
𝒚
𝓕𝐻 𝑼𝑇𝒎 𝒚
𝒙 𝒕
ℰ𝜙
ෝ 𝒚
Concat
𝒓 ~ 𝒰[𝑎, 𝑏]
Concat
𝐷𝜙𝑋
ෝ 𝑞(𝒙|𝒚) = 𝒩 𝒙 𝒟𝜙𝑋 𝒚
Variational Lower Bound as Loss Function
𝐷𝜙𝑇
ෝ 𝑖) 𝑞 𝒕 𝒛, 𝒚 = ෑ Categorical(𝑡𝑖 |𝒟𝜙𝑇 𝒛, 𝒚
max 𝜃,𝜙 𝔼𝑟,𝓜,𝒎,𝒙,𝒕,𝒚 𝔼𝑟(𝒛|𝒙,𝒚)
𝑖=1
𝒛
ℰሚ𝜙
21
𝒟𝜙𝑍
[log 𝑞 𝒕 𝒛, 𝒚 − 𝛽 log 𝑞(𝒙|𝒚) ෝ 𝑞 𝒛 𝒚 = 𝒩 𝒛 𝒟𝜙𝑍 𝒚 −𝐷𝐾𝐿 [𝑟(𝒛|𝒕, 𝒚)| 𝑞 𝒛 𝒚 ]
𝑟 𝒛 𝒕, 𝒚 = 𝒩 𝒛 𝒟𝜙𝑍 (ℰሚ𝝓 (𝒕, 𝓕𝐻 𝑼𝑇𝒎 𝒚))
Fig. 2: Illustration of Algorithm 2 for optimizing (20). At each training iteration, a pair of data (x, t) is sampled from the training set, along with a sampling ratio r drawn from U[a, b]. Using r as the input, the PGN generates a sampling pattern m, which is then
used to create the measurements y. The information in y is encoded into the feature ŷ by passing the zero-filled reconstruction F H Um T y through the measurement encoder EϕY . Subsequently, ŷ is fed into the decoders DϕX , DϕT , and DϕZ to calculate q(x|y), q(t|z, y), q(z|y), and q̂(z|t, y). The variational lower bound (loss function) is then evaluated, and
the entire MRI model comprising the PGN, encoders EϕY , Eϕ̃Y , and decoders DϕX , DϕT , DϕZ is optimized through backpropagation.
tions and introducing an auxiliary latent variable. In this section, we use Classificationadapted CS-MRI as an example to illustrate the steps for maximizing log q(t|y). Specifically, q(t|y) can be directly modeled as a Categorical distribution, with class probabilities produced by DϕT (ŷ): q(t|y) = Categorical t|DϕT (ŷ) .
(26)
Note that, different from (24), (26) does not involve the pixel-wise product of the probability and the latent input z for DϕT . Subsequently, we develop Ĥ(Y ) to estimate H(Y ) with three steps, including i) approximating the marginals of the noisy k-space measurement y, ii) modeling the statistical dependencies due to k-space conjugate symmetry, and iii) estimating the marginal entropy H(Y ). Step i): Approximating marginals of y. We first model the marginals of the underlying noiseless k-space measurement at each position ki = (Fx)i , denoted by q(ki ) for
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
22
i = 1, · · · , N . Since the distribution over the noisy k-space measurement yi given ki is a
Gaussian Nc (yi |ki , σ 2 ) according to (2), we consider modeling q(ki ) as Gaussian5 , which allows for the computationally tractable approximation q(yi ) for the true marginal of yi through analytical Gaussian marginalization: Z q(yi ) = Nc (yi |ki , σ 2 )q(ki )dki .
(27)
Step ii): Modeling statistical dependencies for q(y). From (27), the marginal of the noisy k-space measurement y can be simply formulated as Y q(y) = q(yi ), i = 1, 2, ..., N.
(28)
i:mi =1
However, q(y) can be further analyzed due to the conjugate symmetry property of kspace [53], i.e., statistical dependencies between the sample position i and its conjugate symmetry position i∗ = N − i (also referred to as point-reflected position). Given ki , we have the following relation between two measurements [53]: yi = ki + ni , ni ∼ Nc (0, σ 2 I),
(29)
yi∗ = k̄i + ni∗ , ni∗ ∼ Nc (0, σ 2 I),
(30)
where k̄i is the complex conjugate of ki . Due to (29) and (30), sampling at position i∗ is redundant when sampling at position i. Since the joint distribution of yi , yi∗ given ki is the bivariate Gaussian Nc ((yi , yi∗ )T |(ki , k̄i )T , σ 2 I), we can approximate the joint marginal of yi , yi∗ similarly to (27): Z q(yi , yi∗ ) =
Nc ((yi , yi∗ )T |(ki , k̄i )T , σ 2 I)q(ki )dki .
(31)
Let mi denote the i-th element of current sampling pattern m. Define I = {i : mi = 1, mi∗ = 1} as the set of positions where both the sample and its point-reflected position
are included, and J = {j : mj = 1, mj∗ = 0} the set of positions where the sample is included but its point-reflected position is not. We obtain from (27) and (31) that " # 12 Y Y q(y) = q(y1 , y2 , · · · , yM ) = q(yi , yi∗ ) q(yj ), i∈I
(32)
j∈J
5. Clearly, assuming q(ki ) to be non-Gaussian introduces computational complexity or intractability, as Eq. (27) no longer holds.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
23
where the square root occurs due to the same distributions are multiplied twice, e.g., q(y1 , yN −1 ) = q(yN −1 , y1 ). Compared to (28), (32) takes into account the statistical depen-
dencies arising from the conjugate symmetry property of k-space, resulting in a more accurate estimation of Ĥ(Y ). Step iii): Estimating marginal entropy H(Y ). We estimate H(Y ) by computing the entropy Ĥ(Y ) of q(y). To this end, we first compute q(yj ) and q(yi , yi∗ ) as in (32) by applying Bayes’ theorem for Gaussian variables [54]. Without introducing ambiguity, we define the distribution of a complex scalar random variable through the joint distribution of its real and imaginary parts, i.e., q(yj ) = q(yjr , yjc ), where the superscripts r and c represent the real and imaginary parts of complex variables, respectively. Suppose that the real and imaginary parts, kir and kic , of the noiseless k-space measurement ki are independent for simplicity. Define the mean and variance of kir /kic over the training data set D as µri /µci and Vir /Vic , respectively, i.e., q(ki ) = N (kir |µri , Vir )N (kic |µci , Vic ).
(33)
Note that µri , µci , Vir , Vic are pre-computed before optimizing the MRI model. Since q(yi |ki ) = N (yir |kir , σ 2 )N (yic |kic , σ 2 ) according to (29), we obtain that Z q(yi ) = q(ki )q(yi |ki )d(ki ) = N (yir |µri , σ 2 + Vir )N (yic |µci , σ 2 + Vic ).
(34)
according to (27). For the positions in J , similar to the derivation of (34), we can easily obtain q(yi , yi∗ ) as follows: r T q(yi , yi∗ ) = N ((yir , yi∗ ) |(1, 1)T µri , σ 2 I + (1, 1)T Vir (1, 1)) c T ) |(1, −1)T µci , σ 2 I + (1, 1)T Vic (1, 1)). · N ((yic , yi∗
(35)
From (34) and (35), we obtain Ĥ(Yi ) and Ĥ(Yi , Yi∗ ) according to the analytical form of entropy of a Gaussian distribution. X 1 Ĥ(Yi ) = (1 + log(2π) + log(σ 2 + Vi· )), 2 ·∈{r,c}
Ĥ(Yi , Yi∗ ) =
X
1 1+log(2π)+ (log σ 2 +log(σ 2 +2Vi· )), 2
·∈{r,c}
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
where Vi = Vir + j · Vic , and further compute Ĥ(Y ) according to (32) as follows: X1 X Ĥ(Y ) = Ĥ(Yi , Yi∗ ) + Ĥ(Yj ). 2 i∈I j∈J
24
(36)
For convenience, Ĥ(Y ) can be computed in a point-wise way for all sampled positions P P as Ĥ(Y ) = i ·∈{r,c} Ĥi· , where Ĥi· is defined as follows for · ∈ {r, c} and i ∈ {I, J }: 1 1 (1+log 2π+ log σ 2 +log(σ 2 +2Vi· ))), i ∈ I 2 . (37) Ĥi· = 2 1 (1+log 2π+log(σ 2 +V · )), i ∈ J i 2 Remark 1. Equation (37) mathematically validates some intuitions about how the choice of m affects I(Y ; X). To find m that maximizes I(Y ; X) (which is equivalent to maximizing the marginal entropy H(Y )) when the measurement noise is relatively small, Algorithm 4 suggests the following: •
Avoid redundant sampling. To maximize the marginal entropy H(Y ), one should avoid
allocating acquisition budget to I since log σ 2 → −∞ as σ → 0. •
Sample positions with high uncertainty. Since Hi is monotonically increasing with respect
to k-space variance Vi at position i, one should prioritize measurements in k-space where the data is most uncertain [25]. We summarize the training steps for optimizing (12) in Algorithm 4. Once the model has been properly trained, classification results can be obtained directly with the suppression of MRI reconstruction, as detailed in Algorithm 5. 3.6 3.6.1
Relation to Existing Methods Relation to LI-Net [15]
LI-Net is designed to achieve segmentation of MRI data directly from undersampled kspace measurement, bypassing the MRI reconstruction step. LI-Net consists of two main components: (1) a segmentation auto-encoder (AE) with an encoder Φ and a decoder Ψ, and (2) a measurement encoder Π(y). The training procedure of LI-Net involves two
independent stages. In stage 1, the AE is trained by minimizing the reconstruction error of the segmentation. min Et∼pdata (t) ∥Ψ ◦ Φ(t) − t∥22 , Φ,Ψ
(38)
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
25
Algorithm 4 Optimizing (12) for Classification-Adapted CS-MRI Input: The k-space variance for the real and imaginary parts over the training set Vr , Vc ∈ RN , variance of the measurement noise σ 2 , and initialized encoder EϕY
and decoder DϕT . 1: while ϕ, θ not converge do 2:
Sample x, t from the training set: x, t ∼ pdata (x, t)
3:
Sample m according to Algorithm 1
4:
Compute elements of point-reflected sampling pattern m̃: m̃i = mi∗ for i = 1, 2, · · · , N.
5:
Find positions that are redundantly sampled or not: I = m ⊙ m̃, J = m ⊙ (1 − I).
6:
Compute Ĥ(Y ) in a point-wise way according to (36) and (37)
7:
Obtain y and x̂zf = F H Um T y ∈ CN
8:
Encode x̂zf in the feature ŷ following (21)
9:
Evaluate q(t|y) = Categorical(t|DϕT (ŷ))
10:
Do back-propagation of (12) to compute gradient with respect to ϕ and θ
11: end while
Algorithm 5 Inference for Classification-Adapted CS-MRI Input: Measurement encoder EϕY , classification decoder DϕT , measurement y, sampling pattern m 1: Obtain y and x̂zf = F H Um
T
y ∈ CN
2: Encode x̂zf in the feature ŷ following (21) 3: while obtaining sufficient classification results do 4:
Sample t ∼ q(t|y) = Categorical(t|DϕT (ŷ))
5: end while
Output: A series of classification results ts
where t ∼ pdata (t) denotes sampling true segmentation maps from the training data. Once the AE is well-trained, its parameters are fixed, and the training proceeds to stage 2. In stage 2, the measurement encoder Π(y) is trained to predict the latent representation
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
26
z = Φ(t) of the true segmentation map t by minimizing the MSE in the latent space: min E{y,t}∼pdata (y,t) ∥Π(y) − Φ(t)∥22 . Π
(39)
When Π is trained, segmentation maps are generated directly from the measurements via t ∼ Ψ ◦ Π(y). The scenario that LI-Net addresses corresponds to (20) with β = 0 as follows: max Er,M,m,x,t,y Eq̂(z|t,y) [log q(t|z, y) − DKL [q̂(z|t, y)∥q(z|y)]. θ,ϕ
(40)
Although LI-Net and our method share similarities—such as optimizing encoder-decoder structures and introducing an auxiliary latent variable—our approach offers significant advantages over LI-Net in terms of theoretical interpretation, training procedure, sampling, and task inference. •
Theoretical interpretation: LI-Net is an intuitively designed “deterministic” method,
while our approach is an information-theoretic “statistical” version. Recall that q̂(z|t, y) and q(z|y) in (40) are modeled as Gaussians (see (25) and (23)), with means and variances generated by learnable decoders. Let µq̂ (t, y), Vq̂ (t, y), µq (y), and Vq (y) represent the means and variances of q̂(z|t, y) and q(z|y), respectively. At inference, our method can become fully deterministic by setting the latent variable to the Gaussian mean rather than sampling it (e.g., z = µq̂ (t, y) rather than z ∼ N (µq̂ (t, y), Vq̂ (t, y))). In this deterministic scenario, the optimal measurement encoder µ∗q (y) minimizes the MSE loss between µq (y) and µq̂ (t, y): µ∗q (y) = arg min E∥µq (y) − µr (t, y)∥22 . µq (y)
(41)
Note that (41) resembles (39), with the distinction that the deterministic segmentation encoder µq̂ (t, y) is also conditional on y. Moreover, LI-Net intuitively aligns the measurement encoder Π(y) with the pre-trained Φ(t) and generates segmentation maps via Ψ ◦ Π(y). In contrast, we rigorously derive a principled framework from an information-
theoretic perspective. •
Training procedure: According to (40) and Algorithm 2, the latent decoder DϕZ , task-
inference decoder DϕT , and measurement encoder Eϕ̃Y are jointly optimized in an end-toend manner, providing a holistic training perspective. By contrast, LI-Net decomposes
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
27
training into two separate stages, which may prevent it from achieving optimal parameters. For example, in LI-Net’s second training stage, the AE parameters are fixed and cannot be co-optimized with Π. •
Sampling and Task inference: LI-Net only supports a fixed sampling ratio during the
training stage and is limited to compressive learning. In contrast, our method enables adaptive sampling with a single model and can be adapted to broader scenarios by controlling β . Additionally, our method provides estimates of prediction uncertainty due to the “statistical” nature of variational inference, thereby enhancing robustness and reliability in automated medical diagnosis systems. Experimental results in Section 4.2 demonstrate the superior performance of our method over LI-Net. 3.6.2
Relation to Existing End-to-End CS-MRI Reconstruction Methods
As mentioned in Section 3.4, we model the MR image posterior, q(x|y), as a diagonal Gaussian distribution. In this formulation, maximizing log q(x|y) is equivalent to minimizing the MSE between the signal x and a transformation of y. Specifically, suppose q(x|y) has the mean of Rϕ (x̂zf ) and the variance of σϕ2 (x̂zf ) according to (22). Then maxθ,ϕ Er,M,m,x,t,y log q(x|y) in (20) is equivalent to # "N X (xi −Rϕ (x̂zf )i )2 +log σϕ2 (x̂zf )i , min Er,M,m,x,y 2 θ,ϕ σ (x̂ ) zf i ϕ i=1
(42)
which aligns with existing end-to-end MRI reconstruction methods that employ an MSE loss [4], [5], [6], [7], [8], [9]. Conversely, these end-to-end methods can be viewed as implicitly maximizing I(Y ; X). A key distinction in our method is that it also learns the uncertainty in MRI reconstruction. According to the first-order optimality condition of (42), the optimal variance σϕ2 (x̂zf )i is the minimum mean square estimator (MMSE) of the reconstruction error
given xzf , i.e., E[(xi − Rϕ (x̂zf )i )2 |xzf ], and therefore serves as a reliable predictor of model reconstruction error, indicating uncertainty. For instance, if the reconstruction mean Rϕ (x̂zf ) significantly deviates from the ground truth x, reflecting high uncertainty, the variance σϕ2 (x̂zf )i will increase accordingly to minimize (42). Simultaneously, the second term log σϕ2 (x̂zf ) prevents the variance from growing unboundedly.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
28
Additionally, these variances can be interpreted as automatic weights that dynamically adjust the importance of multi-task losses, as discussed in [55], where, in our context, the multiple tasks correspond to different sampling ratios. Similar variances also appear in the training of diffusion models, where they have been shown to enhance training dynamics [56]. 3.6.3
Relation to Existing Task-Adapted MRI Methods
We clarify the relationship between FSL, Tackle and our work. •
FSL [57]: Similar to SeqMRI, FSL enables sample-level adaptive sampling through its Mask-generating Sub-network Module. It dynamically re-computes a new sampling mask at each iteration of the reconstruction process, conditioned on the intermediate reconstruction. This allows the sampling pattern to adapt in real-time to individual patient characteristics.
•
Tackle [16]: To our best knowledge, Tackle does not provide real-time adaptive sampling during acquisition. However, it introduces ROI-oriented reconstruction and task-oriented optimization for the sampling mask, thereby providing task-level adaptive sampling that tailors the pattern to the specific diagnostic objective.
•
Proposed Method: InfoMRI introduces a distinct and complementary approach. Its primary novelty lies in its information-theoretic foundation, which provides a unified principle for jointly optimizing sampling, reconstruction, and task inference under a single objective: maximizing task-relevant information within a measurement budget. This information-centric principle is orthogonal to existing adaptive sampling strategies and could potentially be combined with them. For instance, sample-level adaptation for InfoMRI could be enabled by leveraging approaches like deep adaptive design [29] or reinforcement learning [58].
Different from FSL and Tackle that offer adaptive sampling strategies at sample and task levels, we provide a new theoretical lens for task-adapted MRI by optimizing information flow throughout the entire imaging-to-diagnosis pipeline.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
4
29
E XPERIMENTS
In this section, we evaluate the proposed methods on various MRI-related tasks. For convenience, we refer to our method as InfoMRI throughout the experiments. To validate the efficacy of InfoMRI, we compare with (1) CS-MRI reconstruction methods, (2) task-adapted CS-MRI methods, and (3) compressed learning methods. Note that all experimental data are simulated using a single-coil forward model. We select widely adopted segmentation and classification tasks [7], [16], [57]. The theoretical generality can be applied beyond segmentation and classification, with clear potential for broader applicability. Experimental results are able to: 1) Highlight the advantages of amortized optimization proposed in Section 3.3, including flexible control of sampling ratios, superior performance, and training efficiency. 2) Demonstrate the robust and reliable prediction and uncertainty quantification by the proposed method to address the uncertainty problem discussed in Section 2.3. 3) Validate the effectiveness of marginal entropy minimization proposed in Section 3.5 for optimizing a sampling pattern that prevents recovering the original image for privacy protection. 4.1
Comparisons to CS-MRI Reconstruction Methods
In this section, we implement pure CS-MRI reconstruction corresponding to β → −∞, to highlight the adaptive sampling via the proposed amortized optimization, demonstrating the flexible sampling ratio control and training efficiency compared to current end-to-end methods, as well as superior performance relative to existing active acquisition methods. The training and inference steps are illustrated in Algorithm 2 and 3, respectively, except for the steps related to task inference. 4.1.1
Experimental Setup
We compare state-of-the-art deep learning-based CS-MRI methods, including the end-toend approaches LOUPE [6] and SeqMRI [8], the active acquisition method PGMRI [11], as well as traditional sampling strategies such as uniform random [60], variable density [61], equispaced sampling [59], and spectrum-based method [62]. The goal of this
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
PSNR (2D Unconstrained)
40.0
30
SSIM (2D Unconstrained)
0.95
37.5 0.90
35.0
SSIM
PSNR
32.5 30.0 27.5 25.0 22.5 20.0 0.00 36.0
0.05
0.80
InfoMRI (Ours) Uniform Random Variable Density Spectrum-based SeqMRI LOUPE 0.10 0.15 0.20 0.25 0.30 Sampling Ratio
0.75 0.70 0.00
PSNR (1D Cartesian)
0.95
34.0
SSIM
PSNR
28.0
22.0
SSIM (1D Cartesian)
0.85
30.0
24.0
0.05
InfoMRI (Ours) Uniform Random Variable Density Spectrum-based SeqMRI LOUPE 0.10 0.15 0.20 0.25 0.30 Sampling Ratio
0.90
32.0
26.0
0.85
InfoMRI (Ours) Uniform Random Equispaced SeqMRI LOUPE PGMRI 0.075 0.100 0.125 0.150 0.175 0.200 0.225 0.250 Sampling Ratio
0.80 0.75 0.70 0.65
InfoMRI (Ours) Uniform Random Equispaced SeqMRI LOUPE PGMRI 0.075 0.100 0.125 0.150 0.175 0.200 0.225 0.250 Sampling Ratio
Fig. 3: The rate-distortion curves of various methods across different sampling ratios in terms of PSNR and SSIM on the fastMRI database [59]. The curves for LOUPE and SeqMRI are plotted at four fixed sampling ratios of 6.25%, 12.5%, 18.75%, and 25%. In contrast, the curves for other methods are plotted using 64 points sampled equidistantly from the sampling rate range of [1%, 30%] for 2D unconstrained sampling and [6.25%, 25%] for 1D Cartesian sampling.
comparison is to demonstrate the superiority of the proposed adaptive sampling method. For a fair comparison, all methods are paired with U-Net [63] for reconstruction. Specifically, the end-to-end methods jointly optimize their sampling modules with the reconstruction U-Net, while the active acquisition method and traditional sampling strategies
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
31
use U-Net to learn the projection from measurements to the target MR images. The proposed InfoMRI employs the same U-Net architecture to outputs the mean and variance of q(x|y). We utilize the pre-trained models of LOUPE and SeqMRI from the open-source repository (https://github.com/tianweiy/SeqMRI), train PGMRI from the open-source repository (https://github.com/Timsey/pg mri), and train the other models from scratch. For the traditional sampling strategies and our PGN, the sampling ratio r is randomly generated from U[0, 0.3] during the training stage. We employ the Adam optimizer with a learning rate of 1 × 10−4 and a weight decay of 1 × 10−4 to optimize the reconstruction U-Nets for the proposed InfoMRI and traditional sampling strategies. For the proposed PGN, we use the Adam optimizer with a learning rate of 1 × 10−2 and no weight decay. The models are trained with a batch size of 16 for 50 epochs. We benchmark the reconstruction performance of all methods on the NYU fastMRI database [59], a widely-used large-scale public dataset for CS-MRI reconstruction research. For the data preprocessing pipeline, we follow the implementation described in [10]. Specifically, we restrict the dataset to single-coil scans, and the slices are cropped to the central 128×128 region in k-space for computational efficiency. Note that cropping the rectangular k-space to a square matrix will result in anisotropic spatial resolution. However, this preprocessing step is a common practice adopted from prior works [6], [8], [11] to standardize input dimensions for network architectures. To provide a fair and controlled environment for comparison, this preprocessing step was applied uniformly to all methods evaluated in our study, including all baselines. Each 2D slice was treated as an independent sample. The proposed framework can be readily extended to 3D MR acquisitions by adapting the network architectures to utilize 3D convolutions and modifying the PGN to output a 3D k-space mask. Critically, the core informationtheoretic principles and the amortized optimization strategy remain identical. The original validation set is further split into a new validation set and a test set, using the same splitting setup as in [10]. The standard deviation σ of the measurement noise is set to 5 × 10−5 . Peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are adopted as performance metrics. For our method, we use the mean of q(x|y)
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
32
as the MRI reconstruction image for calculating the metric values and visualization. Following [6], [8], in addition to the acquisition budget constraint, we conduct further experiments that take into account the physical constraints of real-world MRI scanners. Specifically, the sampling positions are constrained to lie along lines in k-space that are orthogonal to the provided read-out direction, a scheme commonly referred to as 1D Cartesian sampling. In contrast, sampling without this line constraint is referred to as 2D unconstrained sampling in the following sections of the paper. Note that variable density [61] only support 2D unconstrained sampling, while equispaced sampling [59] and PGMRI [11] only support 1D Cartesian sampling. Uniform random [60], LOUPE [6], SeqMRI [8], and our InfoMRI support both scenarios. 4.1.2
Reconstruction Performance
Fig. 3 illustrates the performance of various methods across different sampling ratios through rate-distortion curves. Note that the end-to-end methods LOUPE and SeqMRI do not support flexible sampling ratio control and are thus plotted at four fixed sampling ratios: 6.25%, 12.5%, 18.75%, and 25%. Specifically, LOUPE is trained separately for each ratio, and SeqMRI can only achieve these four sampling ratios using a single trained model. Fig. 4 further visualizes the MRI reconstruction images and the corresponding sampling patterns at the 25% sampling ratio. In the 2D unconstrained sampling scenario, our method achieves comparable performance to LOUPE, and outperforms SeqMRI except at the 25% sampling rate in terms of PSNR. In the 1D Cartesian sampling scenario, our method is slightly inferior to LOUPE at the 6.25% sampling rate. In all other cases, our method demonstrates superior performance over all baselines. This can be attributed to the fact that LOUPE is trained separately for each ratio, and SeqMRI is specifically optimized for the 25% sampling rate. Notably, SeqMRI exhibits a significant drop in reconstruction performance at its first three sampling rate points, as these represent intermediate reconstruction stages for the 25% sampling rate and are not specifically optimized. In contrast, our method supports flexible sampling ratio control via a single trained model. PGMRI, the active acquisition method that also supports flexible sampling ratio control, shows a
1D Cartesian
2D Unconstrained
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
33
Uniform Random
Variable Density
Spectrum
LOUPE
SeqMRI
InfoMRI (Ours)
PSNR: 31.60dB SSIM: 0.9232
PSNR: 34.60dB SSIM: 0.9484
PSNR: 34.84dB SSIM: 0.9634
PSNR: 36.86dB SSIM: 0.9631
PSNR:37.96dB SSIM: 0.9714
PSNR: 36.94dB SSIM: 0.9656
Uniform Random
Equispace
PG-MRI
LOUPE
SeqMRI
InfoMRI (Ours)
PSNR: 28.99dB SSIM: 0.8716
PSNR: 28.99dB SSIM: 0.9013
PSNR: 30.81dB SSIM: 0.8953
PSNR: 30.52dB SSIM: 0.8787
PSNR:28.99dB SSIM: 0.8716
PSNR: 34.13dB SSIM: 0.9353
Ground-Truth
Ground-Truth
Fig. 4: Visualizations of MR reconstruction images and the corresponding sampling patterns at the 25% sampling rate on the fastMRI database. The rightmost column shows the ground truth MR images (top), the amplified regions (middle), and the kspace measurement (bottom). The other columns display the reconstruction images, the amplified areas, and the generated sampling patterns from top to bottom. The PSNR and SSIM values are labeled in the top right corner of the reconstruction images.
substantial decrease in reconstruction performance across all sampling ratios compared to our method. This is because PGMRI solely optimizes the sampling strategy based on a pre-trained MRI reconstruction model, which does not enable joint optimization of sampling and reconstruction.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
4.1.3
34
Performance with Advanced Reconstruction Methods
To demonstrate the modularity and broad applicability of our framework, we evaluate its performance when integrated with various state-of-the-art reconstruction architectures, including Plug-and-Play (PnP) algorithms and deep unrolling networks. The experiments were conducted on the Brain MRI dataset from ADMM-Net [64]. Our framework was integrated with three distinct reconstruction models: •
InfoMRI-PnP: This model utilizes a pre-trained DRUNet [65] as the denoiser prior within a PnP framework, allowing for a direct comparison with PnP-Parallel-MRI [66].
•
InfoMRI-Unroll: This model adopts the architecture of an unrolling network, specifically ISTA-Net+ [2], as its reconstruction backbone.
•
InfoMRI-UNet: This model serves as a baseline using a standard U-Net architecture, enabling comparison with other U-Net-based sampling methods like LOUPE [6].
As shown in Table 1, replacing U-Net with an unrolled backbone yields higher overall reconstruction quality (InfoMRI-Unroll vs. InfoMRI-UNet), as expected. Crucially, the trends reported in the main tables remain unchanged when switching to the unrolled backbone: (i) learned sampling continues to outperform fixed patterns (e.g., InfoMRI-Unroll v.s. ISTA-Net+); and (ii) amortizing training over a range of sampling ratios matches the performance of a single–sampling-ratio model. InfoMRI-Unroll closely matches PUERT in reconstruction fidelity across accelerations while supporting variable sampling ratios within a single trained model. In addition, InfoMRI-UnrollMSE underperforms InfoMRI-Unroll, highlighting the benefit of our uncertainty-aware information objective, which adaptively weights losses across sampling ratios, reduces gradient variance, and improves final performance. 4.1.4
Effectiveness of Amortized Optimization
Fig. 5 compares models trained with r ∼ p(r) = U[0, 0.3] and models individually trained at specific sampling ratios (i.e., with p(r) set as a delta distribution at a specific sampling ratio r0 ) for the 2D unconstrained sampling scenario. The results indicate that models trained with r ∼ U[0, 0.3] achieve nearly identical rate-distortion performance compared to independently trained models, demonstrating the effectiveness of the proposed
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
35
TABLE 1: Results on the Brain MRI dataset from ADMM-Net. InfoMRI-Unroll-MSE means remove the uncertainty weight and train the model solely on MSE.
Method
Sampling Pattern
100 × Acceleration 15
100 × Acceleration 10
100 × Acceleration 5
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
MD-Recon-Net [67]
Radial
37.25
0.9360
35.10
0.9120
31.05
0.8250
PnP-Parallel-MRI [66]
Radial
37.33
0.9388
35.14
0.9111
31.01
0.8221
ADMM-Net [64]
Radial
36.83
0.9306
34.46
0.8972
30.13
0.7958
Self-Supervised MRI [68]
Radial
36.58
0.9218
33.42
0.8924
30.18
0.8068
ISTA-Net+ [2]
Radial
37.07
0.9343
34.73
0.9052
30.64
0.8176
LOUPE [6]
Learned
38.30
0.9473
36.88
0.9381
35.34
0.9245
PUERT [9]
Learned
41.01
0.9655
39.23
0.9565
36.65
0.9399
InfoMRI-PnP
Learned
40.08
0.9529
38.54
0.9441
35.66
0.9215
InfoMRI-UNet
Learned
38.35
0.9480
36.93
0.9391
35.25
0.9232
InfoMRI-Unroll
Learned
40.96
0.9596
39.22
0.9591
36.61
0.9393
InfoMRI-Unroll-MSE
Learned
40.32
0.9590
38.61
0.9492
35.91
0.9293
amortized optimization approach in achieving adaptive sampling and reducing training complexity. Due to the conjugate symmetry property of k-space [53], sampling both a position and its point-reflected counterpart (rotated 180 degrees for 2D images) introduces redundancy (see Section 3.5). We have introduced a quantitative Redundancy Ratio metric based on the conjugate symmetry properties of MRI k-space. We define the redundancy ratio of a sampling mask m as: |I| Redundancy Ratio = |I|+|J , |
(43)
where I and J are defined in Section 3.5. In Table 2, we quantitatively report the redundancy ratio of different sampling patterns, and in Fig. 6, we visualize the sampling patterns produced by our PGN and traditional sampling strategies. Notably, even without manually imposed constraints, PGN adaptively learns to exploit this inherent redundancy in k-space, thereby improving acquisition efficiency, demonstrating that our learned PGN sampling pattern is theoreti-
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
40.0 38.0
PSNR (2D Unconstrained)
0.96
p(r) = (r r0) Unif(0, 0.3)
0.94
34.0 32.0
28.0 0.0
0.90
0.1 0.2 Sampling Ratio
0.3 0.86 0.0
PSNR (1D Cartesian)
0.1 0.2 Sampling Ratio
0.3
SSIM (1D Cartesian)
p(r) = (r r0) Unif(0, 0.3)
0.94
p(r) = (r r0) Unif(0, 0.3)
0.92
34.0
0.90
32.0
SSIM
PSNR
p(r) = (r r0) Unif(0, 0.3)
0.88
30.0
36.0
SSIM (2D Unconstrained)
0.92 SSIM
PSNR
36.0
36
30.0
0.88 0.86 0.84
28.0 0.10
0.15 0.20 Sampling Ratio
0.25
0.10
0.15 0.20 Sampling Ratio
0.25
Fig. 5: Reconstruction performance comparison of models trained with p(r) = U[0, 0.3] versus models trained individually at specific sampling ratios with p(r) = δ(r − r0 ). The performance is evaluated on the fastMRI database.
cally sound and more efficient at acquiring non-redundant spatial frequencies compared to baseline methods.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
𝑟 = 0.10
𝑟 = 0.15
𝑟 = 0.20
𝑟 = 0.25
𝑟 = 0.30
InfoMRI (Ours)
Spectrumbased
Variable Density
Uniform Random
𝑟 = 0.05
37
Fig. 6: Sampling patterns produced by traditional sampling strategies and our PGN as r increases from 0.05 to 0.3. TABLE 2: Redundancy ratios for the sampling masks shown in Fig. 6. Lower values denote higher sampling efficiency. Methods
4.1.5
r = 0.05
r = 0.15
r = 0.25
Uniform random
0.11
0.20
0.28
Variable density
0.51
0.52
0.51
Spectrum
1.0
1.0
1.0
InfoMRI (Ours)
0.07
0.09
0.14
Reconstruction Uncertainty Quantification
The variance of the posterior approximation q(x|y) is calculated for estimating the reconstruction error and evaluating the reconstruction uncertainty, as elaborated in
38
InfoMRI-UNet
InfoMRI-unroll
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
Fig. 7: Visualization of InfoMRI at r = 0.15 for uncertainty quantification on the Brain MRI dataset. InfoMRI-Unroll reconstructs better than InfoMRI-UNet and both variants accurately predict the reconstruction error with learned uncertainty.
Section 3.6. Fig. 7 shows that the pixel-wise variances of q(x|y) highlight regions and textures that closely resemble the ground truth reconstruction error, demonstrating the high accuracy of the uncertainty quantification provided by InfoMRI. This can assist medical professionals in identifying areas where the reconstruction is less clear, ultimately aiding in medical diagnosis. 4.2
Comparisons to Task-adapted CS-MRI Methods
In this section, we implement segmentation-adapted CS-MRI for the case where β < 0, but not approaching −∞, highlighting both the segmentation performance and the
simultaneous prediction of segmentation uncertainty. 4.2.1
Experimental Setup
We compare our method with state-of-the-art task-adapted CS-MRI methods for segmentation using undersampled k-space measurement, including LI-Net [15], Semu-
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
Overlayed Sampling Mask
Sample 0
Sample 1
Sample 2
39
Sample 3
Sample 4
Sample 5
Pixel-wise Prob.
LI-Net
InfoMRI (Ours)
GT
Fig. 8: Comparison of segmentation predictions on QUBIQ 2021 dataset. Pixel-wise probability methods exhibit spatially incoherent segmentation maps. LI-Net produces spatially coherent segmentation map but no variation. InfoMRI successfully recovers the spatial covariance of the segmentations, and produces diverse and spatially coherent samples.
Net [7], and Tackle [16]. Following [16], we also compare with LOUPE-UNet-RS, TradPatUNet-RS, and LOUPE-LI-Net. We categorize the baselines into two groups. •
Implementing segmentation after MRI reconstruction. This category includes SemuNet,
LOUPE-UNet-RS, and TradPat-UNet-RS. In LOUPE-UNet-RS, the model first generates sampling patterns following LOUPE to produce measurements, then performs MRI reconstruction using a U-Net, and finally applies segmentation with another U-Net. TradPat-UNet-RS follows the same configuration as LOUPE-UNet-RS, except that the sampling patterns are generated using traditional sampling strategies, including uniform random [60], variable density [61], and Poisson disk methods [16]. •
Implementing segmentation directly based on measurements. This category includes LI-Net,
Tackle, LOUPE-LI-Net, and our method. LOUPE-LI-Net replaces the sampling pattern of LI-Net with a learned sampling pattern from LOUPE.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
40
TABLE 3: Segmentation performance comparisons on QUBIQ 2021 dataset. Lower GED and Higher DSC indicate better segmentation performance. We use bold and underline to highlight the best and the second best, respectively. Note that 8×, 16×, 24×, and 32× accelerations correspond to sampling ratios of 1/8, 1/16, 1/24 and 1/32, respectively. 8× Acceleration
16× Acceleration
24× Acceleration
32× Acceleration
GED ↓
DSC ↑
GED ↓
DSC ↑
GED ↓
DSC ↑
GED ↓
DSC ↑
Uniform random
0.2396
0.9211
0.4265
0.8566
0.7296
0.7275
0.7685
0.7039
Poisson disk
0.2319
0.9147
0.3124
0.8903
0.4570
0.8343
0.5226
0.8129
Variable density
0.1573
0.9474
0.2324
0.9171
0.2801
0.9047
0.3008
0.8979
Uniform random
0.2516
0.9104
0.3601
0.8737
0.4750
0.8311
0.4872
0.8263
Poisson disk
0.2089
0.9213
0.2443
0.9109
0.3732
0.8707
0.3871
0.8590
Variable density
0.1318
0.9544
0.1585
0.9466
0.2458
0.9150
0.2161
0.9260
LOUPE-UNet-RS
0.1404
0.9516
0.2145
0.9251
0.3054
0.8802
0.3458
0.8696
LOUPE-LI-Net
0.1215
0.9559
0.1368
0.9503
0.1643
0.9428
0.2306
0.9210
0.1166
0.9606
0.1296
0.9556
0.2244
0.9223
0.2549
0.9131
Tackle [16]
0.1009
0.9644
0.1243
0.9570
0.1683
0.9404
0.2216
0.9243
InfoMRI (Ours)
0.0987
0.9496
0.1009
0.9474
0.1086
0.9441
0.1142
0.9429
Method
TradPat-UNet-RS
LI-Net [15]
SemuNet [7]
Sampling Patterns
Learned
The proposed InfoMRI uses the encoder and decoder architecture from latent diffusion models [69] to construct the encoders EϕY , Eϕ̃Y and decoders DϕZ , DϕT , DϕX . We employ the same optimizer, learning rate, and weight decay in Section 4.1.1 to train all the models for 20,000 training steps. Following [70], we adjust the weight parameters in (20) to balance the reconstruction term and KL divergence term. The training objective is modified as max Er,M,m,x,t,y Eq̂(z|t,y) [w1 log q(t|z, y) − w2 DKL [q̂(z|t, y)||q(z|y)] + w3 log q(x|y)], θ,ϕ
(44)
where w1 , w2 , w3 are set to 1, 50, and 1 in our experiments, and ϕ := (ϕY , ϕ̃Y , ϕZ , ϕT , ϕX ). Following [18], [38], [39], we use the Dice Similarity Coefficient (DSC) to evaluate segmentation performance and the Generalized Energy Distance (GED) to measure the similarity between distributions. GED is defined as 2 DGED (q, p) = 2E[dIoU (t̂, t)]− E[dIoU (t̂, t̂′ )]− E[dIoU (t, t′ )],
(45)
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
41
Fig. 9: Visualization of posterior samples from q(t|y) under different acceleration factors for our InfoMRI on QUBIQ 2021 dataset. As the acceleration factor increases, q(t|y) generates more diverse samples to address the increasing uncertainty caused by the reduced amount of information.
where dIoU = 1 − IoU, t̂, t̂′ are independently drawn from q , and t, t′ are independently drawn from p. Here, q and p represent the predicted and ground truth distributions over segmentations, respectively. For deterministic methods, q is set to a delta peak at the predicted segmentation. In our experiments, we draw 32 samples from the predicted distribution q . We use the prostate segmentation dataset from the QUBIQ 2021 challenge [71], which benchmarks methods for addressing uncertainty in medical image segmentation. This dataset consists of MRI images, making it more appropriate for evaluating the proposed InfoMRI method, which is specifically designed for MRI, compared to datasets from other medical imaging modalities (e.g., CT datasets used in [39], [72]). Each prostate image in this dataset includes six labels provided by domain experts, making it ideal for assessing the proposed InfoMRI model’s ability to address aleatoric uncertainty by
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
42
modeling the distribution of segmentations. The standard deviation of the measurement noise σ is set to 5 × 10−2 . 4.2.2
Quantitative Results
Table 3 summarizes the segmentation performance of different methods. It is evident that LI-Net demonstrates an increasing advantage over U-Net for MRI segmentation when using the same sampling patterns. This is observed in the comparisons between LI-Net and TradPat-UNet-RS (both utilizing traditional sampling patterns) and between LOUPE-LI-Net and LOUPE-UNet-RS (both utilizing learned sampling patterns). This advantage can be attributed to LI-Net’s ability to enforce spatial coherence in segmentation, which significantly reduces the solution space and enables correct segmentation to be inferred from fewer measurements. Additionally, Tackle outperforms SemuNet by a noticeable margin. Both methods are based on sequential models that combine a U-Net for reconstruction with a UNet for segmentation from the reconstructed image. The primary distinction lies in their approach to the reconstruction stage: Tackle treats the reconstructed image as an intermediate feature without explicitly optimizing its reconstruction quality, whereas SemuNet optimizes the reconstructed image using the L1 distance to the ground truth. This result suggests that deliberately optimizing MRI reconstruction may not necessarily lead to improved performance in MRI segmentation. For the proposed InfoMRI, the GED performance consistently surpasses all baselines across all acceleration factors, demonstrating a superior match to the ground-truth distribution compared to the baselines. Moreover, InfoMRI exhibits an increasing advantage over the baselines and experiences the smallest performance drop as the acceleration factor increases (in both GED and DSC metrics). This indicates that the proposed method effectively handles information loss in MR subsampling, producing more reliable and stable predictions. 4.2.3
Segmentation Under Uncertainty
Due to the information lost in MR subsampling, there is a considerable uncertainty while implementing segmentation. Fig. 8 demonstrates how different methods handle
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
43
segmentation under an extremely low sampling rate (32× acceleration). As shown, pixelwise prob. methods (abbreviated as methods that model segmentation using pixel-wise probabilities) result in spatially incoherent segmentation when thresholding the pixelwise probabilities. For LI-Net, segmentation performance is significantly better than pixel-wise probability methods due to its approach of decoding segmentation from latent representations, resulting in spatially consistent segmentation. However, segmentation errors remain non-negligible, and LI-Net is unable to provide uncertainty quantification due to its deterministic design. In contrast, the proposed method generates multiple segmentations (see the areas of varying shades of red color) by decoding latent representations from multiple samples of the latent posterior q(z|y), thereby achieving spatially consistent segmentation while also providing uncertainty quantification. Fig. 9 further demonstrates that as the acceleration factor increases, the proposed method produces more diverse segmentations, indicating its ability to provide reasonable uncertainty quantification to address the information loss in MR subsampling. 4.3
Privacy-Preserving Task-Adapted CS-MRI
In this section, we implement privacy-preserving classification-adapted CS-MRI for the case where β ≥ 0, with larger values of β indicating stronger suppression of MRI reconstruction. Many clinical workflows aim to obtain specific diagnostic decisions (e.g., classification of conditions) rather than visually inspectable images. However, MR images inherently contain sensitive, patient-identifiable information. Sharing or processing these images, even for automated analysis, poses significant privacy risks, especially in the context of federated learning or cloud-based diagnostics. Our framework is able to learns a sampling and inference pipeline that is explicitly optimized to be sufficient for the diagnostic task T while being deliberately insufficient for high-fidelity reconstruction of the image X . This creates a form of “privacy-by-design”, where the acquired data is intrinsically difficult to reverse-engineer into a recognizable image, thus protecting patient privacy.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
4.3.1
44
Experimental Setup
The proposed InfoMRI is compared with traditional sampling strategies, including uniform random and variable density, as well as the deep learning-based approach, LOUPE and A-DPS. We utilize ResNet18 [73] as the classification backbone for all traditional sampling methods and LOUPE. All ResNet18 models are trained using the Adam optimizer with a learning rate of 1 × 10−4 , while our PGN is still trained using the Adam optimizer with a learning rate of 1 × 10−2 . We use publicly available code of A-DPS with modifications to enable k-space sampling for fair comparisons. To initially validate our method, we employ the MNIST dataset. This dataset is partitioned into training (50,000 images), validation (5,000 images), and test (5,000 images) subsets. Subsequently, we evaluate our approach in a more clinically relevant setting using the fastMRI knee dataset with slice-level labels provided by [74]. We split the dataset into training (816 volumes), validation (176 volumes), and test (175 volumes) sets. Following [75], we predict the presence of Meniscal Tears and ACL sprains in each slice, introducing an additional ”abnormal” category to encompass less frequent pathologies as described in [75]. To compare task performance, we use accuracy for the MNIST dataset and the area under the receiver operator curve (AUROC) for the fastMRI dataset, given the imbalanced nature of the dataset. For MNIST and fastMRI dataset, we trained for 150 epochs with batch size of 512 and 16, respectively. To assess the effectiveness of privacy protection, we evaluate the performance of reconstructing MR images from the k-space measurements. A lower PSNR implies a lower likelihood of successful reconstruction and indicates better privacy protection. For the fastMRI dataset, we adopt the same Reconstruction U-Net and training configurations described in Section 4.1. For the MNIST dataset, we construct a simple convolutional neural network (CNN) for reconstruction. The architecture of the CNN is Conv2d(2,32,3,3)-ReLU-Conv2d(32,64,3,3)-ReLU-Conv2d(64,1,1,1), where Conv2d(2,32,3,3) denotes a convolutional layer with an input channel size of 2, an output channel size of 32, and a kernel size of 3×3. The CNN is trained to minimize the MSE loss with a batch size of 64 using the Adam optimizer with a learning rate of
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
45
TABLE 4: Privacy preserving classification on MNIST. Higher accuracy and lower PSNR indicate better classification and privacy preserving performance. We use bold to highlight the best and underline for the second best. Methods
32× Acceleration
48× Acceleration
PSNR ↓
Accuracy ↑
PSNR ↓
Accuracy ↑
Uniform random
11.33
0.8236
11.20
0.8061
Variable density
14.05
0.9410
12.74
0.8838
LOUPE
15.61
0.9692
13.73
0.9253
A-DPS
15.10
0.9773
13.02
0.9651
InfoMRI (β = 0)
14.84
0.9776
12.19
0.9587
InfoMRI (β > 0)
11.79
0.9801
10.25
0.9622
TABLE 5: Privacy preserving diagnosis on fastMRI. Higher AUROC and lower PSNR indicate better classification and privacy preserving performance. We use bold and underline to highlight the best and the second best, respectively.
Methods
16× Acceleration
32× Acceleration
PSNR ↓
AUROC ↑
PSNR ↓
AUROC ↑
Uniform random
22.58
0.8968
19.63
0.8848
Variable density
25.33
0.9079
23.81
0.8989
LOUPE
26.20
0.9129
24.85
0.9064
InfoMRI (β = 0)
25.42
0.9206
24.00
0.9157
InfoMRI (β > 0)
24.41
0.9152
23.56
0.9096
1 × 10−4 . Training is terminated when the MSE on the validation set stops decreasing.
4.3.2
Results
Tables 4 and 5 summarize the reconstruction and task performance for all the methods. The optimal value of β > 0 is selected via grid search. An interesting observation is that task performance and reconstruction performance may not be directly correlated, i.e., better reconstruction performance not necessarily leading to better task performance. For instance, Table 4 shows that InfoMRI (β > 0) achieves significantly higher accuracy
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
46
than variable density and LOUPE (by approximately 4% and 2%, respectively) while simultaneously exhibiting worse reconstruction performance (by around 2.3 dB and 3.8 dB, respectively). 4.4
Discussion and Limitation
We provide additional experimental details and results in the supplementary material, including additional reconstruction visualizations on brain MRI data, extended experiments on the SKM-TEA and BRISC 2025 datasets with qualitative analysis of finestructure tumor segmentation, confidence calibration analysis, sensitivity analysis for hyperparameters, inference complexity analysis, training stability analysis, and performance degradation analysis across acceleration factors. While the proposed InfoMRI framework demonstrates promising performance in task-adapted sampling and reconstruction, there exist several limitations in the current study. First, the experiments are primarily conducted on single-coil simulated data. Clinical MRI scanners typically use multi-coil acquisition (parallel imaging), which involves more complex forward models (e.g., sensitivity encoding) and noise characteristics. Extending the proposed information-theoretic optimization to multi-coil settings remains an important future direction. Second, our current methodology focuses on 2D slice-wise acquisition, whereas many clinical protocols employ 3D volumetric imaging. The computational complexity of posterior sampling and entropy estimation would increase significantly in 3D, requiring more efficient approximation techniques. Finally, although we demonstrate the utility of uncertainty maps for flagging difficult cases, integrating this “human-in-the-loop” mechanism into actual clinical workflows requires further validation with radiologists to assess its impact on diagnostic decision-making.
5
C ONCLUSIONS
In this paper, we proposed the first adaptive task-adapted CS-MRI framework via ITMbased optimization that addresses the limitations of existing methods in supporting specific clinical tasks and managing diagnostic uncertainty. By maximizing the mutual
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
47
information between k-space measurements and downstream tasks, our approach enables probabilistic inference and provides a comprehensive solution to the inherent uncertainty in medical diagnosis. Leveraging amortized optimization and constructing tractable variational bounds, we achieved adaptive CS-MRI that jointly optimizes sampling, reconstruction, and task-inference models, allowing flexible control of the sampling ratio within a single end-to-end trained model. Furthermore, our framework unifies two distinct clinical scenarios: enhancing task performance through joint reconstruction and task inference, and implementing tasks with suppressed reconstruction for privacy protection. Extensive experimental results demonstrate the state-of-the-art performance of the proposed framework in terms of reconstruction quality and the accuracy of various clinical tasks.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
48
A PPENDIX A A DDITIONAL E XPERIMENTAL R ESULTS A.1
Additional Results on CS-MRI Reconstruction
To evaluate our method on complex, fine-structured anatomical regions, we show visual reconstruction results on brain MRI data in Figure 10. We further evaluate the Bjøntegaard Delta (BD)-PSNR and BD-Rate of the baselines in Table 6, using the proposed method as the benchmark. Table 6 shows that our method outperforms all baselines except for LOUPE in the 2D unconstrained sampling scenario on the fastMRI dataset. A.2
Additional Results on Task-adapted CS-MRI
A.2.1
More Results
As presented in Table 7, at the 2× and 4× acceleration rates, InfoMRI achieves the lowest GED (0.0940 and 0.0943, respectively). This indicates that its predictions provide the closest distributional match to the ground-truth segmentation posterior among all evaluated methods. To further demonstrate the clinical practicality of the proposed method, we have further evaluated the performance on •
SKM-TEA dataset [76]: A large-scale benchmark of quantitative knee MRI (qMRI) scans with dense manual segmentations, designed for the joint evaluation of MRI reconstruction and analysis tasks.
•
BRISC 2025 dataset [77]: A clinical benchmark providing expert-annotated brain tumor segmentations from realistic MRI scans, targeting robust segmentation performance in clinical settings.
The quantitative results are presented in Table 8 and Table 9, respectively. The proposed method consistently outperforms prior methods on GED metrics and obtains the best DSC scores for high acceleration factors. Figures 11 and 12 visualize our fine-structure segmentation capabilities. Unlike standard deterministic models that yield overconfident and often incorrect pixel-wise prob-
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
49
TABLE 6: BD-PSNR gains (dB) and BD-Rate changes (%) for the baselines compared with InfoMRI on the fastMRI dataset. The results of LOUPE and SeqMRI are averaged over the four sampling ratios of 6.25%, 12.5%, 18.75%, and 25%, whereas those of other methods are calculated using all sampling ratios within the range [6.25%, 25%] (note that there are a total of 100 equidistantly sampled points within the sampling rate range of [1%, 30%]). A negative BD-PSNR indicates that the baseline provides lower quality than InfoMRI at the same sampling rate, while a positive BD-Rate means the baseline requires higher sampling rate than InfoMRI to achieve the same quality. Method
2D Unconstrained
1D Cartesian
BD-PSNR
BD-Rate
BD-PSNR
BD-Rate
Uniform random
-4.22
38.97%
-3.96
32.64%
Variable density
-1.72
26.45%
-
-
Spectrum based
-2.10
30.17%
-
-
Equispace
-
-
-3.06
26.06%
LOUPE
0.03
-0.017%
-0.89
14.22%
SeqMRI
-1.48
17.66%
-3.06
25.61%
PGMRI
-
-
-4.20
28.96%
ability maps under high uncertainty, InfoMRI generates diverse, structurally coherent posterior samples. This enables a highly practical clinical workflow: cases with high sample agreement (Fig. 11, Row 1) indicate reliable segmentations that can bypass manual review; conversely, cases with high inter-sample disagreement (Fig. 11, Row 2) serve as a built-in uncertainty flag, successfully highlighting ambiguous fine-structure boundaries for human expert (radiologist) review to select the correct sample. In contrast, deterministic models rely on pixel-wise probability maps that yield varying segmentation boundaries depending on the chosen threshold. As demonstrated in Row 3 of Fig. 11, under high ambiguity, all such threshold-dependent results remain fundamentally incorrect. Furthermore, Figure 12 demonstrates that our method robustly captures fine, irregular tumor structures across diverse patient subjects. A.2.2
Confidence Calibration
To evaluate whether the methods provide accurate uncertainty estimation, we rely on confidence calibration metrics, Expected Calibration Error (ECE) and Brier Score (BS), for
50
Sampling Ratio 𝒓 = 𝟎. 𝟏𝟓
Sampling Ratio 𝒓 = 𝟎. 𝟎𝟓
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
Fig. 10: Visualizations of MR reconstruction images at the 5% and 15% sampling rates on the Brain MRI dataset. The leftmost column shows the ground truth MR images (top) and the amplified regions (bottom). The other columns display the reconstruction images and the amplified areas. The PSNR and SSIM values are labeled in the reconstruction images.
pixel-wise probabilities. Specifically, for a binary segmentation task, given the ground truth segmentation ti and predicted pixel-wise probabilities q(ti = 1) for each pixel
51
InfoMRI (Ours)
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
InfoMRI (Ours)
High agreement posterior samples, bypass human judger
Selected by human
Pixel-wise prob.
Low agreement posterior samples, request human judger
Thresholding (Thr) unstructured probabilities
All are incorrect segmentations
Fig. 11: Demonstration of practical utility evaluated on the challenging real-world clinical BRISC 2025 dataset. Row 1: A standard clinical case where InfoMRI posterior samples exhibit high consistency, indicating reliable segmentation. Rows 2 and 3: A shared challenging case featuring a small and ambiguous lesion (sharing the same Ground Truth). Row 2: Under high ambiguity, InfoMRI generates diverse, structurally coherent posterior samples. The high inter-sample variance explicitly indicates low confidence, serving as a reliable built-in flag to prompt human expert intervention. Row 3: In contrast, deterministic pixel-wise probability maps fail to capture this diagnostic uncertainty, yielding consistently incorrect segmentation boundaries regardless of the chosen threshold.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
52
TABLE 7: Segmentation performance comparisons on QUBIQ 2021 dataset. Lower GED and Higher DSC indicate better segmentation performance. We use bold and underline to highlight the best and the second best, respectively.
Method
Sampling Pattern
2× Acceleration 4× Acceleration
GED ↓ DSC ↑ GED ↓ Uniform random 0.1372 TradPat-UNet-RS Poisson disk Variable density
DSC ↑
0.9548
0.1666
0.9431
0.1190
0.9572
0.1266
0.9563
0.1300
0.9529
0.2029
0.9266
Uniform random 0.1688
0.9393
0.1868
0.9328
Poisson disk
0.1341
0.9520
0.1347
0.9520
Variable density
0.1411
0.9510
0.1704
0.9421
LOUPE-UNet-RS
0.1150
0.9609
0.1117
0.9618
LOUPE-LI-Net
0.1187
0.9568
0.1211
0.9561
0.1176
0.9565
0.1121
0.9598
Tackle
0.1104
0.9618
0.1073
0.9624
InfoMRI (Ours)
0.0940
0.9557
0.0943
0.9546
LI-Net
SemuNet
Learned
i = 1, . . . , d, the ECE and BS are defined as: ECE =
N X |Dn | n=1
d
|acc(Dn ) − conf(Dn )|,
(46)
d
1X BS = (q(ti = 1) − ti )2 , d i=1
(47)
where d is the dimension of segmentation map, N is the number of bins, Dn is the set of n pixel indices i whose predicted probabilities q(ti = 1) fall within the range n−1 , . The N N terms acc(Dn ) and conf(Dn ) represent the averaged segmentation accuracy and averaged predicted probabilities q(ti = 1) for pixels with indices in Dn , respectively. We set N = 16 in the experiments. For the proposed method, pixel-wise probabilities are obtained by averaging segmentation samples from q(t|y). Table 10 presents the confidence calibration performance of different methods. Notably, although the proposed method does not directly optimize for pixel-wise probabilities, the pixel-wise probabilities obtained by averaging samples from q(t|y) exhibit
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
53
TABLE 8: Segmentation performance comparisons on SKM-TEA dataset. Lower GED and Higher DSC indicate better segmentation performance. We use bold and underline to highlight the best and the second best, respectively. 2× Acceleration
4× Acceleration
8× Acceleration
16× Acceleration
GED ↓
DSC ↑
GED ↓
DSC ↑
GED ↓
DSC ↑
GED ↓
DSC ↑
Uniform random
0.3965
0.4935
0.4028
0.4887
0.4368
0.4742
0.5058
0.4416
Poisson disk
0.3835
0.4996
0.3859
0.4966
0.4126
0.4863
0.4503
0.4679
Variable density
0.3913
0.4944
0.3932
0.4927
0.4236
0.4803
0.4768
0.4571
Uniform random
0.5038
0.4529
0.5096
0.4482
0.5433
0.4331
0.6283
0.3927
Poisson disk
0.5008
0.4524
0.5074
0.4489
0.5115
0.4478
0.5429
0.4332
Variable density
0.5088
0.4496
0.5074
0.4499
0.5147
0.4464
0.5399
0.4343
LOUPE-LI-Net
0.4771
0.4632
0.4792
0.4620
0.4886
0.4573
0.5179
0.4448
LOUPE-UNet-RS
0.3903
0.4968
0.3887
0.4961
0.4136
0.4866
0.4544
0.4686
0.3712
0.5035
0.3721
0.5024
0.3930
0.4932
0.4330
0.4756
SemuNet
0.3741
0.5026
0.3749
0.5015
0.3943
0.4928
0.4345
0.4755
Tackle
0.3645
0.5048
0.3688
0.5033
0.3916
0.4936
0.4321
0.4757
InfoMRI (Ours)
0.3457
0.5013
0.3448
0.5011
0.3424
0.5032
0.3509
0.4949
Method
Sampling Pattern
TradPat-UNet-RS
LI-Net
FSL
Learned
comparable or even superior performance on metrics for confidence calibration. This demonstrates that the uncertainty quantification provided by the proposed method is not only qualitatively reasonable but also quantitatively accurate. A.2.3
Visualization of Privacy-Protected Compressed Learning
We visualize the privacy protection performance in Fig. 13. When the value of β increases, the reconstruction from measurement gradually becomes difficult, and thereby the privacy is better preserved. A.3 A.3.1
Ablation Study Sensitivity analysis for hyperparameters
We have conducted a sensitivity analysis for the key hyperparameters, particularly the weighting parameters w1 , w2 , w3 in our loss function (Eq. 48), to assess their impact on
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
54
TABLE 9: Segmentation performance comparisons on BRISC 2025 dataset. Lower GED and higher DSC indicate better performance. We use bold and underline to highlight the best and the second best, respectively. 8× Acceleration
16× Acceleration
24× Acceleration
32× Acceleration
GED ↓
DSC ↑
GED ↓
DSC ↑
GED ↓
DSC ↑
GED ↓
DSC ↑
Uniform random
0.9109
0.6423
1.1624
0.5132
1.4424
0.3612
1.5017
0.3263
Poisson disk
0.7598
0.7099
0.9390
0.6237
1.1742
0.5018
1.1890
0.4956
Variable density
0.6085
0.7816
0.7307
0.7299
0.8449
0.6770
0.9354
0.6300
Uniform random
1.0508
0.5793
1.2417
0.4783
1.4856
0.3409
1.5214
0.3204
Poisson disk
0.9039
0.6405
1.0580
0.5644
1.2976
0.4347
1.3174
0.4247
Variable density
0.7374
0.7284
0.8590
0.6656
0.9584
0.6164
1.0332
0.5785
LOUPE-LI-Net
0.6473
0.7735
0.7185
0.7389
0.7933
0.7053
0.9290
0.6333
LOUPE
0.5433
0.8091
0.6349
0.7733
0.7290
0.7319
0.8583
0.6680
0.5349
0.8141
0.6562
0.7638
0.7283
0.7336
0.7881
0.7071
Tackle
0.6024
0.7844
0.7303
0.7280
0.8298
0.6839
0.9103
0.6441
InfoMRI (Ours)
0.4993
0.7867
0.5254
0.7649
0.5507
0.7472
0.5677
0.7280
Method
TradPat-UNet-RS
LI-Net
SemuNet
Sampling Patterns
Learned
TABLE 10: Calibration performance comparisons on QUBIQ 2021 dataset. Lower BS and ECE indicate better calibration performance. We use bold and underline to highlight the best and the second best, respectively. Note that 8×, 16×, 24×, and 32× accelerations correspond to sampling ratios of 1/8, 1/16, 1/24 and 1/32, respectively. Method
Sampling patterns
8× Acceleration
16× Acceleration
24× Acceleration
32× Acceleration
ECE ↓
BS ↓
ECE ↓
BS ↓
ECE ↓
BS ↓
ECE ↓
BS ↓
TradPat-UNet-RS
Variable density
0.0121
0.0307
0.0204
0.0425
0.0191
0.0505
0.0318
0.0597
LI-Net
Variable density
0.0228
0.0266
0.0298
0.0337
0.0462
0.0499
0.0415
0.0459
LOUPE-UNet-RS
0.0083
0.0223
0.0129
0.0344
0.0156
0.0397
0.0364
0.0681
LOUPE-LI-Net
0.0299
0.0299
0.0319
0.0319
0.0407
0.0407
0.0532
0.0532
0.0141
0.0205
0.0178
0.0237
0.0155
0.0302
0.0223
0.0368
Tackle
0.0081
0.0176
0.0148
0.0255
0.0251
0.0309
0.0330
0.0421
InfoMRI (Ours)
0.0098
0.0275
0.0127
0.0290
0.0150
0.0310
0.0140
0.0316
SemuNet
Learned
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
GT
Pred Mean
Sample 0
Sample 1
55
Sample 2
Sample 3
Fig. 12: Visual examples of fine-structure tumor segmentation at 16× acceleration on the clinical BRISC 2025 dataset. The proposed method demonstrates a superior capability to accurately delineate intricate lesion boundaries and capture diverse, irregular tumor shapes, thereby maintaining strict structural integrity under significant undersampling.
final performance. max Er,M,m,x,t,y Eq̂(z|t,y) [w1 log q(t|z, y) − w2 DKL [q̂(z|t, y)||q(z|y)] + w3 log q(x|y)], θ,ϕ
(48)
Specifically, we varied w2 and w3 while keeping w1 fixed at 1.0. Varying w2 probes the balance between the KL regularizer DKL [q̂(z|t, y)||q(z|y)] and the task NLL, while w3
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
56
TABLE 11: Sensitivity analysis on hyperparameters of the loss function evaluated on QUBIQ 2021 dataset.
w1
w2
w3
4× Acceleration
8× Acceleration
16× Acceleration
24× Acceleration
GED ↓
DSC ↑
GED ↓
DSC ↑
GED ↓
DSC ↑
GED ↓
DSC ↑
1.0
1.0
1.0
0.7087
0.3466
0.6938
0.3053
0.6372
0.4541
0.6489
0.4356
1.0
5.0
1.0
0.5884
0.4633
0.7075
0.4463
0.6224
0.4880
0.5750
0.5480
1.0
10.0
1.0
0.2388
0.8673
0.2235
0.8826
0.2589
0.8674
0.2352
0.8855
1.0
50.0
1.0
0.0943
0.9546
0.0955
0.9542
0.1054
0.9472
0.1208
0.9380
1.0
100.0
1.0
0.0953
0.9558
0.0900
0.9574
0.1397
0.9199
0.2032
0.8981
1.0
1.0
0.0
0.0997
0.9540
0.0926
0.9557
0.1413
0.9327
0.1384
0.9391
1.0
1.0
0.1
0.0961
0.9567
0.0944
0.9561
0.1128
0.9480
0.1227
0.9448
1.0
1.0
0.5
0.0882
0.9565
0.0983
0.9457
0.0946
0.9529
0.1037
0.9534
1.0
1.0
1.0
0.0943
0.9546
0.0955
0.9542
0.1054
0.9472
0.1208
0.9380
1.0
1.0
10.0
0.1088
0.9507
0.1099
0.9490
0.1367
0.9375
0.1241
0.9437
Fig. 13: Visualization of reconstruction examples under different β values on MNIST. As observed, reconstruction becomes progressively more challenging as β increases. Top: Reconstruction; Bottom: sampling patterns.
controls the trade-off between reconstruction log q(x|y) and task performance. Results are summarized in Table 11. As shown in Table 11, a clear trend emerges as we vary w2 . Increasing w2 from 1 to 50 consistently improves both GED and DSC across all acceleration factors. However, when w2 is increased further to 100, performance at 16× and 24× acceleration exhibits a slight degradation, suggesting over-regularization. This indicates a robust operational
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
57
range for w2 is approximately 10 ∼ 50. With w2 fixed, we observe that a moderate reconstruction weight (w3 ∈ [0, 1]) maintains or slightly improves training stability and performance. Conversely, a large w3 = 10 significantly harms the task-specific metrics, as the optimization objective shifts focus toward pixel fidelity at the expense of the downstream task. This suggests a robust range for w3 is 0.1 ∼ 1.0. The best trade-off between segmentation accuracy and uncertainty calibration is thus achieved with a strong KL regularization term (via w2 ) and a moderate reconstruction term (via w3 ). These trends remained consistent across all acceleration rates. Based on this analysis, we selected the default weights (w1 = 1, w2 = 50, w3 = 1.0) for all segmentation experiments (QUBIQ and SKM-TEA). The fact that our method achieves strong performance on both datasets with these fixed hyperparameters indicates the robustness of our proposed framework. A.3.2
Analysis on Inference Complexity
The inference process for a single pass is feed-forward and efficient. The primary computational cost arises from our method’s ability to draw multiple stochastic samples at test time, which introduces a linear increase in computation. As summarized in the Table 12, increasing the number of stochastic samples systematically improves the GED and stabilizes the Dice score, with diminishing returns observed beyond 8-16 samples. The inference time on an RTX 4090 GPU scales nearly linearly, from 0.006s for one sample to 0.075s for 16 samples. This provides a clinically relevant, tunable accuracylatency trade-off. Compared to single-shot baselines like LOUPE-UNet-RS and SemuNet, InfoMRI achieves markedly better uncertainty calibration (GED) with only a modest increase in latency (e.g., using 4-8 samples), a trade-off we believe is acceptable for many clinical workflows. A.4
Training Stability
A potential challenge in optimizing information-theoretic objectives, such as the one in our framework, lies in the stability of the mutual information (MI) estimator. Both bias in high-dimensional settings and high variance during optimization can impede effective
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
58
TABLE 12: Performance and inference time trade-offs evaluated on QUBIQ 2021 dataset. Performance results are based on subsampling with 16× Acceleration. Inference times are reported based on one NVIDIA GeForce RTX 4090. Method
Number of Samples GED
DSC Infer. Time (s)
1
0.1369 0.9547
0.006
2
0.1146 0.9535
0.011
4
0.0936 0.9576
0.019
6
0.0955 0.9542
0.029
8
0.0869 0.9562
0.038
16
0.0872 0.9563
0.075
LOUPE-UNet-RS
1
0.2145 0.9251
0.004
SemuNet
1
0.1296 0.9556
0.004
Tackle
1
0.1243 0.9570
0.004
InfoMRI (Ours)
training. Our framework, InfoMRI, is architected with specific design choices to address these issues and ensure a stable and reliable training process. To empirically validate the stability of our training process, we present the training dynamics of InfoMRI in Fig. 14. These curves demonstrate that the three sub-losses—the task NLL (log q(t|z, y)), the KL divergence (DKL [q̂(z|t, y)||q(z|y)]), and the reconstruction NLL (log q(x|y))—all converge smoothly as the number of training steps increases to 60k. Simultaneously, key performance metrics on the validation set (e.g., DSC, PSNR) improve and stabilize as expected. This provides direct empirical evidence that our training process is reliable and not compromised by estimator instability. A.4.1
Analysis on Performance vs. Acceleration
We systematically vary the sampling rate and evaluate task performance to identify the “cliff” where performance begins to degrade significantly compared to the fully-sampled case. Specifically, we train an InfoMRI model on QUBIQ 2021 dataset with the sampling ratios sampled from the full range [0,1] and evaluate the GED and Dice performance from 1x acceleration (full sample) to 32x acceleration. The experimental configurations are the same as Table II in the manuscript. As shown in Fig. 15, the performance drop
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
Train Task NLL
Train KL
Val GED
Val DSC
59
Train Reconstruction NLL
Val PSNR
Fig. 14: Training dynamics of InfoMRI. We report the evolution of three sub-losses (Task NLL, KL Divergence, and Reconstruction NLL) and the performance metrics on the validation dataset against the number of training steps. The smooth convergence demonstrates the stability of the optimization process.
is negligible around 1× - 10× acceleration, indicating robust performance across a wide range of acceleration factors.
A PPENDIX B A DDITIONAL E XPERIMENTAL D ETAILS * Implementing segmentation after MRI reconstruction: (1) SemuNet: A two-stage, deterministic pipeline composed of a reconstruction UNet that maps the zero-filled input to a magnitude image, followed by a separate segmentation U-Net. The learned sampling (LOUPE) and both networks are trained end-to-end with a composite objective that combines a reconstruction loss (MSE) and a segmentation loss (cross-entropy). (2) LOUPE-UNet-RS: Maintains the same two-stage architecture (reconstruction UNet → segmentation U-Net), but adopts a staged training protocol: first train LOUPE
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
60
together with the reconstruction U-Net; then freeze the mask and reconstruction U-Net and train only the segmentation U-Net on the reconstructed images. This isolates the segmentation learning while leveraging a learned sampling pattern. (3) TradPat-UNet-RS: Identical two-stage architecture as above, but replaces LOUPE with fixed, handcrafted sampling patterns (uniform random, variable density, or Poisson disk). The reconstruction U-Net is pretrained under each fixed pattern and then frozen; the segmentation U-Net is subsequently trained on the resulting reconstructed images. * Implementing segmentation directly based on measurements: (1) LI-Net: LI-Net consists of (i) a measurement encoder that encodes zero-filled reconstructions and produces a low-dimensional latent code (subsampled k-space data are measured using traditional patterns such as uniform random, Poisson, and variable density), and (ii) a segmentation decoder that decodes this latent code to a segmentation map. The decoder is obtained by pretraining a segmentation autoencoder on groundtruth segmentation; After the autoencoder is trained, the segmentation autoencoder remains frozen, and only the measurement encoder is optimized to align its output latent code with the latent code of the ground-truth segmentation map. At the inference time, the measurement encoder predicts the latent of ground-truth segmentation given the measurement, and then the predicted latent is decoded to the predicted segmentation. (2) LOUPE-LI-Net: Identical task head as LI-Net. The only difference is the sampling pattern used to generate subsampled k-space data, i.e., LOUPE-based learned sampling. (3) Tackle: The task pipeline first applies a reconstruction network to the two-channel measurements (LOUPE-based learned sampling) to obtain a single-channel magnitude image. A subsequent segmentation model (a standard U-shaped convolutional network followed by a sigmoid readout) then predicts the segmentation map from the reconstructed image. Two networks, together with the LOUPE-based learnable sampling, are trained end-to-end for predicting the final segmentation maps from input subsampled k-space data. (4) InfoMRI (Proposed): A probabilistic, task-adapted pipeline that operates directly on undersampled measurements. Similar to LI-Net, InfoMRI consists of (i) a prior encoder that encodes the measurement into a low-dimensional latent code, (ii) a segmentation
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
61
TABLE 13: Training-time trainable parameters by method (task pipeline components). Counts are shown in millions (M) or thousands (K). Method
Trainable Modules
SemuNet
Rec-UNet (31.0 M) + Seg-UNet (31.0 M) + LOUPE (16.4 K) = 62.1 M
TradPat-UNet-RS
Rec-UNet (31.0 M) + Seg-UNet (31.0 M) = 62.0 M
LOUPE-UNet-RS
Rec-UNet (31.0 M) + Seg-UNet (31.0 M) + LOUPE (16.4 K) = 62.1 M
LI-Net
Seg-Enc (34.1 M) + Seg-Dec (49.5 M) + Measure-Enc (34.1 M) = 117 M
LOUPE-LI-Net
Seg-Enc (34.1 M) + Seg-Dec (49.5 M) + Measure-Enc (34.1 M) + LOUPE (16.4 K) = 117 M
Tackle
Rec-UNet (31.0 M) + Seg-UNet (31.0 M) + LOUPE (16.4 K) = 62.1 M
InfoMRI (Ours)
Post-Enc (34.3 M) + Prior-Enc (34.3 M) + Seg-Dec (49.5 M) + Recon-Dec (49.5 M) + PGN (16.4 K) = 167 M
TABLE 14: Inference-time total parameters by method (modules executed at test time). Counts are shown in millions (M) or thousands (K). Method
Total Inference Params
SemuNet
Rec-UNet (31.0 M) + Seg-UNet (31.0 M) + LOUPE (16.4 K) = 62.1 M
TradPat-UNet-RS
Rec-UNet (31.0 M) + Seg-UNet (31.0 M) = 62.0 M
LOUPE-UNet-RS
Rec-UNet (31.0 M) + Seg-UNet (31.0 M) + LOUPE (16.4 K) = 62.1 M
LI-Net
Seg-Dec (49.5 M) + Measure-Enc (34.1 M) = 83.6 M
LOUPE-LI-Net
Seg-Dec (49.5 M) + Measure-Enc (34.1 M) + LOUPE (16.4 K) = 83.6 M
Tackle
Rec-UNet (31.0 M) + Seg-UNet (31.0 M) + LOUPE (16.4 K) = 62.1 M
InfoMRI (Ours)
Prior-Enc (34.3 M) + Seg-Dec (49.5 M) + LOUPE (16.4 K) = 83.8 M
decoder that decodes this latent code to a segmentation map, (iii) a posterior encoder that encodes the ground-truth segmentation map. The difference is that InfoMRI additionally consists of a reconstruction decoder and uses PGN for amortizing optimization across different sampling ratios. * Comparison of the parameter counts: To ensure fairness, we report two complementary quantities: (i) the number of trainable parameters used during task training (training-time task pipeline; sampling policy parameters are excluded unless trained jointly with the task), and (ii) the total param-
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
62
Fig. 15: The GED and Dice performance on the QUBIQ 2021 dataset for InfoMRI trained with the sampling ratios sampled from the full range [0,1]. The experimental configurations are the same as Table II in the manuscript.
eters used at inference for the task pipeline. Exact counts are summarized in Tables 13 and 14. Compared with prior works, InfoMRI carries a larger training-time parameter budget (167M vs. 62–117M; Table 13) because it includes (i) a posterior encoder used only to tighten the variational bound during learning, and (ii) a lightweight reconstruction decoder that stabilizes optimization and improves sample quality. Crucially, both modules are discarded at test time. The inference-time footprint (83.8M; Table 14) is therefore comparable to LI-Net (83.6M), making deployment overhead modest. This extra training capacity yields tangible benefits that we observe consistently across experiments: •
Better uncertainty calibration at the same or similar inference cost: InfoMRI achieves the lowest or among the lowest GED across clinically relevant accelerations (e.g., Table 7 and Table 8) while maintaining competitive DSC, indicating a closer match to the ground-truth posterior than deterministic counterparts.
•
Structured uncertainty for downstream use: sampling from q(t | y) produces anatomically coherent hypotheses that can be marginalized by later modules, a capability not afforded by smaller deterministic heads.
•
Amortized sampling across ratios in a single model: the PGN removes the need
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
63
to train multiple models per acceleration, trading a one-time increase in training parameters for substantial savings in total training runs and easier deployment. In practice, test-time computation involves only the prior encoder, the segmentation decoder, and the PGN. Optional multi-sample evaluation (typically 4–8 samples suffice; Table 12) modestly increases latency but materially improves calibration (GED). Taken together, the larger training-time parameter count is an acceptable and deliberate tradeoff for improved robustness, calibrated uncertainty, and flexible sampling with minimal impact on inference complexity.
R EFERENCES [1]
J. Sun, H. Li, Z. Xu et al., “Deep ADMM-Net for compressive sensing MRI,” in Adv. Neural Inf. Process. Syst. 29, 2016, pp. 10–18.
[2]
J. Zhang and B. Ghanem, “ISTA-Net: Interpretable optimization-inspired deep network for image compressive sensing,” in 2018 IEEE Conf. Comput. Vis. Pattern Recognit., 2018, pp. 1828–1837.
[3]
A. Sriram et al., “End-to-end variational networks for accelerated MRI reconstruction,” in Int. Conf. Med. Image Comput. Comput.-Assist. Interv.–MICCAI 2020, 2020, pp. 64–73.
[4]
I. A. Huijben, B. S. Veeling, and R. J. van Sloun, “Deep probabilistic subsampling for task-adaptive compressed sensing,” in 8th Int. Conf. Learn. Rep., 2020. [Online]. Available: https://research.tue.nl/nl/publications/ 69ddddef-ed02-4293-830a-2cb4e1ec35b3
[5]
H. K. Aggarwal and M. Jacob, “J-MoDL: Joint model-based deep learning for optimized sampling and reconstruction,” IEEE J. Sel. Topics Signal Process., vol. 14, no. 6, pp. 1151–1162, 2020.
[6]
C. D. Bahadir et al., “Deep-learning-based optimization of the under-sampling pattern in MRI,” IEEE Trans. Comput. Imag., vol. 6, pp. 1139–1152, 2020.
[7]
Z. Wang et al., “One network to solve them all: A sequential multi-task joint learning network framework for MR imaging pipeline,” in Int. Workshop Mach. Learn. Med. Image Recon., 2021, pp. 76–85.
[8] [9]
T. Yin et al., “End-to-end sequential sampling and reconstruction for MRI,” in ML4H@ NeurIPS, 2021, pp. 261–281. J. Xie, J. Zhang, Y. Zhang, and X. Ji, “PUERT: Probabilistic under-sampling and explicable reconstruction network for CS-MRI,” IEEE J. Sel. Topics Signal Process., vol. 16, no. 4, pp. 737–749, 2022.
[10] Z. Zhang et al., “Reducing uncertainty in undersampled MRI reconstruction with active acquisition,” in 2019 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019, pp. 2049–2058. [11] T. Bakker, H. van Hoof, and M. Welling, “Experimental design for MRI by greedy policy search,” in Adv. Neural Inf. Process. Syst. 33, 2020, pp. 18 954–18 966. [12] L. Pineda, S. Basu, A. Romero, R. Calandra, and M. Drozdzal, “Active MR k-space sampling with reinforcement learning,” in Int. Conf. Med. Image Comput. Comput.-Assist. Interv.–MICCAI 2020, 2020, pp. 23–33. [13] H. Van Gorp, I. Huijben, B. S. Veeling, N. Pezzotti, and R. J. Van Sloun, “Active deep probabilistic subsampling,” in Proc. 38th Int. Conf. Mach. Learn., 2021, pp. 10 509–10 518.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
64
[14] M. Monteiro et al., “Stochastic segmentation networks: Modelling spatially correlated aleatoric uncertainty,” in Adv. Neural Inf. Process. Syst. 33, 2020, pp. 12 756–12 767. [15] J. Schlemper et al., “Cardiac MR segmentation from undersampled k-space using deep latent representation learning,” in Int. Conf. Med. Image Comput. Comput.-Assist. Interv.–MICCAI 2018, 2018, pp. 259–267. [16] Z. Wu, T. Yin, Y. Sun, R. Frost, A. van der Kouwe, A. V. Dalca, and K. L. Bouman, “Learning task-specific strategies for accelerated MRI,” IEEE Trans. Comput. Imaging, vol. 10, pp. 1040–1054, 2024. [17] J. Caballero, W. Bai, A. N. Price, D. Rueckert, and J. V. Hajnal, “Application-driven MRI: joint reconstruction and segmentation from undersampled MRI data,” in Int. Conf. Med. Image Comput. Comput.-Assist. Interv.–MICCAI 2014, 2014, pp. 106–113. [18] C. F. Baumgartner et al., “PHiSeg: Capturing uncertainty in medical image segmentation,” in Int. Conf. Med. Image Comput. Comput.-Assist. Interv.–MICCAI 2019, 2019, pp. 119–127. [19] T. M. Cover and J. A. Thomas, Elements of Information Theory.
John Wiley & Sons, 2005.
[20] L. Wang et al., “Information-theoretic compressive measurement design,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 6, pp. 1150–1164, 2016. [21] B. Amos et al., “Tutorial on amortized optimization,” Found Trends® Mach Learn, vol. 16, no. 5, pp. 592–732, 2023. [22] C. Mou and J. Zhang, “TransCL: Transformer makes strong and flexible compressive learning,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 4, pp. 5236–5251, 2023. [23] T. Bakker et al., “On learning adaptive acquisition policies for undersampled multi-coil MRI reconstruction,” in Int. Conf. Med. Imag. Deep Learn., 2022, pp. 63–85. [24] M. W. Seeger and H. Nickisch, “Compressed sensing and Bayesian experimental design,” in Proc. 25th Int. Conf. Mach. Learn., 2008, pp. 912–919. [25] S. Ji, Y. Xue, and L. Carin, “Bayesian compressive sensing,” IEEE Trans. Signal Process., vol. 56, no. 6, pp. 2346– 2356, 2008. [26] H. Nickisch, R. Pohmann, B. Schölkopf, and M. Seeger, “Bayesian experimental design of magnetic resonance imaging sequences,” in Adv. Neural Inf. Process. Syst. 21, 2008, pp. 1441–1448. [27] A. Grover and S. Ermon, “Uncertainty autoencoders: Learning compressed representations via variational information maximization,” in Proc. 22nd Int. Conf. Artif. Intell. Stat., 2019, pp. 2514–2524. [28] A. Foster et al., “Variational Bayesian optimal experimental design,” in Adv. Neural Inf. Process. Syst. 32, 2019, pp. 14 036–14 047. [29] A. Foster, D. R. Ivanova, I. Malik, and T. Rainforth, “Deep adaptive design: Amortizing sequential Bayesian experimental design,” in Proc. 38th Int. Conf. Mach. Learn., 2021, pp. 3384–3395. [30] S. Kleinegesse and M. U. Gutmann, “Efficient Bayesian experimental design for implicit models,” in Proc. 22nd Int. Conf. Artif. Intell. Stat., 2019, pp. 476–485. [31] ——, “Bayesian experimental design for implicit models by mutual information neural estimation,” in Proc. 37th Int. Conf. Mach. Learn., 2020, pp. 5316–5326. [32] J. Beck et al., “Fast Bayesian experimental design: Laplace-based importance sampling for the expected information gain,” Comput. Meth. Appl. Mech. Eng., vol. 334, pp. 523–553, 2018. [33] Q. Huang, X. Chen, D. Metaxas, and M. S. Nadar, “Brain segmentation from k-space with end-to-end recurrent attention network,” in Int. Conf. Med. Image Comput. Comput.-Assist. Interv.–MICCAI 2019, 2019, pp. 275–283.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
65
[34] Q. Huang, D. Yang, J. Yi, L. Axel, and D. Metaxas, “FR-Net: Joint reconstruction and segmentation in compressed sensing cardiac MRI,” in Int. Conf. Functional Imaging Model. Heart, 2019, pp. 352–360. [35] L. Sun, Z. Fan, X. Ding, Y. Huang, and J. Paisley, “Joint CS-MRI reconstruction and segmentation with a unified deep network,” in Int. Conf. Inf. Process. Med. Imag., 2019, pp. 492–504. [36] F. Calivá et al., “Breaking speed limits with simultaneous ultra-fast MRI reconstruction and tissue segmentation,” in Proc. 3rd Med. Imag. Deep Learn., 2020, pp. 94–110. [37] F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier-Hein, “nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation,” Nat. Methods, vol. 18, no. 2, pp. 203–211, 2021. [38] S. Kohl et al., “A probabilistic U-Net for segmentation of ambiguous images,” in Adv. Neural Inf. Process. Syst. 31, 2018, pp. 6965–6975. [39] S. A. Kohl et al., “A hierarchical probabilistic U-Net for modeling multi-scale ambiguities,” arXiv preprint arXiv:1905.13077, 2019. [Online]. Available: https://arxiv.org/abs/1905.13077 [40] W. Zhang et al., “PixelSeg: Pixel-by-pixel stochastic semantic segmentation for ambiguous medical images,” in Proc. 30th ACM Int. Conf. Multimedia, 2022, pp. 4742–4750. [41] R. Selvan, F. Faye, J. Middleton, and A. Pai, “Uncertainty quantification in medical image segmentation with normalizing flows,” in 11th Int. Workshop Mach. Learn. Med. Imaging (MLMI), 2020, pp. 80–90. [42] L. Zbinden et al., “Stochastic segmentation with conditional categorical diffusion models,” in 2023 IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 1119–1129. [43] Z. Gao, Y. Chen, C. Zhang, and X. He, “Modeling multimodal aleatoric uncertainty in segmentation with mixture of stochastic experts,” in 11th Int. Conf. Learn. Rep., 2023. [Online]. Available: https://arxiv.org/abs/2212.07328 [44] N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” arXiv preprint physics/0004057, 2000. [Online]. Available: https://arxiv.org/abs/physics/0004057 [45] A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, “Deep variational information bottleneck,” in 5th Int. Conf. Learn. Rep., 2017. [Online]. Available: https://openreview.net/forum?id=HyxQzBceg [46] J. Sijbers and A. Den Dekker, “Maximum likelihood estimation of signal amplitude and noise variance from MR data,” Magn. Reson. Med., vol. 51, no. 3, pp. 586–594, 2004. [47] D. Barber and F. Agakov, “The IM algorithm: a variational approach to information maximization,” in Adv. Neural Inf. Process. Syst. 16, 2004, pp. 201–208. [48] G. E. Forsythe, “Von Neumann’s comparison method for random sampling from the normal and other distributions,” Math. Comput., vol. 26, no. 120, pp. 817–826, 1972. [49] F. Sherry et al., “Learning the sampling pattern for MRI,” IEEE Trans. Med. Imag. vol. 39, no. 12, pp. 4310–4321, 2020. [50] D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in 2nd Int. Conf. Learn. Rep., 2014. [Online]. Available: https://arxiv.org/abs/1312.6114 [51] J. Tomczak and M. Welling, “VAE with a VampPrior,” in Proc. 21st Int. Conf. Artif. Intell. Stat., 2018, pp. 1214–1223. [52] D. M. Blei, A. Kucukelbir, and J. D. McAuliffe, “Variational inference: A review for statisticians,” J. Am. Stat. Assoc., vol. 112, no. 518, pp. 859–877, 2017. [53] A. Defazio, “Offset sampling improves deep learning based accelerated MRI reconstructions by exploiting symmetry,” arXiv preprint arXiv:1912.01101, 2019. [Online]. Available: https://arxiv.org/abs/1912.01101 [54] C. M. Bishop, Pattern Recognition and Machine Learning.
Springer New York, NY, 2006.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
66
[55] A. Kendall, Y. Gal, and R. Cipolla, “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” in 2018 IEEE Conf. Comput. Vis. Pattern Recognit., 2018, pp. 7482–7491. [56] T. Karras, M. Aittala, J. Lehtinen, J. Hellsten, T. Aila, and S. Laine, “Analyzing and improving the training dynamics of diffusion models,” in 2024 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2024, pp. 24 174– 24 184. [57] Z. Wang et al., “Promoting fast MR imaging pipeline by full-stack AI,” iScience, vol. 27, no. 1, article no. 108 608, 2024. [58] T. Blau, E. V. Bonilla, I. Chades, and A. Dezfouli, “Optimizing sequential experimental design with deep reinforcement learning,” in Proc. 39th Int. Conf. Mach. Learn., 2022, pp. 2107–2128. [59] J. Zbontar et al., “fastMRI: An open dataset and benchmarks for accelerated MRI,” arXiv preprint arXiv:1811.08839, 2018. [Online]. Available: https://arxiv.org/abs/1811.08839 [60] E. J. Candès, J. Romberg, and T. Tao, “Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency information,” IEEE Trans. Inf. Theory, vol. 52, no. 2, pp. 489–509, 2006. [61] M. Lustig, D. Donoho, and J. M. Pauly, “Sparse MRI: The application of compressed sensing for rapid MR imaging,” Magn. Reson. Med., vol. 58, no. 6, pp. 1182–1195, 2007. [62] J. Vellagoundar and R. R. Machireddy, “A robust adaptive sampling method for faster acquisition of MR images,” Magn. Reson. Imaging, vol. 33, no. 5, pp. 635–643, 2015. [63] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in 18th Int. Conf. Med. Image Comput. Comput.-Assist. Interv., 2015, pp. 234–241. [64] Y. Yang, J. Sun, H. Li, and Z. Xu, “ADMM-CSNet: A deep learning approach for image compressive sensing,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 3, pp. 521–538, 2018. [65] K. Zhang, Y. Li, W. Zuo, L. Zhang, L. Van Gool, and R. Timofte, “Plug-and-play image restoration with deep denoiser prior,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 10, pp. 6360–6376, 2021. [66] A. Pour Yazdanpanah, O. Afacan, and S. Warfield, “Deep plug-and-play prior for parallel MRI reconstruction,” in 2019 IEEE/CVF Int. Conf. Comput. Vis. Workshops (ICCVW), 2019, pp. 3952–3958 [67] M. Ran et al., “MD-Recon-Net: A Parallel Dual-Domain Convolutional Neural Network for Compressed Sensing MRI,” IEEE Trans. Radiat. Plasma Med. Sci., vol. 5, no. 1, pp. 120–135, 2020. [68] C. Hu et al., “Self-supervised learning for MRI reconstruction with a parallel network training framework,” in Int. Conf. Med. Image Comput. Comput.-Assist. Interv.–MICCAI 2021, 2021, pp. 382–391. [69] R. Rombach et al., “High-resolution image synthesis with latent diffusion models,” in 2022 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 10 684–10 695. [70] I. Higgins et al., “beta-VAE: Learning basic visual concepts with a constrained variational framework,” in 5th Int. Conf. Learn. Rep., 2017. [Online]. Available: https://openreview.net/forum?id=Sy2fzU9gl [71] B. Menze et al., “Quantification of uncertainties in biomedical image quantification challenge 2021,” 2021. [Online]. Available: https://qubiq21.grand-challenge.org/QUBIQ2021/ [72] J. L. Silva and A. L. Oliveira, “Using soft labels to model uncertainty in medical image segmentation,” in Int. MICCAI Brainlesion Workshop, 2021, pp. 585–596. [73] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 770–778.
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE
67
[74] R. Zhao et al., “fastMRI+, Clinical pathology annotations for knee and brain fully sampled multi-coil MRI data,” Sci. Data, vol. 9, no. 152, 2022. [75] C.-Y. Yen et al., “Adaptive sampling of k-space in magnetic resonance for rapid pathology prediction,” in Proc. 41st Int. Conf. Mach. Learn., 2024, pp. 57 018–57 032. [76] A. D. Desai, A. M. Schmidt, E. B. Rubin, C. M. Sandino, M. S. Black, V. Mazzoli, K. J. Stevens, R. Boutin, C. Re, G. E. Gold et al., “Skm-tea: A dataset for accelerated mri reconstruction with dense image labels for quantitative clinical evaluation,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. [77] A. Fateh, Y. Rezvani, S. Moayedi, S. Rezvani, F. Fateh, M. Fateh, and V. Abolghasemi, “Brisc: Annotated dataset for brain tumor segmentation and classification,” Scientific Data, 2026.