ConceptioArchivearXiv CS
arXiv CSopen access

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Smriti Joshia,∗ , Apostolia Tsirikoglouc, , Daniel M. Langd,e, , Richard Osualaa,g, , Noah Márquez Varaa,h , Alejandro Guzmana , Grzegorz Skorupkoa , Sebastian Ibarra Arreguia , Lidia Garruchoa , Akane Ohashii,j , Dimitra Ntoulak , Eugen Divjakl,m , Oğuz Lafcın , Jan C. Peekeng , Julia A. Schnabeld,e,f , Fredrik Strandc , Oliver Diaza and Karim Lekadira,b †

a Departament de Matemàtiques i Informàtica, Universitat de Barcelona, Barcelona, Spain b Institució Catalana de Recerca i Estudis Avançats (ICREA), Barcelona, Spain c Department of Oncology-Pathology, Karolinska Institutet, SE-171 77, Stockholm, Sweden d Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, Munich, Germany e School of Computation, Information and Technology, Technical University of Munich, Munich, Germany f School of Biomedical Engineering and Imaging Sciences, King’s College London, London, United Kingdom

arXiv:2607.29394v1 [cs.CV] 31 Jul 2026

g Department of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of

Munich, Munich, Germany h Department of Computer Science and Engineering, Chalmers University of Technology, Gothenburg, Sweden i Department of Translational Medicine, Diagnostic Radiology, & CIRCE – the Center for Interdisciplinary Research on Cancer and Equity in Women, Lund University, Lund, Sweden j Department of Imaging and Physiology, Skåne University Hospital, Malmö, Sweden k Department of Radiology, Karolinska University Hospital, 17177, Stockholm, Sweden l University of Zagreb, School of Medicine, Zagreb, Croatia m University Hospital Dubrava, Zagreb, Croatia n Department of Biomedical Imaging and Image-Guided Therapy, Medical University of Vienna, Vienna, Austria

ARTICLE INFO

ABSTRACT

Keywords: Contrast Enhancement Breast Cancer Magnetic Resonance Imaging Clinical Reader Study Tumor Segmentation Generative Models

Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast agents (GBCAs) restricts use in contraindicated populations, prolongs scan protocols, and presents environmental toxicity concerns. Contrast synthesis offers a non-invasive alternative; however, existing approaches struggle to balance spatial realism with temporal continuity, suffer from slow iterative sampling, underutilize structural priors, and lack clinical validation. We propose a novel conditioned latent transport framework that predicts contrast enhancement in a single forward pass. By anchoring the latent trajectory to the pre-contrast anatomy and applying continuous time conditioning, the model synthesizes patient-specific contrast evolution at any given acquisition time. The proposed approach outperforms baseline and the stateof-the-art models across spatial, perceptual, temporal, and distributional metrics. Evaluated on an independent external cohort, the method demonstrates robustness to domain shifts induced by scanner noise as well as differing acquisition protocol. Furthermore, our synthetic contrast enhancement significantly improved downstream tumor segmentation performance, yielding a 22.4% relative increase in Dice coefficient (0.60 vs. 0.49 baseline pre-contrast, 𝑝 < 0.01), reducing boundary segmentation error by over 39%, while outperforming all other generative model baselines. Finally, a reader study involving four breast radiologists evaluated the image quality, kinetic fidelity, and diagnostic viability of our synthesized sequences across 40 randomly selected cases. The results demonstrated that in 70% of cases, synthesized images provided sufficient clinical information to support the same management decisions as real DCE-MRI, suggesting a path toward safer and faster contrast-free or contrast-reduced imaging workflows.

1. Introduction Magnetic resonance imaging (MRI) is extensively used in clinical practice due to its high diagnostic sensitivity and soft-tissue contrast. It plays an integral role in breast cancer staging (Selvi et al., 2018), treatment monitoring, and surgical planning (Parvaiz et al., 2016; Kuhl et al., 2017). Unlike other imaging modalities like mammography and ultrasound, which are known to underestimate tumor † These authors contributed equally to this work.

∗ Corresponding author

[email protected] (S. Joshi)

ORCID (s): 0000-0001-8480-023X (S. Joshi)

S. Joshi et al.: Preprint submitted to Elsevier

extent (Dixon et al., 2016; Uematsu et al., 2008), MRI enables the accurate delineation of tumor margins, potentially reducing surgical re-excision rates (Gonzalez et al., 2014). In addition to treatment management, breast MRI is a critical screening modality for high-risk populations, including women carrying genetic mutations, those with a history of chest irradiation before age 30, or those presenting dense breast tissue (Saslow et al., 2007; Monticciolo et al., 2018). The clinical value of MRI in screening is also significant; in one study, single-screening MRI has been shown to depict 18.1 additional cancers per 1,000 women with a history of breast-conserving therapy (BCT) (Gweon et al., 2014). Moreover, the evaluation of time-intensity curves detailing

Page 1 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Figure 1: Overview of the study design with validation framework. Our proposed (left) method leverages a pre-contrast anchor and continuous acquisition time (𝜏) to synthesize temporally consistent post-contrast DCE-MRIs (center). Clinical viability is evaluated via a three part validation strategy (right) comprising: (1) multi-site evaluation for domain robustness, (2) downstream tumor segmentation for biological fidelity, and (3) a clinical reader study to assess diagnostic impact.

contrast enhancement serves as a versatile biomarker, providing, among others, a reliable, non-invasive estimation of tumor malignancy (Kuhl et al., 1999; Partridge et al., 2014). The high sensitivity of MRI is largely owed to contrast agents (Pesapane et al., 2025), most commonly gadoliniumbased contrast agents (GBCAs). However, routine reliance on GBCAs introduces universal procedural burdens, alongside well-recognized toxicity risks (Rogosnitzky and Branch, 2016; Fraum et al., 2017; Alahari and Benstetter, 2017). From an operational perspective, contrast administration requires intravenous administration, which invariably lengthens scan preparation time and carries a risk of access failure requiring specialized staff assistance. Furthermore, safety protocols mandate the immediate physical availability of a supervising physician, imposing geographical and scheduling restrictions on where and when breast MRI can be performed. In terms of patient safety, GBCAs pose significant risks and contraindications for vulnerable populations, including pregnant women (Peterson et al., 2023) and patients with renal failure (Grobner and Prischl, 2007). They have also been associated with retention and accumulation in the brain (Kanda et al., 2015) and bone tissue (Darrah et al., 2009), particularly following exposure to linear agents. Although newer macrocyclic agents appear to reduce tissue retention, they may not completely eliminate it (Murata et al., 2016), and the clinical significance of this persistence remains uncertain. The environmental consequences of GBCAs are also considerable. Because these agents are predominantly excreted renally without undergoing metabolization, they resist

S. Joshi et al.: Preprint submitted to Elsevier

degradation in wastewater treatment plants and are continuously emitted into aquatic ecosystems (Bau and Dulski, 1996; Laczovics et al., 2023). Ecologically, the 𝐺𝑑 3+ ion is capable of mimicking essential cations, including calcium, zinc, magnesium, and iron (Cotruvo, 2019), thereby threatening to disrupt critical cellular and biochemical pathways (Krasznai et al., 2003). The potential for these agents to impact the food chain (Arciszewska et al., 2022) through plants (Lindner et al., 2013), terrestrial as well as aquatic life (Lingott et al., 2016), alongside their unknown longterm ecological effects, constitutes a pressing and unresolved concern. To mitigate these risks and reduce overall scan times, research has focused on synthesizing dynamic contrastenhanced (DCE) MRI from non-contrast, early-phase, or low-dose acquisitions. Explored methods include generative adversarial networks (GANs) (Dar et al., 2019; MüllerFranzes et al., 2023; Osuala et al., 2025; Fonnegra et al., 2025), probabilistic diffusion models (Osuala et al., 2024; Kishore Kumar et al., 2024; Ibarra et al., 2025; Fan et al., 2025; Kong et al., 2026), and deterministic approaches (Lang et al., 2025; Chung et al., 2025a; Chen et al., 2026). However, three primary limitations persist across these paradigms. First, enforcing temporal consistency across DCE sequences often compromises spatial realism. Realistic high-frequency textures are typically generated via stochastic mechanisms, such as noise injection or adversarial sampling; however, this inherent randomness introduces non-physiologicallygrounded variations across sequential frames, disrupting smooth temporal continuity. Second, iterative generative models require extensive inference time, limiting computational efficiency. Third, existing methods underutilize precontrast anatomical priors; requiring a network to synthesize baseline structures from scratch distracts it from accurately modeling true contrast dynamics. Beyond these technical challenges, the true clinical utility of existing models remains difficult to assess due to a widespread lack of downstream task validation and expert reader evaluation. To address these gaps collectively, we propose a framework for dense temporal contrast synthesis, and rigorously evaluate it to position the method in terms of generalizability and clinical utility. A respective overview of this work is presented in Figure 1. Our principal contributions are summarized as follows: • We introduce a novel conditioned latent transport network that synthesizes patient-specific DCE-MRI at any continuous acquisition time (𝜏) from its corresponding pre-contrast image. • We demonstrate that anchoring the generative trajectory to the pre-contrast anatomy enforces macroscopic structural fidelity, while tailored spectral and structural loss functions preserve high-frequency details and perceptual realism.

Page 2 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Figure 2: Overview of the proposed Latent Generative Model Architecture. Pre-contrast (𝑥𝑝𝑟𝑒 ) and post-contrast (𝑥𝑝𝑜𝑠𝑡 ) images are compressed into a high-fidelity latent space via a frozen encoder () of a custom pretrained autoencoder. Then, the intermediate state (𝑧𝑡 ) is constructed by interpolating between the non-enhanced structural anchor (𝑧𝑝𝑟𝑒 ) and the target state (𝑧𝑝𝑜𝑠𝑡 ) with added stochasticity (𝜎). Conditioned on 𝑧𝑝𝑟𝑒 , a sinusoidal pharmacokinetic time embedding (𝜏), and interpolation timestep t, the central U-Net acts as a target-predictor, directly regressing the constant latent subtraction map (𝑧̂ 𝑒𝑛ℎ ≈ 𝑧𝑝𝑜𝑠𝑡 − 𝑧𝑝𝑟𝑒 ). Adding the 𝑧𝑝𝑟𝑒 to the predicted enhancement, 𝑧̂ 𝑝𝑜𝑠𝑡 is obtained to compute the latent supervision loss (𝑀𝑆𝐸 ). Then, 𝑧̂ 𝑝𝑜𝑠𝑡 is passed through a frozen decoder () to upsample the image to pixel space. Structural and perceptual fidelity are strictly enforced via perceptual loss (𝐿𝑃 𝐼𝑃 𝑆 ), while a log-amplitude Fourier loss (𝐹 𝐹 𝐿 ) explicitly penalizes frequency deviations to preserve fine micro-vascular details.

• We evaluate our method on an independent, external validation cohort, showing improvement over the precontrast image and competing state-of-the-art methods, while also systematically analyzing the domain shift and its effect on performance. • We validate the anatomical fidelity of synthesized images through downstream tumor segmentation, showing statistically significant improvements over competing methods across two complementary evaluation paradigms. • We conduct a comprehensive clinical reader study with four expert radiologists to assess the synthesized images, demonstrating their diagnostic viability and potential to reduce reliance on contrast administration in clinical practice. The remainder of this manuscript is organized as follows: Section 2 reviews the current state-of-the-art in contrast synthesis. Section 3 outlines our proposed methodology. Section 4 discusses the implementation details. Section 5 discusses the quantitative and qualitative results. Finally, Section 6 summarizes our core findings and highlights potential directions for future research.

2. Related Work Early methodologies in contrast phase synthesis leveraged GAN based architectures such as pix2pix (Isola et al., 2017), for image-to-image generation of post-contrast phases from pre-contrast sequences (Osuala et al., 2025; MüllerFranzes et al., 2023). GANs have also been applied to crossphase synthesis, demonstrating the feasibility of predicting late post-contrast phases from early-phase acquisitions (Fonnegra et al., 2025; Müller-Franzes et al., 2024). While S. Joshi et al.: Preprint submitted to Elsevier

these models successfully capture high-frequency structural details, they often struggle with training instability and the hallucination of non-existent anatomical features. Diffusion models (Ho et al., 2020; Rombach et al., 2022) provide an alternative to GANs because of their stable training dynamics and high-fidelity sample generation. In the medical imaging domain, diffusion models have been widely adopted across a spectrum of downstream tasks, including image synthesis (Konz et al., 2024; Pinaya et al., 2022), artifact denoising (Gao et al., 2023), and semantic segmentation (Yan et al., 2024a). Specifically, within contrast synthesis, ContrastControlNet (CCNet) (Osuala et al., 2024) trains the latent diffusion model for pre to post contrast synthesis in breast imaging, while parallel frameworks have been optimized for prostate imaging (M et al., 2024). To further constrain the generative process and ensure anatomical fidelity, several approaches have incorporated multimodal conditioning inputs, integrating raw imaging data with clinical metadata and explicit outline guidance (Ibarra et al., 2025; Konz et al., 2024; Fan et al., 2025; Osuala et al., 2024). Probabilistic diffusion models are valuable for modeling data diversity; however, the preference for consistent, repeatable outputs in clinical workflows motivates the shift toward deterministic synthesis. In this realm, works on breast DCE-MRI contrast synthesis rely on hierarchical networks (Zhang et al., 2023), temporal neural cellular automata (Lang et al., 2025), and iterative networks (Chung et al., 2025b). These works promise a strong direction by demonstrating improved anatomical and temporal fidelity, which are frequently sacrificed in favor of perceptual realism within traditional stochastic models. Another line of research are cold diffusion models, which replace stochastic Gaussian perturbations with deterministic degradation operators (e.g. Page 3 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

blurring) and learn the corresponding inverse restoration process (Bansal et al., 2023). Although not yet explored for contrast synthesis, this formulation has demonstrated promising results across a variety of medical imaging applications, including segmentation (Yan et al., 2024b; Zaman et al., 2024; Qi et al., 2025), anomaly detection (Naval Marimont et al., 2024), denoising (Zhang et al., 2025; Tang et al., 2025), and image reconstruction (Shen et al., 2024; Habijan et al., 2025). Importantly, the deterministic degradation trajectory provides a more structured restoration process compared to stochastic sampling, which can facilitate efficient inference. A parallel direction in generative modeling is represented by flow-based approaches, including flow matching and rectified flow formulations, which learn a time-dependent velocity field to transport samples between source and target distributions (Lipman et al., 2022; Liu et al., 2022). These methods have been explored in latent generative frameworks (Dao et al., 2023), Computed Tomography (CT) image synthesis (Wang et al., 2025), multimodal MRI translation (Tur et al., 2026), and more recently for contrast synthesis (Chen et al., 2026), where the task is formulated as a missing modality prediction problem. Despite their impressive generative capabilities, both stochastic diffusion models and flow-based generative methods typically rely on multi-step denoising or numerical integration during inference. While effective for unconstrained image synthesis, repeated integration is not inherently required for contrast enhancement, where the source anatomy is already observed and only the physiological contrast uptake must be modeled, an observation also explored in conditional flow formulations (Tur et al., 2026). Cold Diffusion replaces iterative denoising targets with direct endpoint prediction from deterministically corrupted intermediate states (Bansal et al., 2023). Concurrently, Rectified Flow frames generation as a transport problem learned from interpolated intermediate states, providing supervision along the transformation path rather than only at the endpoints (Lipman et al., 2022). Motivated by these observations, we formulate contrast synthesis as a conditional latent transport problem that combines endpoint prediction with structured interpolation, enabling efficient single-step inference while retaining the richer supervision provided by intermediatestate training.

3. Methodology Our method is presented visually in Figure 2. We detail the mathematical formulation, architectural pipeline, and inference strategy of the proposed method below.

3.1. Problem Formulation Let  ⊂ ℝ𝐶×𝐻 denote the pixel space of 2D MR image slices, where 𝐻 and 𝑊 represent the spatial height and width, respectively. For a given patient, standard clinical protocols acquire a non-enhanced pre-contrast baseline slice 𝑥𝑝𝑟𝑒 ∈  and a target contrast-enhanced slice 𝑥𝑝𝑜𝑠𝑡 (𝜏𝑖 ) ∈  captured at a discrete, protocol-specific physical timepoint S. Joshi et al.: Preprint submitted to Elsevier

𝜏𝑖 . However, while real scanner acquisitions are fundamentally discrete and temporally sparse, physiological contrast enhancement is a continuously evolving process. Therefore, our objective is to learn a continuous generative transport mapping 𝜃 ∶ (𝑥𝑝𝑟𝑒 , 𝜏) ↦ 𝑥̂ 𝑝𝑜𝑠𝑡 (𝜏) that accepts any arbitrary timepoint 𝜏 ∈ ℝ+ . This formulation allows us to accurately synthesize the non-linear pharmacokinetic hemodynamics of localized tumor regions across the entire temporal domain, while strictly preserving the underlying patientspecific anatomical structure and global tissue topology.

3.2. Latent Space Compression and Custom Autoencoder To mitigate the severe computational bottlenecks associated with high-resolution medical imaging, we map the pixel space  into a compressed, low-dimensional latent manifold 𝑐× 𝐻 × 𝑊

 ⊂ ℝ 𝑓 𝑓 using a Variational Autoencoder (VAE), where 𝑓 denotes the spatial downsampling factor. Let  and  denote the encoder and decoder, such that 𝑧 = (𝑥) and 𝑥 ≈ (𝑧). Standard stable diffusion VAEs (Rombach et al., 2022) typically employ an 𝑓 = 8 spatial downsampling factor, achieved through three consecutive convolutional blocks with a stride of 2. While computationally efficient, we empirically observed that this aggressive compression discards the fine-grained micro-vasculature and parenchymal textures essential for clinical diagnostics (Appendix A.3). To address this spatial bottleneck, we train a custom AutoencoderKL1 architecture from scratch, strictly constrained to a milder 𝑓 = 4 spatial downsampling factor. We optimize this encoder-decoder pair on our pre-contrast and post-contrast sequences using the standard VAE objective: 𝑉 𝐴𝐸 = 𝑟𝑒𝑐𝑜𝑛 (𝑥, ((𝑥))) + 𝛽𝐾𝐿 (𝑞𝜙 (𝑧|𝑥)||𝑝(𝑧)) , where 𝑟𝑒𝑐𝑜𝑛 is a composite reconstruction loss combining Mean Squared Error (MSE) for signal fidelity and Learned Perceptual Image Patch Similarity (LPIPS) (Zhang et al., 2018) for structural sharpness, and 𝐾𝐿 is the KullbackLeibler divergence used to regularize the latent manifold (Kingma and Welling, 2013) by forcing the learned posterior 𝑞𝜙 (𝑧|𝑥) to approximate a standard normal prior 𝑝(𝑧). We observe that this 4× configuration balances computational efficiency with the retention of micro-scale structures required for accurate contrast synthesis. Once trained,  and  are frozen. The resulting latent representations (𝑧𝑝𝑟𝑒 and 𝑧𝑝𝑜𝑠𝑡 ) provide the mathematical domain for the contrast synthesis network.

3.3. Forward Process and Temporal Conditioning The core objective is to transport the non-enhanced baseline latent 𝑧𝑝𝑟𝑒 to the target enhanced state 𝑧𝑝𝑜𝑠𝑡 (𝜏). We frame this generation as a conditioned latent process governed by a linear interpolant. The network, parameterized as an acquisition-time conditioned U-Net 𝜃 , is trained to map a corrupted intermediate state 𝑧𝑡 to the fully enhanced target 1 https://huggingface.co/docs/diffusers/en/api/models/autoencoderkl

Page 4 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

state. The forward process corrupts the target state directly toward the patient’s non-enhanced baseline 𝑧𝑝𝑟𝑒 : 𝑧𝑡 = (1 − 𝑤𝑡 )𝑧𝑝𝑜𝑠𝑡 + 𝑤𝑡 𝑧𝑝𝑟𝑒 + (𝑤𝑡 ⋅ 𝜎𝑛𝑜𝑖𝑠𝑒 )𝜖 , where 𝑤𝑡 ∈ [0, 1] dictates the temporal degradation schedule, 𝑡 ∈ [0, 1000] is the timestep, and 𝜖 ∼  (0, 𝐼) introduces stochasticity to prevent deterministic collapse. The degradation weight 𝑤𝑡 governs the structural denoising process and is strictly decoupled from the physical acquisition time 𝜏. The stochastic noise variance scales linearly with 𝑤𝑡 , maximizing at the baseline (𝑤𝑡 = 1) and decaying to zero at the target (𝑤𝑡 = 0). We treat the physical acquisition time 𝜏 as a continuous conditional prior. It is encoded via Sinusoidal Positional Embeddings (Vaswani et al., 2017) and passed through a Multi-Layer Perceptron (MLP) to match the generative network’s feature dimensionality. This condition, alongside the timestep 𝑡, is injected into the U-Net residual blocks via Adaptive Group Normalization (AdaGN) (Dhariwal and Nichol, 2021). Conditioning explicitly on 𝜏 enables the model to learn spatially-varying intensity mappings, assigning distinct physiological kinetics to different biological structures. The spatial input is constructed via channel-wise concatenation, yielding 𝑧𝑖𝑛 = [𝑧𝑡 ∥ 𝑧𝑝𝑟𝑒 ]. The generative mapping is thus formalized as 𝑧̂ 𝑝𝑜𝑠𝑡 = 𝜃 (𝑧𝑖𝑛 , 𝑡, 𝜏).

3.4. Residual Parameterization and Training Objective Because the non-enhanced anatomical topology remains strictly static during the localized acquisition window, the pharmacokinetic transformation can be expressed in residual form as Δ𝑧 = 𝑧𝑝𝑜𝑠𝑡 − 𝑧𝑝𝑟𝑒 . To improve representational efficiency, we parameterize the latent U-Net 𝜃 to predict this residual component directly, yielding Δ𝑧̂ = 𝜃 (𝑧𝑖𝑛 , 𝑡, 𝜏). This residual formulation decoupling the representation of the transformation (via 𝑧𝑡 ) from the prediction objective (via Δ𝑧) prevents the network from allocating capacity to redundant anatomical reconstruction to simplify optimization. The final enhanced latent is reconstructed as 𝑧̂ 𝑝𝑜𝑠𝑡 = 𝑧𝑝𝑟𝑒 + Δ𝑧. ̂ We train the model using a composite objective function to ensure signal fidelity, perceptual realism, and the retention of fine frequency details: 𝑡𝑜𝑡𝑎𝑙 = 𝜆𝑀𝑆𝐸 𝑀𝑆𝐸 + 𝜆𝐿𝑃 𝐼𝑃 𝑆 𝐿𝑃 𝐼𝑃 𝑆 + 𝜆𝐹 𝐹 𝐿 𝐹 𝐹 𝐿 Signal Fidelity (𝑀𝑆𝐸 ): We apply a standard MSE loss between the predicted and ground-truth latents, defined as ||𝑧̂ 𝑝𝑜𝑠𝑡 − 𝑧𝑝𝑜𝑠𝑡 ||22 , to guarantee macroscopic pharmacokinetic accuracy and correct global intensity scaling. Perceptual Realism (𝐿𝑃 𝐼𝑃 𝑆 ): To overcome the deterministic blurring inherent to purely pixel-wise objectives, we evaluate LPIPS (Zhang et al., 2018) loss between the decoded images 𝑥𝑝𝑜𝑠𝑡 and 𝑥̂ 𝑝𝑜𝑠𝑡 . By comparing deep feature activations, this term enforces high-level structural coherence and prevents regression-to-the-mean artifacts. Spectral Fidelity (𝐹 𝐹 𝐿 ): Standard spatial losses often fail to resolve chaotic, high-frequency micro-vasculature. S. Joshi et al.: Preprint submitted to Elsevier

We incorporate a Focal Frequency Loss (FFL) (Jiang et al., 2021) to explicitly penalize discrepancies in the Fourier domain. Let  denote the 2D orthogonal Fast Fourier Transform. We map the predicted and ground-truth latents through the frozen decoder and penalize the logarithmic difference of their amplitude spectra: ‖ ‖ 𝐹 𝐹 𝐿 = ‖log(1 + |((𝑧̂ 𝑝𝑜𝑠𝑡 ))|) − log(1 + |((𝑧𝑝𝑜𝑠𝑡 ))|)‖ ‖ ‖1 This logarithmic scaling ensures that the network is mathematically incentivized to recover peripheral high-frequency clinical textures rather than being overpowered by the zerofrequency (DC) component. The values of 𝜆𝑀𝑆𝐸 , 𝜆𝐿𝑃 𝐼𝑃 𝑆 , and 𝜆𝐹 𝐹 𝐿 were empirically determined on the validation set as 1, 5, and 50, respectively.

3.5. One-Step Inference Strategy During inference, given a pre-contrast anchor 𝑧𝑝𝑟𝑒 and a requested target temporal phase 𝜏, the model evaluates the state at the maximum degradation timestep 𝑡𝑚𝑎𝑥 (effectively resulting in pre-contrast sequence with added noise 𝜎). To preserve the generative capacity and mimic realistic scanner textures while ensuring smooth temporal consistency across continuous samples of 𝜏, we inject a fixed, patient-level latent noise map 𝜖𝑓 𝑖𝑥𝑒𝑑 . The noisy input state is constructed as: 𝑧𝑖𝑛𝑝𝑢𝑡 = 𝑧𝑝𝑟𝑒 + (𝑤𝑡𝑚𝑎𝑥 ⋅ 𝜎𝑛𝑜𝑖𝑠𝑒 )𝜖𝑓 𝑖𝑥𝑒𝑑 , where 𝑤𝑡𝑚𝑎𝑥 represents the terminal degradation weight and 𝜎𝑛𝑜𝑖𝑠𝑒 controls the scale of the injected generative variance. The U-Net receives this noisy input state 𝑧𝑖𝑛𝑝𝑢𝑡 , concatenated with the clean anatomical anchor 𝑧𝑝𝑟𝑒 , the timestep 𝑡𝑚𝑎𝑥 , and the continuous temporal condition 𝜏. In a single predictive pass, the network outputs the predicted residual Δ𝑧, ̂ reconstructing the fully enhanced target endpoint as 𝑧̂ 𝑝𝑜𝑠𝑡 = 𝑧𝑝𝑟𝑒 + Δ𝑧. ̂ Finally, the synthesized latent representation is mapped back to the high-resolution pixel space via the frozen VAE decoder, yielding the synthetic image 𝑥̂ 𝑝𝑜𝑠𝑡 = (𝑧̂ 𝑝𝑜𝑠𝑡 ).

4. Implementation Details This section describes the datasets, evaluation metrics, and baseline methods. Additional information, including data preprocessing, and network hyperparameters, can be found in Appendix A.

4.1. Datasets We use two publicly available datasets as well as one private dataset for training and validating our method. MAMA-MIA: As detailed in (Garrucho et al., 2025), the MAMA-MIA dataset is a multi-center cohort of 1,506 patients featuring pre-treatment T1-weighted DCE breast MRIs. Compiled from the ISPY-1 (Newitt et al., 2016), ISPY-2 (Li et al., 2022), NACT (Newitt and Hylton, 2016), and Duke-Breast Cancer MRI (Saha et al., 2021) collections, the dataset demonstrates high technical diversity, including Page 5 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

(a) Intensity Histogram

(b) Peak Enhancement Time

(c) Spectral Difference

Figure 3: Characterization of the domain shift between the internal and external validation cohorts. (a) Log-scaled marginal intensity histograms of the post-contrast phase exhibit general macroscopic alignment, indicating relatively comparable global contrast distributions. (b) However, the probability density function of peak enhancement time (𝜏) reveals a severe temporal domain shift; the external cohort demonstrates significantly faster pharmacokinetic wash-in dynamics compared to the broader, delayed uptake of the internal dataset. (c) The 2D log power spectral difference between the cohorts highlights high-frequency spatial discrepancies, characteristic of differing scanner hardware, resolution, and acquisition protocols.

various scanner vendors (GE, Siemens, Philips), magnetic field strengths (1.5T and 3T) and multiple acquisition planes (axial, sagittal). More importantly, MAMA-MIA provides expert tumor segmentations as well as harmonized acquisition times extracted in a uniform format from all the collections. Following the provided protocol, we maintain the designated split, utilizing 300 cases for the internal validation set. Duke-Breast-Cancer-MRI (Saha et al., 2021): This dataset was collected from 2000-2014 and contains MRI scans of 922 patients acquired from two vendors, namely GE and Siemens, with 7 scanners of 1.5T and 3T magnetic field strengths from a single center in the United States. The acquisition plane is axial. MAMA-MIA includes 259 cases from this study. We incorporate the additional 663 cases from this collection in our training set. Karolinska Instituet: This private dataset was collected from 2012-2020 and contains MRI scans of 192 patients, acquired from a single vendor GE with 1.5T and 3T magnetic field strengths from a single center in Sweden. The acquisition plane is axial. We use this dataset to perform external validation and evaluate our model under domain shift.

4.2. Metrics We evaluate the synthetic post-contrast MRIs using eight complementary metrics: MSE quantifies pixel-wise reconstruction fidelity. Peak Signal-to-Noise Ratio (PSNR) measures the overall reconstruction quality relative to the reference image. Structural Similarity Index Measure (SSIM) (Wang et al., 2004) evaluates the preservation of local structural information, accounting for luminance, contrast, and texture consistency. LPIPS (Zhang et al., 2018) measures perceptual similarity using deep feature representations. To assess distribution-level realism, we report DINOv2-based Fréchet Distance (FID-DINOv2) (Heusel et al., 2017; Oquab et al., 2023). Specifically, we replace Inception features with self-supervised DINOv2 embeddings to better capture semantic and structural similarity in medical images. We also report the Fréchet Radiomics Distance (FRD) (Konz S. Joshi et al.: Preprint submitted to Elsevier

et al., 2026), computed from radiomic features extracted within the expert-annotated tumor masks. Unlike imagelevel metrics, FRD evaluates whether the generated images preserve clinically relevant tumor characteristics, including intensity statistics, texture, and shape-related descriptors, providing an assessment of radiomic fidelity. Furthermore, the quantitative evaluation of temporal dynamics in contrast synthesis remains largely absent from current literature. To address this gap and validate the temporal realism of the generated pharmacokinetic curves, we evaluate all methods using two targeted metrics:

Time-to-Peak Error (PTE): In clinical BI-RADS as-

sessment, the Time-to-Peak (TTP) is a critical diagnostic biomarker. To quantify temporal shift, we define the Peak Timing Error (PTE) as the mean absolute difference in seconds between the predicted and ground-truth time-to-peak across all temporal curves in the validation set. Let 𝑇 = {𝜏1 , 𝜏2 , … , 𝜏𝐾 } be the set of acquisition times, and 𝑐(𝜏) be the mean intensity of the tumor region at time 𝜏. The error is defined as: 𝑁 | 1 ∑ || | (𝑖) (𝑖) TTPE = |arg max 𝑐𝑝𝑟𝑒𝑑 (𝜏) − arg max 𝑐𝑔𝑡 (𝜏)| | 𝑁 𝑖=1 || 𝜏∈𝑇 𝜏∈𝑇 |

A value of zero indicates that the model correctly identified the peak enhancement phase, while non-zero values correspond to a misalignment of one or more DCE phases (typically multiples of ∼ 80–100 s, depending on the acquisition protocol).

Dynamic Time Warping (DTW): To measure shape

preservation independent of localized temporal distortions, we employ normalized Dynamic Time Warping (DTW)(Chung et al., 2025a) on the mean intensity value inside the ground truth mask. For a predicted sequence 𝑃 and ground truth sequence 𝐺, the optimal alignment path 𝜋 is computed to minimize the cumulative 𝐿1 cost, normalized by the Page 6 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table 1 Quantitative evaluation of pharmacokinetic synthesis across pixel-level accuracy, structural fidelity, perceptual realism, and temporal alignment metrics. Baseline refers to the results on the pre-contrast image. MSE, DTW and DTW-ROI are reported in 10−2 scale. pix2pix model does not consider acquisition time 𝜏 during training. Therefore, temporal metrics are not reported for this method. Best results are highlighted in bold, second-best are underlined. Statistical significance was evaluated using the Wilcoxon signed-rank test for all metrics, with the exception of PTE, which was assessed using a binomial sign test. Unless otherwise indicated, all reported values are statistically significant at 𝑝 < 0.001. Markers denote lower significance levels: * indicates 𝑝 < 0.01, ** indicates 𝑝 < 0.05, and † indicates a statistically insignificant difference. Method

MSE ↓

PSNR ↑

SSIM ↑

LPIPS ↓

FID ↓

FRD ↓

PTE ↓

DTW ↓

DTW-ROI ↓

18.24 (3.25) 20.84 (2.81) 20.36 (2.21) 18.60 (2.43) 20.90 (2.69) 21.61 (2.77)

0.70 (0.11) 0.71 (0.11) 0.73 (0.08) 0.58 (0.10) 0.73 (0.11) 0.71 (0.11)

0.19 (0.05) 0.19 (0.05) 0.18 (0.04) 0.26 (0.05) 0.21 (0.05) 0.18 (0.05)

166.68 184.83 274.56 205.92 181.66 154.56

5.12 5.34 4.50 7.40 5.57 4.98

265.30 (114.43) 41.24 (83.58) † – 101.46 (108.67) 74.47 (110.09) ∗ 44.57 (81.59)

2.91 (1.39) 0.84 (0.77) ∗∗ – 1.21 (0.83) 1.03 (0.86) 0.68 (0.56)

18.55 (5.46) 6.18 (4.91) – 8.84 (0.05) 5.97 (4.94) 3.76 (3.33)

17.20 (3.73) 19.75 (2.71) 19.55 (1.90) 17.68 (2.29) 20.12 (2.40) 20.32 (2.57)

0.69 (0.09)∗ 0.70 (0.09) 0.73 (0.06) 0.57 (0.09) 0.72 (0.08) 0.69 (0.09)

0.20 (0.04) 0.20 (0.04) 0.19 (0.03)† 0.26 (0.04) 0.21 (0.04) 0.19 (0.04)

220.34 257.74 345.82 254.07 244.23 226.26

5.03 4.93 4.57 5.44 5.22 4.78

237.24 (104.89) 43.59 (82.23)† – 106.12 (95.11) 40.29 (77.88)† 37.86 (70.13)

3.31 (1.53) 1.11 (0.84)† – 1.27 (0.65) 1.00 (0.62)† 0.98 (0.68)

19.44 (4.54) 7.32 (4.67)† – 9.60 (4.60) 8.25 (4.65) 6.28 (3.95)

Internal Validation Baseline U-Net pix2pix CCNet TeNCA Ours

1.93 (1.39) 1.01 (0.67) 1.05 (0.60) 1.60 (0.87) 0.99 (0.65) 0.84 (0.55)

External Validation Baseline U-Net pix2pix CCNet TeNCA Ours

2.52 (1.62) 1.26 (0.71) † 1.22 (0.55) 1.95 (1.01) 1.11 (0.56)† 1.08 (0.57)

sequence lengths: DTW(𝑃 , 𝐺) =

5. Results ∑ 1 min |𝑃 − 𝐺𝑗 | |𝑃 | + |𝐺| 𝜋 (𝑖,𝑗)∈𝜋 𝑖

A lower DTW score indicates that the generative model successfully learned the continuous pharmacokinetic manifold, synthesizing biologically plausible enhancement trajectories even if subtly shifted in absolute time.

4.3. Baseline Methods To rigorously evaluate our proposed framework, we compare it against a diverse set of baseline and state-ofthe-art techniques spanning multiple generative paradigms. First, we establish a standard U-Net (Ronneberger et al., 2015; Schreiter et al., 2024) which models sequential postcontrast sequences through different output channels. Second, we evaluate the adversarial pix2pix (Osuala et al., 2025) architecture. Because this model does not incorporate continuous acquisition time conditioning, we exclude it from temporal metric evaluations and assess it solely on spatial image fidelity. Third, we compare against CCNet (Osuala et al., 2024), a latent diffusion model conditioned on both pre-contrast anatomy and acquisition time. Because CCNet synthesizes contrast by sampling from random noise, it exhibits high stochasticity and frame-to-frame variance. Finally, we benchmark against Temporal Neural Cellular Automata (TeNCA) (Lang et al., 2025), a deterministic approach based on neural cellular automata that generates smooth, dense temporal trajectories, originally designed to overcome the stochastic limitations of diffusion models like CCNet. U-Net and the proposed method rely on residual prediction, while the remainder of the methods directly predict post-contrast phases.

S. Joshi et al.: Preprint submitted to Elsevier

This section presents results on in-domain dataset, examines out-of-domain generalization, and ablates different components of the proposed method. The quantitative results are summarized in Table 1.

5.1. In-Domain Performance At the pixel level, our method strictly dominates traditional reconstruction metrics, achieving the lowest MSE (0.84 x 10−2 ) and highest PSNR (21.61). While deterministic continuous-time models (TeNCA) and 2D adversarial networks (pix2pix) marginally outperform our framework on SSIM (0.73 vs. 0.71), this stems from inherent algorithmic biases: TeNCA favors smooth, regression-to-the-mean approximations, whereas pix2pix relies on hyper-realistic, albeit hallucinatory, textural synthesis (see Figure 4). This adversarial optimization also rewards pix2pix with lowest FRD of 4.50 on the test set, followed by our method at 4.98. Our formulation prioritizes true high-frequency structural coherence. This is quantitatively validated by our superior performance on LPIPS (0.18) and global feature distribution via FID-Dinov2 (154.56). Ultimately, the most significant advantage of our framework lies in its preservation of pharmacokinetic dynamics; our method drastically outperforms all baselines in temporal alignment, achieving a DTW of 0.68 (x 10−2 ) and a highly localized DTW-ROI of 3.76 (x 10−2 ). Our framework establishes a new state-of-the-art across the internal MAMA-MIA validation cohort, uniquely balancing spatial fidelity, perceptual realism, and temporal dynamics where existing baselines force a compromise. Dense temporal contrast synthesis results can be viewed in Supplementary Videos 1 - 3.

Page 7 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Figure 4: Qualitative comparison of synthetic contrast generation. This figure presents highly challenging cases where tumor margins are unclear in the pre-contrast image, complicating accurate contrast injection for all models. The images displayed are selected and correspondingly synthesized at the earliest acquisition time available in the ground truth. Additional comparisons are provided in Appendix Figure A.1. Columns from left to right: non-enhanced pre-contrast input, baseline methods (U-Net, pix2pix, CCNet, TeNCA), our method, and Ground Truth (GT). U-Net and TeNCA exhibit over-smoothing. pix2pix and CCNet demonstrate unpredictable behavior, capturing the structure in some instances but altering the underlying anatomy in others. Although our method also struggles to replicate exact tumor boundaries compared to the GT, it improves spatial fidelity and tumor localization. Note that the figure presents magnified views centered directly on the tumor for better visibility, Images are normalized to [0, 255] for display.

5.2. Out-of-Domain Performance To fully contextualize the performance on the external KI cohort, we analyze the fundamental domain shifts between the two institutions in Figure 3. First, the global post-contrast intensity distributions (Figure 3a) exhibit nearperfect alignment. This confirms the absence of macroscopic intensity covariate shifts; the physiological boundaries of contrast uptake remain strictly consistent across domains. This macroscopic alignment directly explains our model’s robust preservation of structural realism (LPIPS: 0.19) and overall pixel accuracy. However, density estimation of the acquisition priors (Figure 3b) exposes a severe temporal concept shift. The external KI protocol captures peak enhancement significantly earlier (𝜏 ≈ 280s) than the internal MAMA-MIA training distribution (𝜏 ≈ 450s). While our temporal alignment metrics exhibit a predictable degradation compared to S. Joshi et al.: Preprint submitted to Elsevier

our internal validation baseline (DTW: 0.68 vs. 0.98), our method still achieves the best absolute temporal performance (Table 1). Under this severe domain shift, the differences in global temporal alignment (PTE and DTW) between our approach, U-Net, and TeNCA do not reach statistical significance (𝑝 > 0.05). However, our framework maintains a statistically significant advantage over CCNet globally. Furthermore, within the localized tumor regions (DTWROI), our approach significantly outperforms both TeNCA and CCNet (𝑝 < 0.001), demonstrating robust preservation of tumor-specific kinetic fidelity despite the accelerated, outof-distribution injection protocol. Finally, while macroscopic statistics align, the 2D Spectral Difference Map (Figure 3c) reveals a highly directional high-frequency covariate shift. The distinct spectral lobes indicate that the external scanner possesses a fundamentally

Page 8 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table 2 Ablation study. Evaluating the impact of predicting subtraction, stochastic regularization and different loss components. Sub. refers to using subtraction between post-contrast and pre-contrast images as the prediction target, instead of the post-contrast image. 𝜖 refers to the added noise. 𝐿𝑃 𝐼𝑃 𝑆 and 𝐹 𝐹 𝐿 refer to the perceptual and spectral losses, respectively. Best results are bolded. MSE, DTW and DTW-ROI are reported in 10−2 scale. Sub.

𝜖

𝐿𝑃 𝐼𝑃 𝑆

𝐹 𝐹 𝐿

MSE ↓

PSNR ↑

SSIM ↑

LPIPS ↓

FID-Dinov2 ↓

FRD ↓

PTE ↓

DTW ↓

DTW-ROI ↓

0.93 (0.53) 1.06 (0.63) 0.89 (0.54) 0.86 (0.54) 0.84 (0.55)

20.98 (2.49) 20.53 (2.65) 21.31 (2.66) 21.47 (2.74) 21.61 (2.77)

0.71 (0.11) 0.68 (0.12) 0.71 (0.11) 0.71 (0.11) 0.71 (0.11)

0.18 (0.05) 0.22 (0.06) 0.18 (0.05) 0.18 (0.05) 0.18 (0.05)

162.83 357.52 152.14 145.06 154.56

5.17 7.81 5.10 5.05 4.98

45.35 (77.37) 48.94 (80.26) 47.87 (78.60) 45.65 (83.22) 44.57 (81.59)

0.94 (0.64) 1.70 (1.07) 0.80 (0.61) 0.77 (0.62) 0.68 (0.56)

3.48 (3.47) 3.54 (3.32) 4.02 (3.61) 3.65 (3.23) 3.76 (3.33)

✓ ✓ ✓ ✓ ✓

✓ ✓

✓ ✓ ✓

Figure 5: Qualitative evaluation of added ablation components. These images are normalized to range [0, 255] for display.

different microscopic noise floor and k-space reconstruction profile. Therefore, the degradation observed in deepfeature metrics like FID (226.26) on the external cohort is likely an artifact of localized hardware noise and disparate institutional statistics, rather than a failure of the network’s physiological synthesis.

5.3. Ablation Study 5.3.1. The Impact of Loss Components Our ablation study (Table 2) and corresponding qualitative visualizations (Fig. 5) demonstrate the critical, additive benefits of integrating structural, perceptual, and frequencybased constraints into the synthesis pipeline. Training with 𝑀𝑆𝐸 maps the macro-level contrast uptake, but visually results in a highly over-smoothed, blurry lesion devoid of internal heterogeneity. Introducing perceptual loss mitigates this blurring, recovering structural dimension and sharpening macroscopic tumor boundaries. The subsequent addition of stochastic regularization (+ noise) injects essential micro-textural realism into the tissue, visually harmonizing the generated distribution and quantitatively improving deep-feature semantic fidelity (optimizing FID-Dinov2 to 145.06). Finally, the integration of Fourier-domain guidance proves essential for recovering high-frequency morphological details. As observed visually, the Fourier constraint improves contrast in fine vascular structures surrounding the tumor area, closely aligning the generated output with the ground truth. Quantitatively, this comprehensive formulation yields the best overall pixel fidelity (MSE: 0.0084, PSNR: 21.61), radiomic texture preservation (FRD: 4.97), and temporal pharmacokinetic accuracy (DTW: 0.0068).

5.3.2. Effect of pre-contrast concatenation Table 3 evaluates the effect of the explicit morphological anchor (𝑧𝑝𝑟𝑒 ) on generative fidelity. The unanchored model S. Joshi et al.: Preprint submitted to Elsevier

Table 3 Ablation Study. Effect of Pre-conditioning. Best results are highlighted in bold. MSE, DTW and DTW-ROI are reported in 10−2 scale. Metric MSE ↓ PSNR ↑ SSIM ↑ LPIPS ↓ PTE ↓ FID-Dinov2 ↓ FRD ↓ DTW ↓ DTW-ROI ↓

w/o pre conditioning

w pre conditioning

1.04 (0.66) 20.67 (2.73) 0.71 (0.11) 0.18 (0.05) 58.94 (92.79) 141.87 5.67 1.08 (1.01) 4.45 (3.73)

0.84 (0.55) 21.61 (2.77) 0.71 (0.11) 0.18 (0.05) 44.57 (81.59) 154.56 4.98 0.68 (0.56) 3.76 (3.33)

achieves a lower FID-Dinov2 score (141.91 vs. 155.17) and a similar SSIM. By synthesizing both baseline anatomy and contrast through a single pathway, the unanchored network produces smooth deep-feature representations. However, explicitly conditioning the vector field on 𝑧𝑝𝑟𝑒 improves the clinical and temporal metrics. The anchored network focuses on contrast synthesis, achieving better pixel-level calibration (MSE: 0.0084), radiomic texture preservation (FRD: 4.97), and reduced trajectory error (DTW: 0.0068 vs. 0.0098). Furthermore, while both CCNet and our approach use stochastic noise to model output distributions, the precontrast anchor and the amount of injected noise dictate its specific function in image generation. CCNet initializes from pure noise conditioned on the pre-contrast image and acquisition time. As shown in Figure 6, while this conditioning preserves the macroscopic structure of the breast, different random noise result in varying tumor extents and anatomical hallucinations, across four predictions at a given acquisition time. In contrast, our method uses the pre-contrast image Page 9 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table 4 Quantitative downstream segmentation performance on the MAMA-MIA internal validation set. Results evaluate the performance on synthesized images against real pre-contrast (Baseline) and real post-contrast (Upper Bound) targets. Statistical significance was computed via a paired Wilcoxon signed-rank test (𝑝 < 0.01). Results that did not reach statistical significance are marked with an asterisk (∗). Method

Dice Coefficient ↑

HD95 ↓

Segmentation model trained on post-contrast Baseline U-Net pix2pix CCNet TeNCA Ours Upper Bound Figure 6: Qualitative comparison of multiple stochastic predictions. Columns refer to pre-contrast image, synthesized images from CCNet (Osuala et al., 2024) and our method respectively, and GT (real post-contrast image). Four rows correspond to four different random noise initializations. In CCNet, the injected noise leads to structurally different anatomical predictions, including the shape of the tumor region. In the proposed method, anatomy is strictly preserved and noise accounts for subtle contrast variations and inherent MRI noise. Note that the figure presents magnified views centered directly on the tumor for better visibility, These images are normalized to range [0, 255] for display.

as an anchor to retain the macroscopic integrity of both the anatomy and the pathology. Consequently, the injected noise models epistemic uncertainty, manifesting as realistic variations in contrast and standard MRI noise. The relevance of this property for uncertainty estimation is discussed in Section 5.5.

5.3.3. How to add noise for temporal continuity? Standard clinical practice typically acquires 4-5 postcontrast phases during MRI exam. Transitioning from this sparse sampling to modeling a dense temporal trajectory requires balancing spatial realism with temporal continuity. While injecting latent noise mimics uncertainty in contrast and scanner texture, independent random sampling across frames causes non-physiological flickering. We compared two noise injection strategies during inference, Independent random noise per phase restores spatial texture but disrupts temporal continuity. Patient-level noise holds a single noise map constant across all phases for a given patient, ensuring the temporal condition solely drives image changes. While we compare CCNet with independent noise to replicate the original setting of the paper, we also demonstrate that patient-level noise can improve the method to yield smooth trajectories. Specifically, this can be done because CCNet is based on denoising diffusion implicit models, which do not

S. Joshi et al.: Preprint submitted to Elsevier

0.17 (0.30) 0.44 (0.36) 0.22 (0.33) 0.45 (0.35) 0.30 (0.35) 0.51 (0.37) 0.68 (0.33)

160.69 (104.48) 82.20 (104.68)* 147.79 (108.24) 76.81 (102.32) 122.31 (110.69) 68.78 (99.62) 38.59 (77.73)

Segmentation model trained on pre-contrast Baseline U-Net pix2pix CCNet TeNCA Ours Upper Bound

0.49 (0.37) 0.56 (0.35) 0.44 (0.35) 0.44 (0.35) 0.51 (0.36) 0.60 (0.33) 0.63 (0.34)

71.48 (99.90) 53.05 (89.01) 70.10 (96.55) 78.51 (102.96) 62.59 (93.86) 43.38 (80.12) 47.46 (84.74)

require noise injection at each denoising step. In Supplementary Videos 4 - 7, we show the comparison for both CCNet and the proposed method with different noise strategies. We observe that patient-level noise results in smoother trajectories for both methods. In contrast to our method, which maintains reasonable consistency across different random noise seeds, CCNet exhibits a high degree of variance and becomes unreliable due to its inherent dependence on noise, reinforcing the observations made in the previous section.

5.4. Downstream Tumor Segmentation To evaluate the preservation of biological structures, we compare downstream tumor segmentation performance on real and generated scans (Table 4). Specifically, two 2D nnUNet (Isensee et al., 2021) models are trained on the MAMA-MIA training set, on (a) first post-contrast images only, and (b) pre-contrast images only. We evaluate downstream segmentation using the Dice coefficient and the 95th percentile Hausdorff Distance (HD95). These metrics are reported exclusively for the first postcontrast phase, as the ground truth masks were delineated on this specific timepoint. Training details, extended evaluations across all subsequent temporal sequences, as well as comprehensive results on the external validation cohort, are detailed in the Appendix B. Can a segmentation model trained on real post contrast images find the tumor on virtually contrast injected images?

Page 10 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Figure 7: Qualitative downstream segmentation results. Each row depicts a distinct patient case across real and synthesized images. Note that the figure presents magnified views centered directly on the tumor for better visibility, These images are re-normalized to range [0, 255] for display. The accompanying Kernel Density Estimation (KDE) maps (far right) illustrate the intensity distribution within each ground truth masks. The top row features a large, highly lobulated malignant lesion exhibiting limited signal on the pre-contrast scan. The pix2pix model artificially spikes pixel intensity (evidenced by the sharp green peak in the KDE plot) in and around the tumor, resulting in poor boundary discrimination. CCNet hallucinates high level of texture and overestimates the tumor region. While U-Net underestimates the tumor extent and TeNCA demonstrates moderate localization, our model achieves the closest alignment with the ground truth (GT) in both morphological boundary and intensity distribution. The middle row presents a diffuse, lower-contrast malignant lesion, visible merely as a darker hypointense region in the baseline scan but still, correctly identified by the pre-contrast trained model. Here, CCNet, and U-Net underestimates the region, pix2pix fails to provide meaningful enhancement, and TeNCA suffers from a spatial hallucination (false positive). Conversely, our model successfully recovers the broad enhancement profile, matching both the GT boundaries and the true KDE distribution. The bottom row depicts a distinct malignant focal mass. Although pix2pix achieves a KDE profile closer to the GT, it suffers from severe image degradation and, similar to TeNCA, massively overestimates the tumor boundary. While U-Net, CCNet and our proposed method accurately localize the mass, our method demonstrates superior statistical consistency with the true post-contrast KDE profile.

On the pre-contrast baseline, the post-contrast trained segmentation network effectively fails to localize the lesion (Dice: 0.17), establishing its critical reliance on enhancement. While the competing methods including pix2pix and TeNCA previously demonstrated high SSIM, their corresponding downstream segmentation performances only show a small improvement over the non-enhanced baseline, In contrast, our proposed method explicitly preserves biological boundaries, achieving the highest segmentation performance (Dice: 0.51, HD95: 68.78) and recovering the vast majority of the true clinical signal (Upper Bound Dice: 0.67). If the goal is to avoid contrast, why not simply rely entirely on pre-contrast scans? In other words, why synthesize at all? Segmentation networks trained directly on post-contrast data rely highly on the magnitude of pixel intensity (Joshi et al., 2024), as evident by low performance on pre-contrast data. Consequently, they also have higher sensitivity to variations in contrast dynamics, for instance, degree of contrast S. Joshi et al.: Preprint submitted to Elsevier

uptake. Conversely, as shown in Table 4, a network explicitly trained on pre-contrast data learns robust morphological and textural priors and already localizes tumor better in absence of contrast signal (Dice: 0.49). When we route these images through the generative frameworks, synthetic contrast surfaces the latent physiological structures, and all deterministic methods improve over the baseline. Our method performs the best with Dice of 0.60 and an HD95 of 43.38 (Upper bound: 0.63), thereby maximizing the algorithmic utility of raw scans while bypassing the physiological risks of actual gadolinium. Visual inspection (Figure 7) shows three representative examples where our model produces intensity distributions (pink) that closely mirror the true post-contrast profiles (green), and yields automated segmentation boundaries that conform to the ground truth. Additionally, our proposed framework exhibited the lowest absolute failure rate across all evaluated methods and training paradigms, yielding complete missed detections in only 29 of 300 cases. This performance was second only to the true post-contrast ground truth Page 11 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Figure 8: Clinical evaluation across varying case complexities. The figure contrasts a representative low-complexity case (Top) with a high-complexity case (Bottom) from the reader study. For both cases, the image panels display the pre-contrast baseline, the predicted synthetic time-series, the corresponding pixel-wise uncertainty maps (𝜎), and the ground truth (GT) time-series. The kinetic plots illustrate the enhancement dynamics, comparing the predicted mean tumor intensity against the ground truth over time. The consensus tumor characterization across synthetic and real images, and corresponding diagnostic impact are summarized for each case in text format.

(16/300). A detailed analysis of catastrophic segmentation failures (Dice = 0) is provided in the Appendix Section B.4 and qualitative examples are added to Appendix Figure B.2.

5.5. Clinical Reader Study To systematically evaluate the diagnostic equivalence, realism, and clinical utility of the synthetic contrast-enhanced MRI images, we conducted a comprehensive multi-part clinical reader study with four radiologists, with 5, 8, 13, and 15 years of experience respectively, evaluating randomly selected forty cases across three stratification variables: tumor size, tumor shape and enhancement pattern. The study design, reader profile, statistical tests and additional results are added in Appendix C. Readers can directly access platform here: https://smriti-joshi.github.io/reader-stu dy-contrast-synthesis/. Which method do the readers prefer? In a direct comparative evaluation against established temporal contrast synthesis baselines (UNet and TeNCA), readers demonstrated an strong preference for the proposed method. Specifically, out of 160 total evaluations (40 cases x 4 readers), our method was selected as the superior image in 83.8% of cases (134 votes). In contrast, the compared architectures tied for a second, with UNet and TeNCA each capturing 8.1% (13 votes) of the total preference. Statistical

S. Joshi et al.: Preprint submitted to Elsevier

analysis confirmed that this dominance was highly significant (𝑝 < 0.001) and not due to chance. This preference for our method was consistent across all four individual readers, while the two alternative methods were statistically indistinguishable from one another (𝑝 = 1.0). How do real and synthetic image characteristics compare? The reader study comparing synthetic and ground truth images demonstrated a robust overall mean agreement rate of 69.8% (median 66.7%). As seen in Table 5, coarse categorical judgments showed substantial GT-to-synthetic agreement, whereas margin and internal-enhancement assessments were only moderate. Readers noted a divergence in ordinal assessments such as image quality (56.2% agreement) and kinetic plausibility (48.8%), where real images maintained a clear scoring advantage. Nonetheless, for both criterion, synthetic and real images are scored between 2 and 3, suggesting that the readers found the kinetic curves plausible and the image quality acceptable on average (further demonstrated in Appendix Figure B.3). How does synthetic data impact patient care? In the side-by-side review of real and synthetic images, readers judged that relying on the synthetic image would produce a major change in clinical management in 30.0% of assessments. However, 34.4% of evaluations resulted in minor diagnostic deviations that did not alter the clinical Page 12 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table 5 Lesion characterization: GT vs. synthetic. Top: categorical tasks (Cohen’s 𝜅 with bootstrap 95% CI). Bottom: ordinal tasks (GT/synthetic means, mean difference Δ, signed rank-biserial 𝑟, Wilcoxon 𝑝). All 𝑝-values survive Benjamini– Hochberg correction. Categorical task

N

% agree

Cohen’s 𝜅

95% CI

Lesion Type Shape Margins Enhancement

160 119 153 160

86.2 89.9 77.1 68.8

0.72 0.73 0.45 0.43

[0.62, 0.83] [0.57, 0.86] [0.30, 0.59] [0.31, 0.55]

Ordinal task

GT

Synth

Δ

𝑟

𝑝

Kinetic Plausibility Image Quality

2.71 2.81

2.27 2.37

-0.43 -0.44

-0.62 -0.77

<0.001 <0.001

Table 6 Predictors of clinical impact. Proportional-odds ordinal logistic regression; standardized predictors. OR > 1 indicates higher odds of greater clinical impact. Predictor

OR

95% CI

𝑝

Perceived synthetic appearance Diagnostic complexity Image quality Reader experience Tumor size (large) Shape (irregular) Peak enhancement (late)

3.60 2.14 0.86 0.54 0.82 1.24 0.84

2.26–5.74 1.46–3.12 0.56–1.33 0.38–0.77 0.59–1.14 0.89–1.72 0.60–1.18

<0.001 <0.001 0.507 <0.001 0.235 0.197 0.318

management plan, and 35.6% resulted in no change at all. In other words, in 70.0% of the cases evaluated, replacing real images with synthetic ones would not negatively alter patient care. Further analysis using a proportionalodds ordinal logistic regression model identified perceived synthetic appearance (OR 3.60, 𝑝 < 0.001) and diagnostic complexity (OR 2.14, 𝑝 < 0.001) as significant independent predictors of greater clinical impact (i.e., a higher likelihood of altered patient management). Conversely, higher reader work experience served as a significant protective factor against these deviations (OR 0.54, 95% CI 0.38–0.77, 𝑝 < 0.001). Image quality was not independently predictive of clinical impact once perceived appearance and complexity were accounted for in the model (OR 0.86, 𝑝 = 0.51; Table 6). How does uncertainty map affect reader confidence? When presented with synthetic image and corresponding uncertainty maps, readers indicated that they would avoid or defer physical contrast injection in 64% of the evaluated cases. To help guide these critical management decisions, the integration of uncertainty served as a valuable, though imperfect, adjunct. This map was computed by adding random noise for 10 different iterations to produce slightly different outputs. As seen in Figure 8, the areas where the predictions agreed were displayed with high model confidence (black) while disagreement indicated uncertainty (yellow). In the reader study, map provided actionable guidance in 49% of evaluations overall, meaningfully shifting radiologist confidence by either validating trustworthy synthetic images (increasing confidence in 32% of cases) or S. Joshi et al.: Preprint submitted to Elsevier

flagging unreliable ones (decreasing confidence in 17%). These shifts directly influenced downstream actions: when the map decreased confidence, it acted as an effective safety mechanism, prompting readers to request true contrast injection in 88% of those instances. Conversely, when it increased confidence in unchanged assessments, 82% of readers proceeded without requesting real contrast. Furthermore, statistical analysis revealed that the map’s directional influence is significantly modulated by case complexity (Kendall’s 𝜏 = 0.218, 𝑝 = 0.003). While the map predominantly increased confidence in low-complexity cases (39% increase vs. 11% decrease), it functionally inverted in high-complexity scenarios, where it more frequently decreased confidence (41% decrease vs. 14% increase) to flag unreliable generations. Consequently, its overall informative rate peaked during these highly complex evaluations (55%). This utility also demonstrated a trending association with lesion type (𝑝 = 0.059), proving especially beneficial for challenging non-mass enhancement (NME) lesions, which achieved a 67% informative rate. However, these metrics must be interpreted with distinct clinical caution. Inter-reader agreement regarding the map’s effect was notably low (Fleiss’ 𝜅 = 0.02–0.10), indicating that reliance on the map remains highly subjective and reader-dependent. Furthermore, qualitative feedback highlighted specific limitations with the map’s reliability, noting critical instances where it incorrectly displayed high confidence despite the underlying synthetic image being clinically inaccurate. This susceptibility to overconfident errors underscores that while the uncertainty map is a helpful decision-support tool in the aggregate, it is not a definitive fail-safe, and must continue to be scrutinized carefully to prevent diagnostic missteps.

6. Discussion and Conclusion This work presents a conditioned latent transport framework for single-step DCE-MRI contrast synthesis. Quantitative evaluations demonstrate that the approach effectively balances spatial fidelity with precise pharmacokinetic temporal alignment, outperforming baseline models on internal data. Independent external validation evaluates the model’s adaptability to varying scanner noise profiles and temporal acquisition shifts. While the framework demonstrates promising resilience, the performance degradation under differing clinical protocols indicates that cross-institutional generalization remains an ongoing challenge Furthermore, the ablation studies evaluate the design choices required for temporal contrast synthesis. Specifically, anchoring to a pre-contrast image enforces physiological adherence, allowing the model to focus on temporal synthesis, individual loss components improve visual and temporal fidelity, while a fixed noise strategy maintains continuity in densely sampled predictions. In the downstream tumor segmentation task, our proposed framework outperforms competing state-of-the-art methods across two distinct evaluation strategies. Success on a network natively trained on post-contrast images confirms Page 13 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

that the synthesized contrast uptake accurately mirrors the ground truth. Superior performance on a pre-contrast trained network demonstrates that the synthesized data strictly preserves the underlying morphological and textural characteristics. These quantitative successes translate directly to clinical viability. In a blinded reader study conducted by expert breast radiologists, our synthesized images were preferred over competing baselines in approximately 84% of cases. Highlighting the framework’s potential to safely reduce contrast burden, the study further confirmed that substituting real post-contrast MRI with our synthetic counterparts would result in no major negative deviation to patient management in 70% of cases. Qualitative feedback from the reader study indicates clinical optimism regarding the potential of synthetic contrastenhanced MRI to reduce patient contrast burden. At the current maturity level of the state-of-the-art in contrast synthesis, conventional DCE-MRI remains indispensable for precise preoperative surgical planning. However, radiologists acknowledged that upon rigorous clinical validation, the technology shows promise as a diagnostic adjunct. Specifically, in their opinion, integrating synthetic contrast with high-resolution diffusion-weighted imaging (DWI) or mammography could elevate baseline screening confidence without necessitating immediate contrast administration. This aspect is not explicitely tested in this study and can make for an interesting future direction. In addition, further model refinements must address specific morphological and kinetic blind spots. Targeted improvements should prioritize enhancing margin sharpness in dense breast tissue, accurately reproducing internal tumor heterogeneity (e.g., central necrosis), and reliably capturing the true spatial extent of small lesions and non-mass enhancement (NME). Furthermore, it is important to increase robustness in failure detection through uncertainty estimation and other complementary methods. We add one failure case from the reader study in Appendix Figure C.6 which obtained identical diagnostic assessment, but the tumor is completely mislocalized. Strictly in terms of methodology, our work has several limitations. First, our current pipeline relies on qualitative assessments of pre- and post-contrast image registration, meaning that patient motion between acquisitions was not strictly quantified or corrected. Second, the training dataset exhibits a long-tailed temporal distribution, with sparse representation of early (< 100 s) and late (> 500 s) acquisition phases (Appendix Figure D.7). Consequently, the dataset contains fewer examples of tumors captured during the rapid wash-in period or exhibiting late-phase enhancement. We explored sampling techniques to address this imbalance, but they failed to improve synthesis quality in these lowdensity temporal regions without simultaneously degrading performance on the broader test set (Appendix Table D.7). Addressing this domain gap to ensure kinetic fidelity in lowdata regimes remains a critical area for future investigation. Third, in our current formulation, the model’s integration timestep 𝑡 and the physical acquisition time 𝜏 are fully decoupled. Given the highly non-linear nature of contrast S. Joshi et al.: Preprint submitted to Elsevier

uptake in malignant tumors, characterized by a rapid washin phase and a gradual wash-out phase, future work could explicitly couple these variables by utilizing a physiological pharmacokinetic (PK) model as the interpolation function. While a strict PK-driven trajectory could enforce rigorous physical priors in highly controlled, single-center protocols, it presents a unique challenge for generalized models. Because the MAMA-MIA dataset aggregates data from over 25 institutions with widely heterogeneous injection protocols and temporal resolutions, enforcing a single, idealized PK curve could act as an overly restrictive inductive bias. Future architectures must balance these physiological priors with the flexibility required to model multi-center clinical variance. Finally, owing to limited data availability, we did not explicitly evaluate synthetic contrast uptake in benign lesions. Because benign and malignant masses typically exhibit distinct pharmacokinetic enhancement profiles, extending this framework to reliably model and differentiate between these varying pathologies remains an essential objective for future clinical translation. In conclusion, the aim of this work is to contextualize the current progress of contrast synthesis breast DCE-MRI. We hope that our evaluation framework serves an anchor for future research, ensuring that progress in this domain remains transparent and clinically meaningful. We also hope that the community builds upon the gaps identified in the study to advance temporal contrast generation.

CRediT authorship contribution statement Smriti Joshi: Conceptualization, Data curation, Methodology, Investigation, Project administration, Formal Analysis, Software, Writing - original draft, Writing - review & editing. Apostolia Tsirikoglou: Data curation, Formal analysis, Writing - original draft, Writing - review & editing. Daniel M. Lang: Formal analysis, Writing - review & editing. Richard Osuala: Conceptualization, Writing - review & editing. Noah Márquez Vara: Software. Alejandro Guzman: Data curation. Grzegorz Skorupko: Software. Sebastian Ibarra Arregui: Writing - review & editing. Lidia Garrucho: Writing - review & editing. Akane Ohashi: Data curation, Validation. Dimitra Ntoula: Data curation, Validation. Eugen Divjak: Validation, Writing - review & editing. Oğuz Lafcı: Validation, Writing - review & editing. Jan C. Peeken: Writing - review & editing. Julia A. Schnabel: Writing - review & editing. Fredrik Strand: Resources, Writing - review & editing. Oliver Diaz: Supervision, Writing - review & editing. Karim Lekadir: Supervision, Resources, Funding acquisition, Writing - review & editing.

Acknowledgements This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 101057699 (RadioVal). Additionally, this work was partially supported by the project FUTURE-ES (PID2021-126724OB-I00) and project AIMED Page 14 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

(PID2023-146786OB-I00) from the Ministry of Science and Innovation of Spain. D.M.L. and J.A.S. received funding from HELMHOLTZ IMAGING, a platform of the Helmholtz Information and Data Science Incubator. G.S. received funding from the European Union’s Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart). The authors acknowledge the role of Gemini 3.1 Pro for enhancing clarity in the text, as well as Claude Sonnet 4.6 and DeepSeek V4 Flash for assistance in coding. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.

7. Ethical Approval Statement For the publicly available datasets used in this study, formal ethics committee approval and written informed consent were not required. For external validation data, ethical approval was obtained from the Swedish Ethical Review Authority (Etikprövningsmyndigheten) (Approval Number: 2020-00488, with subsequent amendment 2022-06777-02). The requirement for written informed consent was waived due to the retrospective nature of the study.

8. Declaration of competing interest Given their role as Medical Image Analysis Associate Editor, Julia Schnabel had no involvement in the peer-review of this article and has no access to information regarding its peer-review. Full responsibility for the editorial process for this article was delegated to another journal editor. The authors declare that they have no additional competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

9. Data Availability The MAMA-MIA dataset (Garrucho et al., 2025) and DUKE-Breast-Cancer-MRI dataset (Saha et al., 2021) are publicly available on Synapse2 and TCIA3 , respectively. The external validation KI dataset is private and the authors do not have permission to release it.

10. Code Availability The source code will be made publicly available upon acceptance.

Appendix A

Proposed Model

The implementation details and hyperparameter configuration of the proposed architecture is detailed in this section. Figure A.1 shows additional qualitative comparison between different methods. 2 https://www.synapse.org/Synapse:syn60868042/wiki/628716

3 https://www.cancerimagingarchive.net/collection/duke-breast-

A.1

Preprocessing

This work follows the same preprocessing pipeline established in TeNCA (Lang et al., 2025), with resampling to a uniform voxel spacing of 1 mm and intensity values are linearly rescaled between zero and one based on the 0.02 and 99.98 percentiles of the respective pre-contrast image. Additionally, images are cropped to patches of size 168×168.

A.2

Network and Hyperparameters

Variational autoencoder: The model is trained with a composite loss function with MSE weighted at 1.0, perceptual (LPIPS) loss weighted at 5.0, complemented by a small KL regularization term (1e-6), a Sobel gradient-based edgesharpness loss (0.5). Training runs for 200 epochs with a batch size of 16, a learning rate of 1e-4, and gradient clipping at 1.0. This yields a 4-channel latent space derived from a VAE with a 4× spatial compression and a latent scaling factor of 1.0259. Latent UNet: The core latent space architecture utilizes DiffusionModeUNet (from MONAI Generative (Pinaya et al., 2023)), with channel configuration of [128, 256, 512, 512], incorporating two residual blocks per resolution stage and spatial self-attention at the three highest-resolution levels. The noise level fixed at 𝜎 = 0.1. We employ classifierfree guidance, applying a conditioning dropout rate of 0.1 during training and a guidance scale of 1.0 during inference. Training is conducted over 200 epochs with a batch size of 8 using the AdamW optimizer (learning rate: 1 × 10−4 , weight decay: 3 × 10−4 ). To ensure training stability and efficiency, we apply gradient clipping at 1.0 and utilize mixed-precision training. The composite objective function comprises an MSE loss in the latent space (𝜆 = 1.0), a perceptual loss (𝜆 = 5.0), and a focal frequency loss (𝜆 = 50). Finally, an exponential moving average (EMA) of the network weights is maintained with a decay rate of 0.999.

A.3

The Impact of VAE reconstruction

Table A.1 evaluates the effect of the autoencoder (VAE) architecture on reconstruction fidelity (Phase 1) and temporal contrast synthesis (Phase 2). Using a pretrained autoencoder (Stability VAE, 8x downsampling)4 creates a spatial information bottleneck. Despite medical-domain finetuning, the 8x compression degrades high-frequency anatomical structures, resulting in a Phase 1 SSIM of 0.79. A custom VAE, trained from scratch with a 4x downsampling factor, preserves these details and improves structural reconstruction (SSIM: 0.92, PSNR: 30.80). This reconstruction performance directly impacts the Phase 2 generative temporal synthesis. The 4x latent space retains finer details, enabling the network to map the pharmacokinetic trajectory while maintaining spatial boundaries. This yields improved pixel fidelity (MSE: 0.84×10−2 ) and spatiotemporal alignment across the full image and isolated lesions (DTW: 0.68, DTW-ROI: 3.76). The 8x latent 4 https://huggingface.co/stabilityai/sd-vae-ft-mse

cancer-mri/

S. Joshi et al.: Preprint submitted to Elsevier

Page 15 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Figure A.1: Additional Qualitative Results for Synthetic Contrast Generation.

S. Joshi et al.: Preprint submitted to Elsevier

Page 16 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table A.1 Ablation Study. Comparing autoencoder reconstruction fidelity (Phase 1) and its downstream impact on temporal synthesis (Phase 2). Best results are highlighted in bold. MSE is reported in 10−2 scale. Metric

Stability VAE (Zero-Shot)

Stability VAE (Fine-tuned)

VAE 4x (Scratch)

Phase 1: Pure VAE Reconstruction MSE ↓ PSNR ↑ SSIM ↑ LPIPS ↓ FID-Dinov2 ↓

0.33 (0.21) 25.82 (3.25) 0.78 (0.07) 0.12 (0.02) 79.13

0.31 (0.20) 26.13 (3.27) 0.79 (0.07) 0.11 (0.02) 78.18

0.11 (0.08) 30.80 (3.14) 0.92 (0.04) 0.10 (0.03) 120.99

Phase 2: Synthesis MSE ↓ PSNR ↑ SSIM ↑ LPIPS ↓ FID-Dinov2 ↓ FRD ↓ DTW ↓ DTW-ROI ↓

– – – – – – – –

0.97 (0.60) 20.93 (2.67) 0.70 (0.11) 0.18 (0.05) 131.19 5.72 0.74 (0.65) 4.07(3.72)

0.84 (0.55) 21.61 (2.77) 0.71 (0.11) 0.18 (0.05) 154.56 4.98 0.68 (0.56) 3.76 (3.33)

space yields better perceptual deep-feature scores (FIDDinov2: 131.19 vs. 154.56), likely due to its closer alignment with the natural-image distributions expected by the pretrained metric. However, the 4x architecture performs better on temporal and localized metrics, including FRD (4.98 vs. 5.72).

limitation in the evaluation paradigm: the reference ground truth (GT) segmentation masks were clinically delineated based strictly on the first post-contrast sequence. As the contrast agent naturally diffuses into surrounding tissues during later acquisition phases, the physiological enhancement boundaries shift relative to this static, early-phase GT mask. Consequently, the multi-phase metrics reported here accurately reflect sustained temporal contrast identifiability, but serve only as a proxy for absolute boundary precision in late-phase dynamics.

B.3

B.4

Appendix B B.1

Segmentation

Training Details

For downstream segmentation, we trained a 2D network using the standard nnUNet framework (Isensee et al., 2021), utilizing both the pre-contrast and first post-contrast images of the MAMA-MIA training set. The automated configuration heuristic generated a deep 2D U-Net architecture tailored to the median image dimensions and spacing of the dataset. Preprocessing included resampling all modalities to an in-plane resolution of 0.703 × 0.703 mm (using third-order spline interpolation for image data and nearestneighbor for segmentation masks) alongside global Z-score normalization. Model optimization was executed on 256 × 256 pixel patches using a batch size of 49 and a batchaggregated Dice loss function.

B.2

Performance on all post contrast phases

Table B.2 extends the downstream segmentation evaluation presented in the main manuscript by reporting performance across all synthesized post-contrast sequences, rather than exclusively the first acquisition phase. This extended analysis provides strong evidence that our generative framework successfully synthesizes an identifiable and persistent contrast enhancement trajectory as temporal acquisition progresses. However, we note an expected slight degradation in the absolute metrics, including the real post-contrast Upper Bound, which decreases from a Dice of 0.68 to 0.62 on the post-contrast trained network. This is due to a structural S. Joshi et al.: Preprint submitted to Elsevier

External Validation

Table B.3 presents the downstream tumor segmentation performance evaluated on the independent, multiinstitutional external validation cohort. As anticipated, the absolute quantitative metrics are lower compared to internal validation, including the real post-contrast upper bound (Dice: 0.60), which reflects the inherent domain gap when applying a fixed segmentation network to external data. Despite this overall reduction in absolute performance, the relative hierarchical trends established in the internal validation remain consistent, where our proposed method maintained robust biological feature amplification, significantly outperforming all synthetic baselines (Dice: 0.43 and 0.46 with post-contrast and pre-contrast trained segmentation networks respectively).

Failure Cases

An analysis of complete segmentation failures (Dice = 0) further highlights the robustness of our method. Natively trained post-contrast networks proved sensitive to any deviation from post-contrast distribution, yielding 238 unique patient failures across all methods, compared to 150 for the robust pre-contrast network. As seen in Table B.4, our proposed framework exhibited the lowest catastrophic failure rate (46 cases), and even surpassing the real post-contrast ground truth (54 cases) when tested with pre-contrast segmentation network. Overall, we have the lowest failure rate across synthesis methods where 29/300 case were missed by both segmentation networks. We add representative failure cases of the proposed method in Figure B.2.

Appendix C

Clinical Reader Study

To systematically evaluate the diagnostic equivalence, realism, and clinical utility of the synthetic contrast-enhanced MRI images, we conducted a comprehensive multi-part reader study. Readers can directly access platform here: https://smriti- joshi.github.io/reader- study- contras t-synthesis/.

C.1 Study Design C.1.1 Selection of Cases To ensure a balanced evaluation, a subset of 40 cases was selected from the internal test set using a stratified random sampling approach. Stratification was based on three objective metrics extracted directly from the groundtruth segmentations: tumor volume, boundary sphericity, Page 17 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Figure B.2: Failure Cases in Downstream Tumor Segmentation Task.

and peak enhancement latency. To establish the phenotypic subgroups, each metric was dichotomized using a robust median split approach, where the median value across the entire test-set cohort served as the definitive cutoff threshold. Specifically, tumor volume was assessed using the maximal cross-sectional area in pixels, categorizing cases as Small S. Joshi et al.: Preprint submitted to Elsevier

(less than or equal to the median) or Large (greater than the median). Boundary sphericity was quantified using a standard circularity metric, mathematically defined as 4𝜋×Area , with cases classified as Irregular (less than or Perimeter2 equal to the median) or Regular (greater than the median). Finally, peak enhancement latency was measured as the Page 18 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table B.2 Downstream Segmentation Performance on Internal Validation Data (extended with evaluation on all sequences). Baseline and Upper Bound refer to inference on pre-contrast images and real post-contrast images, respectively. Statistical significance was computed via a paired Wilcoxon signed-rank test (𝑝 < 0.01). Results that did not reach statistical significance are marked with an asterisk (∗). Method

First Post-Contrast Dice ↑

HD95 ↓

All Post-Contrast Dice ↑

HD95 ↓

0.17 (0.30) 0.62 (0.32) 0.44 (0.35) 0.22 (0.33) 0.40 (0.30) 0.25 (0.32) 0.48 (0.36)

160.69 (104.48) 51.91 (79.03) 83.32 (98.23)* 147.79 (108.24) 84.67 (87.48) 145.04 (94.63) 75.48 (97.23)

0.49 (0.37) 0.63 (0.31) 0.57 (0.33) 0.44 (0.35) 0.46 (0.31) 0.53 (0.34) 0.60 (0.33)

71.48 (99.90) 41.33 (71.74) 49.98 (83.36) 70.09 (96.55) 67.66 (83.13) 59.52 (86.86) 44.26 (79.37)

Segmentation model trained on post-contrast Baseline Upper Bound U-Net pix2pix CCNet TeNCA Ours

0.17 (0.30) 0.68 (0.33) 0.44 (0.36) 0.22 (0.33) 0.45 (0.35) 0.30 (0.36) 0.52 (0.36)

160.69 (104.48) 38.59 (77.73) 82.20 (104.68)* 147.79 (108.24) 76.81 (102.32) 122.31 (110.69) 63.07 (95.85)

Segmentation model trained on pre-contrast Baseline Upper Bound U-Net pix2pix CCNet TeNCA Ours

0.49 (0.37) 0.63 (0.34) 0.56 (0.35) 0.44 (0.35) 0.44 (0.35) 0.51 (0.36) 0.60 (0.33)

71.48 (99.90) 47.46 (84.74) 53.05 (89.01) 70.10 (96.55) 78.51 (102.96) 62.59 (93.86) 43.38 (80.12)

Table B.3 Downstream Segmentation Performance on External Validation Data. Baseline and Upper Bound refer to inference on pre-contrast images and real post-contrast images, respectively. Statistical significance was computed via a paired Wilcoxon signed-rank test (𝑝 < 0.01). Results that did not reach statistical significance are marked with an asterisk (∗). Method

First Post-Contrast Dice ↑

HD95 ↓

All Post-Contrast Dice ↑

HD95 ↓

0.07 (0.21) 0.54 (0.33) 0.38 (0.31)* 0.09 (0.23) 0.33 (0.28) 0.19 (0.29) 0.40 (0.32)

201.92 (81.43) 67.61 (89.76) 83.30 (97.52)* 191.52 (89.96) 89.29 (81.12) 153.31 (97.06) 76.32 (91.76)

0.36 (0.35) 0.52 (0.33) 0.41 (0.33) 0.33 (0.31) 0.37 (0.31) 0.37 (0.30) 0.46 (0.32)

107.26 (110.60) 65.45 (86.69) 77.12 (94.96) 95.46 (107.05) 82.77 (86.70) 93.90 (87.15) 62.07 (90.57)

Segmentation model trained on post-contrast Baseline Upper Bound U-Net pix2pix CCNet TeNCA Ours

0.07 (0.21) 0.60 (0.35) 0.38 (0.33)* 0.09 (0.23) 0.36 (0.32) 0.19 (0.30) 0.43 (0.33)

201.92 (81.43) 53.55 (90.56) 84.19 (105.11)* 191.52 (89.97) 76.45 (101.29) 152.64 (107.67) 59.28 (90.86)

Segmentation model trained on pre-contrast Baseline Upper Bound U-Net pix2pix CCNet TeNCA Ours

(a) Paired Kinetic

0.36 (0.35) 0.51 (0.37) 0.39 (0.34) 0.33 (0.31) 0.34 (0.33) 0.41 (0.33) 0.46 (0.33)

107.26 (110.60) 76.18 (101.95) 88.55 (105.97) 95.46 (107.05) 93.74 (107.18) 80.45 (101.59) 65.86 (95.69)

(b) Paired Quality

Figure B.3: Overall composite caption describing the kinetic analysis, quality metrics, and case agreement.

S. Joshi et al.: Preprint submitted to Elsevier

Page 19 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table B.4 Analysis of catastrophic segmentation failures (Dice = 0). The table reports the absolute number of unique patient cases where the segmentation networks completely failed to localize the lesion, evaluated across both pre-contrast and postcontrast training paradigms. "Overlap" indicates the number of cases that failed simultaneously under both paradigms. Method

Pre-Contrast Network Failures

Post-Contrast Network Failures

Overlap (Both)

86 54 60 86 92 63 46

203 40 96 184 87 153 82

79 16 45 78 51 53 29

Pre-contrast Post-Contrast U-Net pix2pix CCNet TeNCA Ours

Inter-Reader Agreement: Fleiss' with 95% Bootstrap CI (GT vs Synthetic) 0.6 0.5

Fleiss'

0.4 =0.2 (slight) =0.4 (fair) =0.6 (moderate)

0.3 0.2

GT Synthetic

0.1 0.0

ion

Les

type

pe

Sha

Ma

s

rgin

rn

atte

nt p

me

nce

a Enh

etic

Kin

y

it sibil

plau

ge

Ima

lity

qua

Figure B.4: Inter-reader agreement on GT and synthetic images (Fleiss’ 𝜅 with bootstrap 95% CI) across tasks.

acquisition time in seconds required for the mean signal intensity inside the tumor mask to reach its maximum across all post-contrast phases, categorizing cases as Early (less than or equal to the median) or Late (greater than the median). This categorization yielded eight distinct phenotypic subgroups. To prevent evaluation bias toward any single, frequently occurring tumor presentation, exactly five cases were randomly sampled from each of the eight subgroups. The detailed distribution is provided in Table C.5.

C.1.2 Lesion Characterization and Image Quality In the first section, readers were presented with isolated images (either real or synthetic, presented in a blinded fashion) and asked to characterize the breast tumor and assess the overall image quality. The characterization was based on the standard lexicon, requiring readers to classify the lesion type as a mass, non-mass enhancement (NME), focus/foci, or indeterminate. For lesions identified as masses, readers further specified the shape (round, oval, or irregular) and margins (circumscribed, irregular, spiculated, or indeterminate). The enhancement pattern was categorized as homogeneous, heterogeneous, rim enhancement, dark internal septations, or not assessable. Beyond standard characterization, readers evaluated the kinetic plausibility of the enhancement using a 5-point scale ranging from 1 (implausible/non-physiologic pattern) to 5 (fully physiologic). Finally, readers rated the overall quality of the MRI image, independent of the tumor, on a 5-point Likert scale ranging from 1 (non-diagnostic due to severe S. Joshi et al.: Preprint submitted to Elsevier

artifacts) to 5 (excellent high-quality image with minimal artifacts).

C.1.3

Clinical Decision-Making and Diagnostic Equivalence In the second section, readers performed a side-byside comparative evaluation of the synthetic image against the corresponding real acquisition to determine diagnostic equivalence. Readers assessed the potential impact of the synthetic image on clinical decision-making, categorizing it as causing no change, a minor change (no impact on clinical management), or a major change (affecting clinical management or diagnosis). If a major change was indicated, readers specified the primary discrepancy, such as incorrect tumor localization, inaccurate extent/size estimation, implausible enhancement kinetics, image quality issues, or missed/false detections. Additionally, readers rated the inherent diagnostic complexity of the case (low, moderate, or high) and evaluated the perceived synthetic appearance of the generated image on a 4-point scale from "None" (fully natural) to "Strong" (clearly synthetic). During this phase, readers were also provided with an AI-generated anomaly map indicating epistemic uncertainty. Readers reported how this maps influenced their diagnostic confidence (increased, decreased, or no change) and selected their next logical clinical step if presented with this data in a real workflow (e.g., proceed normally, review with caution, or request standard contrast injection). C.1.4

Method Preference and Comparative Benchmarking In the final phase, the proposed generation method was benchmarked against alternative state-of-the-art methods for temporal contrast synthesis. Readers were presented with the actual MRI acquisition alongside images generated by competing models, UNet (Ronneberger et al., 2015) and TeNCA (Lang et al., 2025). They were tasked with selecting their preferred synthetic image based on a holistic assessment of tumor characterization accuracy, enhancement quality, and overall image fidelity. An optional free-text field was provided for readers to leave additional qualitative comments regarding their selection. C.1.5 Broader Implications Upon completion of the image-specific evaluations, readers participated in a final post-study survey designed to capture their broader perspectives on the current state and future utility of AI-based contrast synthesis. To assess the perceived maturity of the technology, readers were asked to evaluate how close the field is to solving the problem of contrast enhancement synthesis for clinical use, selecting from four levels of progress: very far (fundamental limitations remain), moderate progress (significant gaps remain), getting close (most cases are convincing), or nearly solved (clinically equivalent in most scenarios). The remainder of the survey consisted of open-ended, free-text questions aimed at gathering qualitative insights to Page 20 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table C.5 Stratified case selection from the internal test set based on tumor volume, boundary sphericity, and peak enhancement latency. Tumor Volume

Boundary Sphericity

Peak Enhancement

Large Large Large Large Small Small Small Small

Irregular Irregular Regular Regular Irregular Irregular Regular Regular

Early Late Early Late Early Late Early Late

Inter-reader 𝜅

Δ

0.72 0.73 0.45 0.43

0.43 0.31 0.31 0.12

+0.30 +0.41 +0.14 +0.31

Lesion Type Shape Margins Enhancement

GT inter-reader (Fleiss ) GT Synthetic (Cohen )

0.6 0.5 0.4 0.3 0.2 0.1 0.0

guide future technical development and clinical implementation. Readers were prompted to identify where current AIbased synthesis methods are most lacking and to specify which technical improvements they would prioritize (e.g., spatial resolution, temporal consistency, kinetic accuracy, or artifact reduction). Furthermore, the survey explored the readers’ clinical vision for the technology, specifically asking for their thoughts on integrating virtual contrast MRI with standard mammography for breast cancer screening workflows, as well as its potential utility when combined with other unenhanced MRI sequences (such as diffusionweighted imaging) for comprehensive diagnostic assessments. A final open-text field was provided to capture any remaining general feedback or observations regarding the study.

C.2

Participants

Four board-certified breast-imaging radiologists from three academic centers participated in the reader study. Subspecialty experience ranged from 5 to 15 years (median 10.5, values 5, 8, 13, 15). Each reader evaluated all 40 cases across all four sections, yielding 2,927 analyzable responses (overall item-level missingness 12.9%, predominantly the conditional and optional items).

C.3

Statistical analysis

Analyses were paired at the case×reader level wherever the design permitted. Categorical agreement between GT and synthetic characterization was quantified with Cohen’s 𝜅 (unweighted for nominal tasks, linear-weighted for ordinal tasks) and percent agreement; 95% confidence intervals were obtained by case-level bootstrap resampling (1,000 iterations). Ordinal ratings (kinetic plausibility, image quality) were compared with the Wilcoxon signed-rank test, with a signed rank-biserial correlation 𝑟 (computed as (𝑇+ − 𝑇− )∕(𝑇+ + S. Joshi et al.: Preprint submitted to Elsevier

33 58 34 24 31 28 52 39

0.7

Kappa

GT↔Synth 𝜅

Total Pool

5 5 5 5 5 5 5 5

Paired GT Synthetic agreement exceeds inter-reader agreement

Table C.6 Paired vs. inter-reader agreement. GT-to-synthetic Cohen’s 𝜅 exceeds GT inter-reader Fleiss’ 𝜅 for every categorical task. Task

Selected Cases

Lesion

Type

Shape

s

Margin

rn

t Patte

cemen

Enhan

Figure C.5: Paired agreement exceeds inter-reader agreement. For every categorical task, GT-to-synthetic Cohen’s 𝜅 (blue) exceeds GT inter-reader Fleiss’ 𝜅 (gray).

𝑇− ), negative values indicating lower synthetic ratings) as the effect size, and Kendall’s 𝜏 for association. Inter-reader agreement on GT and on synthetic images was summarized with Fleiss’ 𝜅 and bootstrap CIs. The clinical-impact outcome (no change / minor / major) was modeled with a proportional-odds ordinal logistic regression on standardized predictors (perceived synthetic appearance, diagnostic complexity, image quality, reader experience, and the three stratification axes); odds ratios with 95% CIs are reported. Method preference was tested globally with Cochran’s 𝑄 (the appropriate 𝐾=3 extension for related binary outcomes) and post-hoc with exact McNemar tests; preference against chance used the binomial test. All families of tests were corrected for multiple comparisons within each analysis phase using the Benjamini–Hochberg false-discovery-rate procedure at 𝑞 = 0.05; significance (𝛼) was set at 0.05.

C.4 Additional results C.4.1 Synthetic images add less variability than radiologists themselves For every categorical task, paired GT-to-synthetic Cohen’s 𝜅 exceeded the GT inter-reader Fleiss’ 𝜅 (Figure C.5). In other words, the discrepancy a synthetic image introduces relative to its own GT is smaller than the disagreement that already exists between radiologists reading the same real images.

Page 21 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

Figure C.6: Case demonstrating a major impact on clinical management. Although the synthetic lesion is characterized identically to the ground truth, its incorrect spatial localization fundamentally alters the diagnostic outcome.

C.4.2

Real images show have better image quality and kinetic plausibility Extending the analysis in the main manuscript on this theme, Figure B.3 show absolute number of cases binned in each category. Most real images are classified as having good kinetic curves, while for majority of synthetic curves are deemed plausible. Similarly, while the distribution of cases is similar in real and synthetic cases, the mean score is lower due to higher number of non-diagnostic and poor quality cases. C.4.3 Inter-reader agreement Inter-reader agreement was moderate for lesion type (Fleiss’ 𝜅 = 0.43 GT, 0.36 synthetic) and decreased for finer tasks, reaching near-zero for the two ordinal scales on both image types, indicating that kinetic-plausibility and quality ratings are inherently reader-subjective regardless of image origin. C.4.4 Failure mode Blinded uncertainty propagated to clinical concern: when a reader had answered “cannot determine” in the blinded phase, their later side-by-side impact rating was markedly higher indicating worse impact (mean 1.48 vs. 0.85; 65% vs. 24% “major”; Kendall’s 𝜏=0.25, 𝑝 < 0.001). Four cases drew unanimous (4/4) major-impact verdicts. Three were explained by visible quality degradation, but one (case 31) was rated indistinguishable from GT (quality Δ ≈ 0) yet drew unanimous major-impact verdicts because of spatial errors, specifically incorrect tumor localization and extent, corroborated by three independent readers (Case 31; Fig. C.6). This dissociation between perceptual realism and spatial reliability is the study’s cautionary finding: a synthetic image can look entirely natural while placing or sizing disease incorrectly.

Appendix D

MISCELLANEOUS

Training Data Distribution: Figure D.7 are referenced in the main manuscript Section IV to highlight lower sample sizes early and late in acquisition period. S. Joshi et al.: Preprint submitted to Elsevier

Figure D.7: Temporal distribution of training samples and ground truth contrast enhancement. The overlaid histogram (grey bars, right axis) illustrates the number of available training samples across the acquisition time 𝜏, highlighting the long-tailed distribution with sparse representation in early (< 100 s) and late (> 500 s) phases. The solid blue line and shaded region (left axis) denote the population mean and ±1 standard deviation of the ground truth enhancement trajectory, respectively.

Sampling: Table D.7 shows an ablation with and without using sampling strategy to increase representation in early (< 100𝑠) and late (> 500𝑠) acquisition period. This is referenced in main manuscript Section IV. We employ a stochastic temporal latent augmentation strategy during training. For a given fraction of the batch, the method generates synthetic training targets by sampling new acquisition times (𝜏𝑛𝑒𝑤 ) located within sparsely populated temporal region. Once 𝜏𝑛𝑒𝑤 is sampled, the algorithm identifies the two closest available real acquisition phases that bracket this time point. The corresponding images are encoded into the latent space, and a linear interpolation is performed between these two bracketing latents based on the relative temporal position of 𝜏𝑛𝑒𝑤 . The original post-contrast training target is then replaced with this newly interpolated latent and its corresponding timestamp. This data-driven augmentation provides the model with a smooth, continuous supervisory signal across the entire temporal range, enforcing physically plausible kinetic transitions in low-density regions.

References Alahari, A., Benstetter, M., 2017. Ema’s final opinion confirms restrictions on use of linear gadolinium agents in body scans. Medical Writing 26, 52. Arciszewska, Ż., Gama, S., Leśniewska, B., Malejko, J., NalewajkoSieliwoniuk, E., Zambrzycka-Szelewa, E., Godlewska-Żyłkiewicz, B., 2022. The translocation pathways of rare earth elements from the environment to the food chain and their impact on human health. Process safety and environmental protection 168, 205–223. Bansal, A., Borgnia, E., Chu, H.M., Li, J., Kazemi, H., Huang, F., Goldblum, M., Geiping, J., Goldstein, T., 2023. Cold diffusion: Inverting arbitrary image transforms without noise. Advances in Neural Information Processing Systems 36, 41259–41282. Bau, M., Dulski, P., 1996. Anthropogenic origin of positive gadolinium anomalies in river waters. Earth and Planetary Science Letters 143,

Page 22 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table D.7 Ablation Study. For a fraction of training samples, we replace the real post contrast objective with a synthetic sample where tau falls in a sparsely-populated region of the acquisition-time distribution. The post-contrast latent is linearly interpolated between the two nearest available phases (including pre at tau=0). This aims to give the model smooth, learnable supervisory signal across the entire 𝜏 range, especially in regions where real training data is scarce. Using this sampling strategy, the model still produces contrast and shows improvement over baseline. However, it is not able to learn physiologically accurate temporal kinetics, with marked increase in peak timing error. Best results are highlighted in bold. MSE, DTW and DTW-ROI are reported in 10−2 scale. Method

MSE ↓

PSNR ↑

SSIM ↑

LPIPS ↓

FID - Dinov2 ↓

FRD ↓

PTE ↓

DTW ↓

DTW-ROI ↓

18.19 (4.29) 21.61 (2.77)

0.71 (0.11) 0.71 (0.11)

0.18 (0.05) 0.18 (0.05)

161.65 154.56

5.29 4.98

60.79 (83.63) 44.57 (81.59)

0.74 (0.60) 0.68 (0.56)

3.76 (3.33) 3.76 (3.33)

MAMA-MIA (Internal Validation) With Sampling Without Sampling

0.88 (1.55) 0.84 (0.55)

245–255. Chen, Y., Yin, F., Chen, H., Wu, J., Li, C., 2026. Contrast-x: A multimodal contrast image synthesis benchmark and universal modality flow matching. URL: https://arxiv.org/abs/2601.15884, arXiv:2601.15884. Chung, W., Kang, J., Park, G.E., Kim, S.H., Nam, Y., 2025a. Synthesizing delayed-phase contrast-enhanced breast mr images from early-phase images using an iterative deep network, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 583–592. Chung, W., Kang, J., Park, G.E., Kim, S.H., Nam, Y., 2025b. Synthesizing delayed-phase contrast-enhanced breast mr images from early-phase images using an iterative deep network, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 583–592. Cotruvo, J., 2019. The chemistry of lanthanides in biology: recent discoveries, emerging principles, and technological applications. acs cent sci 5: 1496–1506. Dao, Q., Phung, H., Nguyen, B., Tran, A., 2023. Flow matching in latent space. arXiv preprint arXiv:2307.08698 . Dar, S.U., Yurt, M., Karacan, L., Erdem, A., Erdem, E., Cukur, T., 2019. Image synthesis in multi-contrast mri with conditional generative adversarial networks. IEEE transactions on medical imaging 38, 2375–2388. Darrah, T.H., Prutsman-Pfeiffer, J.J., Poreda, R.J., Ellen Campbell, M., Hauschka, P.V., Hannigan, R.E., 2009. Incorporation of excess gadolinium into human bone from medical contrast agents. Metallomics 1, 479– 488. Dhariwal, P., Nichol, A., 2021. Diffusion models beat gans on image synthesis, in: Advances in Neural Information Processing Systems, pp. 8780–8794. Dixon, J., Newlands, C., Dodds, C., Thomas, J., Williams, L., Kunkler, I., Bing, A., Macaskill, E.J., 2016. Association between underestimation of tumour size by imaging and incomplete excision in breast-conserving surgery for breast cancer. Journal of British Surgery 103, 830–838. Fan, X., Xu, L., Zheng, B., Zeng, X., Li, W., Huang, Z., 2025. Pre-to postcontrast medical image synthesis with outline-guide accelerate diffusion model. Neural Networks , 107851. Fonnegra, R.D., Hernández, M.L., Caicedo, J.C., Díaz, G.M., 2025. Synthesizing late-stage contrast enhancement in breast mri: A comprehensive pipeline leveraging temporal contrast enhancement dynamics. Computers in Biology and Medicine 196, 110660. Fraum, T.J., Ludwig, D.R., Bashir, M.R., Fowler, K.J., 2017. Gadoliniumbased contrast agents: a comprehensive risk assessment. Journal of Magnetic Resonance Imaging 46, 338–353. Gao, Q., Li, Z., Zhang, J., Zhang, Y., Shan, H., 2023. Corediff: Contextual error-modulated generalized diffusion model for low-dose ct denoising and generalization. IEEE Transactions on Medical Imaging 43, 745– 759. Garrucho, L., et al., 2025. A large-scale multicenter breast cancer dce-mri benchmark dataset with expert segmentations. Scientific Data 12, 453. Gonzalez, V., Sandelin, K., Karlsson, A., Åberg, W., Löfgren, L., Iliescu, G., Eriksson, S., Arver, B., 2014. Preoperative mri of the breast (pomb) influences primary treatment in breast cancer: a prospective, randomized, multicenter study. World journal of surgery 38, 1685–1693.

S. Joshi et al.: Preprint submitted to Elsevier

Grobner, T., Prischl, F., 2007. Gadolinium and nephrogenic systemic fibrosis. Kidney international 72, 260–264. Gweon, H.M., Cho, N., Han, W., Yi, A., Moon, H.G., Noh, D.Y., Moon, W.K., 2014. Breast mr imaging screening in women with a history of breast conservation therapy. Radiology 272, 366–373. Habijan, M., Krpić, Z., Perić, J., Galić, I., 2025. Diffusion models for mri reconstruction: A systematic review of standard, hybrid, latent and cold diffusion approaches. Electronics 15, 76. Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S., 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30. Ho, J., Jain, A., Abbeel, P., 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33, 6840–6851. Ibarra, S., del Riego, J., Catanese, A., Cuba, J., Cardona, J., Leon, N., Infante, J., Lekadir, K., Diaz, O., Osuala, R., 2025. Comparing conditional diffusion models for synthesizing contrast-enhanced breast mri from pre-contrast images, in: Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, Springer. pp. 226–236. Isensee, F., Jaeger, P.F., Kohl, S.A., Petersen, J., Maier-Hein, K.H., 2021. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18, 203–211. Isola, P., Zhu, J.Y., Zhou, T., Efros, A.A., 2017. Image-to-image translation with conditional adversarial networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1125–1134. Jiang, L., Dai, B., Wu, W., Loy, C.C., 2021. Focal frequency loss for image reconstruction and synthesis, in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 13919–13929. Joshi, S., Osuala, R., Garrucho, L., Tsirikoglou, A., del Riego, J., Gwoździewicz, K., Kushibar, K., Diaz, O., Lekadir, K., 2024. Leveraging epistemic uncertainty to improve tumour segmentation in breast mri: an exploratory analysis, in: Medical Imaging 2024: Image Processing, SPIE. pp. 292–300. Kanda, T., Fukusato, T., Matsuda, M., Toyoda, K., Oba, H., Kotoku, J., Haruyama, T., Kitajima, K., Furui, S., 2015. Gadolinium-based contrast agent accumulates in the brain even in subjects without severe renal dysfunction: evaluation of autopsy brain specimens with inductively coupled plasma mass spectroscopy. Radiology 276, 228–232. Kingma, D.P., Welling, M., 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 . Kishore Kumar, M., Ramanarayanan, S., Sadhana, S., Sarkar, A., Gayathri, M.N., Ram, K., Sivaprakasam, M., 2024. Dce-diff: Diffusion model for synthesis of early and late dynamic contrast-enhanced mr images from non-contrast multimodal inputs. in 2024 ieee, in: CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 5174–5183. Kong, J., He, Y., Xia, C., Ge, R., Li, S., 2026. Mri contrast enhancement kinetics world model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1288–1299. Konz, N., Chen, Y., Dong, H., Mazurowski, M.A., 2024. Anatomicallycontrollable medical image generation with segmentation-guided diffusion models, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 88–98.

Page 23 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Konz, N., Osuala, R., Verma, P., Chen, Y., Gu, H., Dong, H., Chen, Y., Marshall, A., Garrucho, L., Kushibar, K., et al., 2026. Fréchet radiomic distance (frd): A versatile metric for comparing medical imaging datasets. Medical Image Analysis , 103943. Krasznai, Z., Morisawa, M., Krasznai, Z.T., Morisawa, S., Inaba, K., Bazsáné, Z.K., Rubovszky, B., Bodnár, B., Borsos, A., Márián, T., 2003. Gadolinium, a mechano-sensitive channel blocker, inhibits osmosisinitiated motility of sea-and freshwater fish sperm, but does not affect human or ascidian sperm motility. Cell motility and the cytoskeleton 55, 232–243. Kuhl, C.K., Mielcareck, P., Klaschik, S., Leutner, C., Wardelmann, E., Gieseke, J., Schild, H.H., 1999. Dynamic breast mr imaging: are signal intensity time course data useful for differential diagnosis of enhancing lesions? Radiology 211, 101–110. Kuhl, C.K., Strobel, K., Bieling, H., Wardelmann, E., Kuhn, W., Maass, N., Schrading, S., 2017. Impact of preoperative breast mr imaging and mr-guided surgery on diagnosis and surgical outcome of women with invasive breast cancer with and without dcis component. Radiology 284, 645–655. Laczovics, A., Csige, I., Szabó, S., Tóth, A., Kálmán, F.K., Tóth, I., Fülöp, Z., Berényi, E., Braun, M., 2023. Relationship between gadoliniumbased mri contrast agent consumption and anthropogenic gadolinium in the influent of a wastewater treatment plant. Science of the Total Environment 877, 162844. Lang, D.M., Osuala, R., Spieker, V., Lekadir, K., Braren, R., Schnabel, J.A., 2025. Temporal neural cellular automata: Application to modeling of contrast enhancement in breast mri, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 604–614. Li, W., et al., 2022. I-SPY 2 breast dynamic contrast enhanced MRI trial (version 1) [data set]. The Cancer Imaging Archive URL: https: //doi.org/10.7937/TCIA.D8Z0-9T85, doi:10.7937/TCIA.D8Z0-9T85. Lindner, U., Lingott, J., Richter, S., Jakubowski, N., Panne, U., 2013. Speciation of gadolinium in surface water samples and plants by hydrophilic interaction chromatography hyphenated with inductively coupled plasma mass spectrometry. Analytical and bioanalytical chemistry 405, 1865–1873. Lingott, J., Lindner, U., Telgmann, L., Esteban-Fernández, D., Jakubowski, N., Panne, U., 2016. Gadolinium-uptake by aquatic and terrestrial organisms-distribution determined by laser ablation inductively coupled plasma mass spectrometry. Environmental Science: Processes & Impacts 18, 200–207. Lipman, Y., Chen, R.T., Ben-Hamu, H., Nickel, M., Le, M., 2022. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747 . Liu, X., Gong, C., Liu, Q., 2022. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003 . M, K.K., Ramanarayanan, S., S, S., Sarkar, A., Gayathri, M.N., Ram, K., Sivaprakasam, M., 2024. Dce-diff: Diffusion model for synthesis of early and late dynamic contrast-enhanced mr images from non-contrast multimodal inputs, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 5174–5183. Monticciolo, D.L., Newell, M.S., Moy, L., Niell, B., Monsees, B., Sickles, E.A., 2018. Breast cancer screening in women at higher-than-average risk: recommendations from the acr. Journal of the American College of Radiology 15, 408–414. Müller-Franzes, G., Huck, L., Bode, M., Nebelung, S., Kuhl, C., Truhn, D., Lemainque, T., 2024. Diffusion probabilistic versus generative adversarial models to reduce contrast agent dose in breast mri. European radiology experimental 8, 53. Müller-Franzes, G., Huck, L., Tayebi Arasteh, S., Khader, F., Han, T., Schulz, V., Dethlefsen, E., Kather, J.N., Nebelung, S., Nolte, T., et al., 2023. Using machine learning to reduce the need for contrast agents in breast mri through synthetic images. Radiology 307, e222211. Murata, N., Gonzalez-Cuyar, L.F., Murata, K., Fligner, C., Dills, R., Hippe, D., Maravilla, K.R., 2016. Macrocyclic and other non–group 1 gadolinium contrast agents deposit low levels of gadolinium in brain and bone tissue: preliminary results from 9 patients with normal renal function.

S. Joshi et al.: Preprint submitted to Elsevier

Investigative radiology 51, 447–453. Naval Marimont, S., Siomos, V., Baugh, M., Tzelepis, C., Kainz, B., Tarroni, G., 2024. Ensembled cold-diffusion restorations for unsupervised anomaly detection, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 243–253. Newitt, D., Hylton, N., 2016. Single site breast DCE-MRI data and segmentations from patients undergoing neoadjuvant chemotherapy (version 3) [data set]. The Cancer Imaging Archive URL: https://doi.org/10.793 7/K9/TCIA.2016.QHsyhJKy, doi:10.7937/K9/TCIA.2016.QHsyhJKy. Newitt, D., et al., 2016. Multicenter breast DCE-MRI data and segmentations from patients in the I-SPY 1/ACRIN 6657 trials. The Cancer Imaging Archive URL: https://doi.org/10.7937/K9/TCIA.2016.HdHpgJL K, doi:10.7937/K9/TCIA.2016.HdHpgJLK. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al., 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 . Osuala, R., Joshi, S., Tsirikoglou, A., Garrucho, L., Pinaya, W.H., Lang, D.M., Schnabel, J.A., Diaz, O., Lekadir, K., 2025. Simulating dynamic tumor contrast enhancement in breast mri using conditional generative adversarial networks. Journal of Medical Imaging 12, S22014–S22014. Osuala, R., Lang, D.M., Verma, P., Joshi, S., Tsirikoglou, A., Skorupko, G., Kushibar, K., Garrucho, L., Pinaya, W.H., Diaz, O., et al., 2024. Towards learning contrast kinetics with multi-condition latent diffusion models, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 713–723. Partridge, S.C., Stone, K.M., Strigel, R.M., DeMartini, W.B., Peacock, S., Lehman, C.D., 2014. Breast dce-mri: influence of postcontrast timing on automated lesion kinetics assessments and discrimination of benign and malignant lesions. Academic radiology 21, 1195–1203. Parvaiz, M.A., Yang, P., Razia, E., Mascarenhas, M., Deacon, C., Matey, P., Isgar, B., Sircar, T., 2016. Breast mri in invasive lobular carcinoma: a useful investigation in surgical planning? The breast journal 22, 143– 150. Pesapane, F., Sorce, A., Battaglia, O., Mallardi, C., Nicosia, L., Mariano, L., Rotili, A., Dominelli, V., Penco, S., Priolo, F., et al., 2025. Contrast agents in breast mri: State of the art and future perspectives. Biomedicines 13, 829. Peterson, M.S., Gegios, A.R., Elezaby, M.A., Salkowski, L.R., Woods, R.W., Narayan, A.K., Strigel, R.M., Roy, M., Fowler, A.M., 2023. Breast imaging and intervention during pregnancy and lactation. Radiographics 43, e230014. Pinaya, W.H., Graham, M.S., Kerfoot, E., Tudosiu, P.D., Dafflon, J., Fernandez, V., Sanchez, P., Wolleb, J., Da Costa, P.F., Patel, A., et al., 2023. Generative ai for medical imaging: extending the monai framework. arXiv preprint arXiv:2307.15208 . Pinaya, W.H., Tudosiu, P.D., Dafflon, J., Da Costa, P.F., Fernandez, V., Nachev, P., Ourselin, S., Cardoso, M.J., 2022. Brain imaging generation with latent diffusion models, in: MICCAI workshop on deep generative models, Springer. pp. 117–126. Qi, M., Shi, R., Cai, Y., He, L., Wang, W., Ma, L., 2025. A boundaryaware cold-diffusion model for electron microscopy segmentation, in: International Conference on Medical Image Computing and ComputerAssisted Intervention, Springer. pp. 13–23. Rogosnitzky, M., Branch, S., 2016. Gadolinium-based contrast agent toxicity: a review of known and proposed mechanisms. Biometals 29, 365–376. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B., 2022. Highresolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695. Ronneberger, O., Fischer, P., Brox, T., 2015. U-net: Convolutional networks for biomedical image segmentation. URL: https://arxiv.org/abs/1505 .04597, arXiv:1505.04597. Saha, A., et al., 2021. Dynamic contrast-enhanced magnetic resonance images of breast cancer patients with tumor locations [data set]. The Cancer Imaging Archive URL: https://doi.org/10.7937/TCIA.e3sv-re9 3, doi:10.7937/TCIA.e3sv-re93.

Page 24 of 25

Dense Temporal Contrast Synthesis via Conditioned Latent Transport Saslow, D., Boetes, C., Burke, W., Harms, S., Leach, M.O., Lehman, C.D., Morris, E., Pisano, E., Schnall, M., Sener, S., et al., 2007. American cancer society guidelines for breast screening with mri as an adjunct to mammography. CA: a cancer journal for clinicians 57, 75–89. Schreiter, H., Eberle, J., Kapsner, L.A., Hadler, D., Ohlmeyer, S., Erber, R., Emons, J., Laun, F.B., Uder, M., Wenkel, E., et al., 2024. Virtual dynamic contrast enhanced breast mri using 2d u-net architectures, in: Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, Springer. pp. 85–95. Selvi, V., Nori, J., Meattini, I., Francolini, G., Morelli, N., Di Benedetto, D., Bicchierai, G., Di Naro, F., Gill, M.K., Orzalesi, L., et al., 2018. Role of magnetic resonance imaging in the preoperative staging and work-up of patients affected by invasive lobular carcinoma or invasive ductolobular carcinoma. BioMed research international 2018, 1569060. Shen, G., Li, M., Farris, C.W., Anderson, S., Zhang, X., 2024. Learning to reconstruct accelerated mri through k-space cold diffusion without noise. Scientific Reports 14, 21877. Tang, R., Zhang, X., Guo, P., Qiang, X., 2025. Residual pre-training assisted cold diffusion for denoising low-dose computed tomography images, in: Eighth International Conference on Artificial Intelligence and Pattern Recognition (AIPR 2025), SPIE. pp. 1127–1135. Tur, Y., Stojkovic, M., Bagci, U., 2026. Wfm: 3d wavelet flow matching for ultrafast multi-modal mri synthesis, in: Medical Imaging with Deep Learning. Uematsu, T., Yuen, S., Kasami, M., Uchida, Y., 2008. Comparison of magnetic resonance imaging, multidetector row computed tomography, ultrasonography, and mammography for tumor extension of breast cancer. Breast cancer research and treatment 112, 461–474. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017. Attention is all you need. Advances in neural information processing systems 30. Wang, J., Reynaud, H., Erick, F.X., Kainz, B., 2025. Ctflow: Videoinspired latent flow matching for 3d ct synthesis, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6750– 6758. Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P., 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13, 600–612. Yan, P., Li, M., Zhang, J., Li, G., Jiang, Y., Luo, H., 2024a. Cold segdiffusion: A novel diffusion model for medical image segmentation. Knowledge-Based Systems 301, 112350. Yan, P., Li, M., Zhang, J., Li, G., Jiang, Y., Luo, H., 2024b. Cold segdiffusion: A novel diffusion model for medical image segmentation. Knowledge-Based Systems 301, 112350. Zaman, F.A., Jacob, M., Chang, A., Liu, K., Sonka, M., Wu, X., 2024. Surf-cdm: Score-based surface cold-diffusion model for medical image segmentation, in: 2024 IEEE International Symposium on Biomedical Imaging (ISBI), IEEE. pp. 1–5. Zhang, J., Liu, G., Chen, J., Cheng, Y., 2025. Multi-scale adaptive residual cold diffusion model for low-dose ct denoising. Expert Systems with Applications 294, 128817. Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O., 2018. The unreasonable effectiveness of deep features as a perceptual metric, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595. Zhang, T., Han, L., D’Angelo, A., Wang, X., Gao, Y., Lu, C., Teuwen, J., Beets-Tan, R., Tan, T., Mann, R., 2023. Synthesis of contrast-enhanced breast mri using t1-and multi-b-value dwi-based hierarchical fusion network with attention mechanism, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 79–88.

S. Joshi et al.: Preprint submitted to Elsevier

Page 25 of 25

Record · ID 422296 · SHA-256 84e62b2a71d1fe7d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.