ConceptioArchivearXiv CS
arXiv CSopen access

A novel unsupervised machine learning strategy to handle multimodal cardiac PET/MRI data

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

A Novel Unsupervised Machine Learning Strategy to Handle Multimodal Cardiac PET/MRI Data Brunnhilde Ponsia,b , Thomas Carliera,b , Lara Marteaua,c , Aurélien Monnetd , Thomas Eugènea,b , Jean-Michel Serfatya,e , Nicolas Pirioua,c , Hatem Neciba,b,∗ a Nantes Université, CHU Nantes, F-44000 Nantes, France b CRCI2NA, INSERM UMR 1307, Nantes, France c Nantes Université, CHU Nantes, INSERM, Cardiology Department, INSERM UMR 1307, CIC 1413, l’institut du Thorax, F-44000 Nantes, France d Siemens Healthineers France, Courbevoie, France e Nantes Université, CHU Nantes, INSERM, Radiology Department, l’institut du Thorax, F-44000 Nantes, France

Abstract

arXiv:2607.13936v1 [cs.CV] 15 Jul 2026

Arrhythmogenic left ventricular cardiomyopathy is a genetic myocardial disease difficult to diagnose due to the lack of gold standard criteria. Simultaneous PET/MR imaging, combined with multiparametric quantitative analysis, could facilitate the identification of different profiles related to the phenotype and progression of cardiomyopathy. This preliminary study focuses on a methodological strategy for dealing with PET/MRI data, including inter-patient data linkage and regional analysis. Two-step clustering was applied to T1 and T2 maps, LGE, and 18F-FDG-PET images of 99 patients genetically diagnosed with arrhythmogenic left ventricular cardiomyopathy. Each patient’s images were independently z-scored and summed into a single volume, which was clustered into supervoxels. Thirty-two inter-patient groups of supervoxels were obtained by spectral clustering. An "abnormality" score was assigned to each cluster and modality, and used to visualise abnormal regions likely associated with disease. They enabled the generation of automated textual and bullseye health reports for each patient, which were compared with cardiac imager assessments using balanced accuracy in repeated nested cross-validation. This approach was further validated on a larger cohort of 167 numerical phantoms. The reports generated by clustering accurately identified most of the cardiac physicians’ observations (BA = 0.76 ± 0.04 in repeated nested cross-validation on patients, and BA ≥ 0.8 on phantoms). Furthermore, the identified abnormal clusters closely matched their visual observations, facilitating the identification of varying degrees of fibrosis or inflammation on the images. This approach enables a more systematic handling of multimodal PET/MRI data to characterise myocardial heterogeneity in arrhythmogenic left ventricular cardiomyopathy patients. Keywords: PET/MRI, Arrhythmogenic Cardiomyopathy, Clustering, Multimodal data, Patient profiles

1. Introduction Arrhythmogenic cardiomyopathy (ACM) is an inherited myocardial disease characterized by fibrofatty replacement of the ventricular myocardium. Since 2010, it has emerged that ACM, which was originally believed to affect only the right ventricle (RV) can also affect the left ventricle (LV), either completely or partially with isolated LV fibrosis and normal LV function. In 2020, the Padua criteria were introduced to redefine ACM and its diagnostic criteria [1, 2]. These criteria propose a multiparametric approach to diagnosis, including specific criteria for arrhythmogenic left ventricular cardiomyopathy (ALVC). They emphasise the importance of cardiac MRI (CMR), which is now considered a necessary diagnostic tool as it allows tissue characterization by late-gadolinium enhancement (LGE) as a major ACM diagnostic criterion [2]. LGE techniques are extremely valuable in assessing myocardial fibrosis, but they are limited by the fact that they are qual∗ Corresponding author

Email address: [email protected] (Hatem Necib)

itative, which reduces their usefulness as an inter-patient diagnostic tool. Additionally, they fail to address diffuse fibrosis in the cardiac muscle, which is an early indicator of disease that could be mitigated if detected early. Therefore, using alternative sequences such as T1 and T2 mapping is beneficial, as it provides additional quantitative information on the presence of diffuse fibrosis or edema [3]. Although CMR is the reference diagnostic tool for identifying myocardial lesions, it is often inadequate for detecting inflammation outside the acute phase of cardiomyopathy [4]. However, some patients with ACM may exhibit myocardial inflammation, which is a poor prognostic marker for cardiomyopathy, although treatment can improve event-free survival. Studies have shown the value of FDG-PET in identifying markers of myocardial inflammation, which is prevalent in a significant proportion of patients with ALVC [4, 5]. These studies demonstrated the complementary role of FDG-PET compared with cardiac MRI, highlighting the need for further investigations into the impact of using PET to characterize and diagnose ACM [6]. Then, while CMR provides valuable anatomical insights and

information on fibrosis and edema, FDG-PET can bring complementary functional information. Acquiring PET and MRI simultaneously using a PET-MRI system thus has the potential to significantly improve ACM diagnosis. Multi-parametric PET-MRI may deepen our understanding of ACM, providing a more precise and straightforward diagnosis. It could also lead to the identification of different patient profiles with advanced cardiomyopathy, enabling the development of targeted therapeutic strategies that will be unique to each patient. In particular, double-step clustering approaches based on unsupervised machine learning techniques applied in two successive stages have already been successfully utilized in several oncology studies for prediction, prognosis, and tumor segmentation [7, 8, 9, 10]. Supervoxel clustering offers a balance between data analysis methods disregarding spatial context by averaging and pixel-wise analyses, which are both computationally costly and vulnerable to noise and registration inaccuracies [9, 10]. In this study, we propose a two-step clustering approach to handle multimodal PET-MRI volumes from ALVC patients. This method automatically identifies all possible combinations of high and low states across the multiple modalities and detects the corresponding regions of interest. The approach involves grouping voxels from each patient into supervoxels, followed by inter-patient clustering of the resulting supervoxels. A key innovation of our method is the introduction of a scoring system that not only characterizes cluster-level abnormalities but also bridges the gap to clinical practice through automated report generation. These clusters have the potential to reveal new ALVC biomarkers that could help determine new phenotypes and identify different patient profiles. The method was quantitatively evaluated on patient data using a nested cross-validation framework. It was further validated on phantoms and illustrated through a representative patient example.

Figure 1: Study flowchart.

Figure 2: Workflow of the two-step multimodal PET-MR imaging clustering.

2. Method

Definition of the region of interest (ROI): For each patient, a volume of interest (VOI) was manually defined on the T2 or T1 map. This VOI corresponds to a rectangular volume surrounding the entire heart, as in Fig. 3 Step a, and was used for the resampling and registration steps.

2.1. Database The database comprised data from 115 patients who were diagnosed with an ACM genetic variant. All patients were enrolled in the prospective trial "CharACTPET-MR" (NCT05450783). PET-MRI data were acquired using the Siemens PET-MR mMR Biograph. Analysis was performed using quantitative T1 and T2 maps, PET images with attenuation correction (PET-AC), and LGE-PSIR short-axis acquisitions. As described in Fig. 1, the final cohort comprised 99 patients following the exclusion of those with missing or poorquality data.

Data resampling: All other sequences were resampled in all 3 dimensions on the heart VOI with BSpline interpolation (Fig. 3 Step b). Segmentation of the left ventricle (LV): The left ventricle was semi-automatically segmented on the LGE images (Fig. 3 Step c) by a cardiologist on the Circle software by CVI [13]. Segmentations of the epicardium and endocardium were then exported and merged into a single LV segmentation.

2.2. Workflow of the method The workflow of this two-step clustering method is summarized in Fig. 2 by eight steps from A to G2.

Data registration: Although all MRI sequences were acquired at the same time point in the cardiac cycle, misalignment between MRI images can occur due to cardiac, respiratory, and patient movement. Additionally, since PET images are averaged over multiple cardiac cycles, they are not always perfectly

2.2.1. Step A. Data pre-processing and preparation Data was pre-processed (Fig. 2 Step A) for each patient individually using the 3D Slicer software [11, 12], as described in the next paragraphs and summarized in Fig. 3. 2

2.2.3. Step C. Removal of supervoxels with medically impossible relaxation times Both longitudinal (T1) and transversal (T2) relaxation times have been well-established in healthy myocardium [16]. Some studies have also investigated elevated T1 and T2 values in diseased myocardium. Knowledge of these values makes it possible to establish thresholds above or below which relaxation times are deemed implausible in the myocardium. Such values could be attributed to factors such as fat or MRI artifacts, rendering them unsuitable for consideration in our study. Consequently, extreme (minimum and maximum) values were defined for T1 and T2, and any supervoxel with T1 or T2 values exceeding one of these thresholds was excluded from subsequent analyses (Fig. 2 Step C). Plausible T1 and T2 values for the myocardium were defined as follows: Figure 3: Workflow of the data pre-processing steps (Fig. 2 Step A).

aligned with MRI images. All modalities were then registered on the LGE images (Fig. 3 Step d), either using the Elastix module [14] implemented in 3D Slicer or via manual registration if registration failed. All registrations were validated by a senior radiologist.

1000 ms ≤ T1myocardium ≤ 1800 ms

(1a)

20 ms ≤ T2myocardium ≤ 80 ms

(1b)

2.2.4. Step D. Data transformation As data in its original form is not always optimal for statistical analysis, it is common practice to re-express the data in a modified format that is more suitable for analysis. These modified representations can exhibit characteristics such as increased dispersion and improved symmetry, or simply become more interpretable and manipulable. Power transformations, in particular, are widely used in literature due to their simplicity and favorable properties, such as continuity, differentiability, and preservation of data rank [17]. Of these transformations, the Box-Cox transformation [18] is widely used to normalize feature distributions by reducing skewness and stabilizing variance. By making the data more Gaussian-like, it improves the behavior of distance-based metrics and enhances the performance of clustering algorithms, ultimately leading to better-defined clusters [10]. The Box–Cox transformation was therefore applied to the pooled supervoxels from all patients in an adapted version that included data with negative values (Fig. 2 Step D). The modified form is expressed as follows [18]:  (y+λ )λ1 −1   if λ1 , 0  λ2 1 (λ) (2) y =  log(y + λ2 ) if λ1 = 0

2.2.2. Step B. First clustering: Intra-patient clustering, from voxels to supervoxels A first clustering phase (Fig. 2 Step B) was used for each patient to group adjacent and similar voxels into supervoxels. This reduced registration issues, computational complexity, and aberrant voxel values. This initial clustering was applied on a new ROI surrounding the LV, as in the example of Fig. 4. The mean value of each of the four PET or MRI modalities within this ROI was subtracted, and the result was divided by the standard deviation to ensure equal contribution from all modalities. These z-scores were then summed to create a single cumulative “Z-volume”. Supervoxels were extracted from this Z-volume using SLIC (a k-means-based algorithm) [15]. The chosen parameters are described in Appendix A. The LV segmentation mask was applied to select only supervoxels containing at least 20% of their volume within the LV, thereby eliminating those caused by registration errors (oversegmentation). Each supervoxel was assigned T1, T2, LGE and PET values corresponding to their median pixel values within the LV. All data outside the LV segmentation was masked. While the T1 and T2 values remained quantitative, only the zscore values of each supervoxel were saved for the LGE and PET values. In addition, the position of each supervoxel within the 17 cardiac segments was recorded. Where a supervoxel contained voxels spanning several cardiac segments, the segment containing the majority of the supervoxel’s voxels was selected. Finally, standard data augmentation techniques (rotations and flips) were used on steps A and B to reduce the sensitivity of the method to imaging variations as described in Appendix B.

where y is the combined supervoxels data before transformation, y(λ) is the transformed data, λ2 is chosen so that yi + λ2 > 0 for all yi ∈ y , and λ1 is defined to maximize the log-likelihood, to approach a normal distribution of the transformed data. We defined λ1 and λ2 independently for each modality, on the combined supervoxels data of all the patients, and we took = |min y(mod) | + 10−5 to ensure positivity in the logarithm. λ(mod) i 2 2.2.5. Step E. Second clustering: Inter-patient clustering of the supervoxels Spectral clustering was used on Box-Cox transformed data to group all supervoxels into inter-patient clusters (Fig. 2 Step E). This method was chosen because it produced good results in our context and performed better than hierarchical clustering 3

in the presence of noise, as has already been demonstrated [10]. To facilitate the construction of the spectral clustering matrix, duplicate supervoxels were retained only once during the clustering process. Various sets of parameters were tested to group the supervoxels into clusters, as described in the later sections.

where mod is the considered modality between PET or LGE, xCmod is the mean value of the considered cluster for the considi mod mod ered modality, and xmax and xmin respectively are the maximum and minimum mean values within the clusters in the considered modality. The T1 and T2 abnormality scores of a cluster were quantitatively assessed, based on the myocardial T1 and T2 values at 3T given by Von Knobelsdorff-Brenkenhoff [16], and confirmed by senior cardiologists. The following definitions were used, based on a piecewise affine function defining a zero abnormality score for T1 and T2 values within clinically defined healthy ranges:

2.2.6. Step F. Abnormality score After the supervoxels were grouped into clusters, the statistical significance of these clusters was evaluated, and a score was attributed to each cluster and modality (Fig. 2 Step F). This was done to identify abnormal regions and compare them with physicians’ reports.

  −0.99     T 1Ci −1050−100     1000    sCT i1 =  0    T 1Ci −1300+100     1000    1

ANOVA and Tukey tests: To assess significant differences between clusters, an ANOVA test was conducted to examine the substantial differences in cluster medians for each modality independently. If the difference was significant (p < 0.05), a Tukey test was performed for each modality to establish statistical differences between each pair of clusters. We then examined the clusters that were statistically different from the reference cluster, as described in the following sections.

  −0.99     T 2Ci −38−10        100 sCT i2 =  0    T 2Ci −50+10     100    1

Choice of a reference cluster: As LGE and PET values are not strictly quantitative, a reference cluster had to be defined. This reference cluster, C p , comprised supervoxels that were presumed to be healthy, and was used as a baseline for comparison with the median values of the other clusters. As all subjects presented with ALVC, this cluster of healthy supervoxels had to be identified among the patients’ supervoxels. The reference cluster was defined as the “healthiest” among “valid” clusters, which were defined as those that met the inclusion criteria indicative of healthy tissue. “Valid clusters” had to contain more supervoxels than the average number of supervoxels per cluster and to exhibit median relaxation times within the normal ranges (1050 ms < T 1 < 1300 ms; 38 ms < T 2 < 50 ms) [16]. Additionally, the median LGE and PET values had to exceed the 5th percentile of their respective global distributions, while the median PET value remained below the 60th percentile. Within this subset of valid clusters, the reference cluster was defined as the one with the lowest mean LGE value, in accordance with the Padua criteria that identify LGE hypersignal as the main marker of ACM. These constraints ensured that C p could not correspond to a small region or be characterized by LGE/PET hyposignal or abnormally high PET activity.

xCmod − xCmod p i mod mod xmax − xmin

if 1000 ms ≤ T 1 < 1050 ms if 1050 ms < T 1 < 1300 ms

(4)

if 1300 ms ≤ T 1 < 1800 ms if T 1 > 1800 ms if T 2 < 20 ms if 20 ms ≤ T 2 < 38 ms if 38 ms < T 2 < 50 ms

(5)

if 50 ms ≤ T 2 < 100 ms if T 2 > 100 ms

where T 1Ci and T 2Ci are the mean T1 and T2 values of cluster Ci . Clusters not statistically different from the reference cluster in a given modality, based on the Tukey test, were assigned a score of 0 for that modality. 2.2.7. Step G1. Visualisation of the abnormal clusters on the images: To visualise the pathological zones detected on the medical images (Fig. 2 Step G1), the computed abnormality scores were superimposed on the original MRI and PET images. This highlighted areas with elevated LGE, T1, T2, or PET values. Supervoxels that had been previously excluded due to implausible T1 or T2 values were also superimposed on the original images to aid the identification of regions affected by poor registration, artifacts, or fat. 2.2.8. Step G2. Quantitative report generation and evaluation Obtaining a medical report for each patient: In order to summarise the cluster-derived information succinctly in a clinically useful format and to assess the performance of the clustering in relation to the diagnoses of cardiologists and radiologists, the cluster scores were converted into reports that are comparable with those produced by cardiac imagers (Fig. 2 Step G2). Each cluster was categorized as having normal or high values for each of its four modalities, based on thresholding of the cluster’s score within each modality. The threshold value was set at 0.1 for both T1 and T2 since the formulation of their scores was designed to define 0.1 as the upper boundary of normality. Since the PET and LGE scores were defined as

Scoring of the clusters for each modality: Each cluster was assigned an abnormality score for each modality. The method of determination of this score varied based on the modality under consideration. To assess the PET or LGE abnormality of a cluster Ci , the score sCmod was calculated relative to the median value xCmod of i p the reference cluster C p using the formula: sCmod = i

if T 1 < 1000 ms

(3) 4

distances to the reference cluster, their score thresholds were defined as hyperparameters in the nested cross-validation described in the next section. Subsequently, each cluster was associated with one of two categories for each of its four imaging modalities: high or normal. Patients were then assigned to the proportion of their cluster combinations, resulting in a personalized report detailing the distribution of various regions within their myocardium. Evaluation of the reports compared to the physicians’ reports: The method was evaluated by comparing the PET/LGE/T1/T2 combinations identified by clustering with those observed by physicians. To ensure consistency with the clustering-based report, physicians assessed each myocardial segment, noting the presence of high-intensity voxels in each of the four imaging modalities. A repeated nested cross-validation framework was used to compare the results. The score was defined as the mean balanced accuracy (BA) across all patients, averaged across all outer folds and repetitions. The dataset was split into three outer folds, each of which was subjected to a grid search with twofold inner cross-validation. The grid search process included multiple sets of hyperparameters, including variations in the spectral clustering parameters and the high-threshold values for PET and LGE.

Figure 4: Automatic detection workflow of the abnormal PET and MRI regions, illustrations on one basal slice of a patient from the cohort. The yellow square defines the LV VOI on which the clustering steps were applied. The multimodal volumes (a) are normalized and summed to one single ‘Z-volume’ (b), on which intra-patient clustering is applied to identify supervoxels (c). Supervoxels from all patients are then combined to identify 32 clusters that can be reported to the images (d). Lastly, the clusters are scored based on the intensity of their voxels, and clusters with high scores can be reported on the images to visualise abnormal regions for each modality (e).

3. Results 3.1. Quantitative evaluation of the method

Representation of the clustering reports with Bull’s eye plots: The clustering reports were summarized using two types of Bull’s Eye plot. The first type represented the main PET/LGE/T1/T2 combination observed per segment. The second type showed the percentage of high modality per segment. First, the Bull’s eye plots of the main combination per segment show the proportion of main imaging combinations (formed by normal or abnormally high T1, T2, LGE, and PET values) for each patient and segment. If a modality has a high value in a combination, it is represented in pink; otherwise, it is represented in white. This approach enables specific myocardial tissue to be characterized based on multimodal imaging. Secondly, the Bull’s eye plots of the ratios of abnormally high value voxels per segment per modality illustrate the ratio of high-value voxels per segment and modality, offering a quantitative perspective on the distribution of abnormalities within cardiac segments.

The model applied to the ninety-eight patients achieved a mean balanced accuracy of 0.76±0.04 in repeated nested crossvalidation across three outer folds and 56 repetitions. To further validate the method, it was applied to the 167 phantoms, using the hyperparameters chosen via nested CV (Appendix D). Balanced accuracies of 0.88, 0.80, and 0.80 were obtained, respectively, when the phantoms were created with noise levels of 0, 5%, and 10% (Appendix C). 3.2. Qualitative evaluation of the method To qualitatively evaluate the abnormal PET and MRI regions detection, the model was finally retrained on the full patient dataset using the chosen hyperparameters, giving a balanced accuracy of 0.81, with a mean sensitivity and specificity of 0.76 and 0.86. This two-step clustering workflow automatically highlights regions of interest in multimodal images by detecting areas associated with abnormal scores. By performing inter-patient clustering on normalized supervoxels, the method may reveal subtle pathological patterns invisible to visual inspection, which relies primarily on contrast differences. Fig. 4 presents an example of this automated detection workflow using a basal LGE slice from a representative patient. As shown in Fig. 4, the patient exhibits complete extinction of physiological myocardial fixation and no hypermetabolism in the left ventricle. The visualisation method highlighted no high voxel values for both PET and T2-map. However, it detected a large area of late contrast uptake under the epicardium

2.3. Validation of the method on numerical phantoms The clustering method was finally validated using a set of 167 numerical phantoms that simulated LGE, PET, T1, and T2 cardiac volumes. The process of creating the phantom is described in Appendix C. Double-step clustering was applied to the phantoms using the same parameters selected for the patients through nested cross-validation. The only difference was that the number of SLIC supervoxels was limited to 1000, due to the simpler anatomical representation of the phantoms and computational considerations. Various levels of noise η were tested: η ∈ 0%, 5%, 10%. 5

including glioma grading or classification [7, 8], identification of tumor habitats correlated with specific outcomes [9], and segmentation [10]. They showed the interest of using supervoxels for getting spatial information while being resistant to outliers and reducing computational costs compared to voxelbased methods. We particularly focused on the workflow proposed by Hansen et al. [10] for segmentation, as they experimented with several variations of this two-step method. Although our final goals were different, our work can also be considered a type of segmentation task, aiming to identify separate PET/MRI habitats in the left ventricle (LV), with the additional task to identify these habitats as either abnormal or normal regions to correlate with traditional medical reports. This additional identification task was completed by defining a scoring metric attributing to each cluster abnormality score for each modality. As T1 and T2 maps provide quantitative, medically documented values, the T1 and T2 scores were simply defined for these two modalities as piecewise linear functions centered on clinical reference values. To take advantage of the interpatient information, LGE and PET scores of each cluster were defined on the normalized PET and LGE values of the clusters by comparing the clusters to a healthy reference cluster. One challenge was thus the absence of healthy control patients, and therefore of clearly healthy clusters to use as a reference. We made the reasonable assumption that not all regions in each patient were pathological, making it possible to identify a reference cluster presumed to be healthy, against which the others could be compared. This reference cluster was built based on the hypothesis that it would not be too small, and that it would be among the clusters with the lowest PET and LGE values, while having T1 and T2 within clinically healthy ranges. Score thresholds were finally defined to dichotomize between ‘normal’ or ‘abnormal’ clusters for each modality, thus allowing for visualisation on the input volumes and comparison with the clinical routine reports. Additionally, some lighter modifications and additions were made to the original framework to correspond with our cardiac data and medical-related priors, while improving its robustness. Indeed, to force the clustering to focus on the LV myocardium, an extra selection of the supervoxel was added, by keeping only supervoxels with at least 20% of their volume within the LV segmentation. This reduced the overall complexity of the model while enabling a focused identification of clusters of interest within the LV. Excluding supervoxels with less than 20% of their surface area within the LV segmentation reduced the risk of oversegmentation by allowing the algorithm to discard poorly segmented supervoxels. The robustness of the method to imaging variations was also improved by the data augmentation techniques. The Box-Cox transformation was modified to include negative values. One main limitation of the proposed approach is that several methodological choices required the manual selection of parameters. Whenever possible, these parameters were validated using grid search within the nested cross-validation framework. However, some aspects, such as the exact definition of the reference cluster, could not be optimised in this way and required empirical choices. These decisions were nevertheless carefully

Figure 5: Generated Bull’s Eyes plots. In a), the main combination of the four modalities (formed by T1, T2, LGE, and PET having either normal or abnormally high voxel values) is illustrated. In b), the proportion of high voxels per segment is given per modality.

of the lateral wall. This same area exhibits T1 values greater than 1300 ms and was highlighted by clustering. 3.3. Automatisation of the multimodal image analysis To summarise the patient observations per segment in one scheme and for ease of visualisation, bull’s eye plots were created, as illustrated in Fig. 5. This was generated using the same patient as in Fig. 4. Two types of Bull’s Eye plot were generated. For each segment, a total of 24 = 16 possible combinations of T1, T2, LGE, and PET (either normal or high) could be observed. Fig. 5a shows the main combination of the four modalities detected in the segment and its proportion. For instance, the most frequently observed combination in the anterolateral basal segment was normal PET and T2 values alongside abnormally high LGE and T1 values. This combination accounted for 61% of the segment’s total volume. Fig. 5b illustrates the proportion of high voxels per segment in each modality using the Bull’s eye plot. Thus, the same segment was found to comprise 0% high PET voxels, 61% high LGE voxels, 89% high T1 voxels, and 0% high T2 voxels. 4. Discussion This unsupervised double-step clustering method has successfully automated the process of analysing and identifying multimodal imaging combinations from PET/MRI, achieving a balanced accuracy of 0.76 in nested cross-validation on 99 patients with ALVC. It allowed for precise multimodal visualisation of abnormal zones, while summarising the reports using bull’s eye plots. This provided segment-by-segment quantification of abnormal regions across the four imaging modalities at a level of precision that would be impossible for humans to achieve in practice. This supervoxel-based clustering workflow was inspired by several oncology studies [7, 8, 9, 10] with various purposes, 6

considered and discussed in close collaboration with senior cardiologists and radiologists. One of the challenges in this study relates to the absence of fully comprehensive ground truth data for evaluation. To address this, clustering reports were compared with segment-wise assessments provided by physicians, who manually reported the presence of high-value voxels in each modality for each cardiac segment, serving as a reference standard. While these expert annotations provide valuable clinical insight, they inherently reflect the complexity of visual interpretation, including inter-observer variability. Moreover, they primarily focus on the presence of high values within each segment, without explicitly capturing low values or intra-segment heterogeneity, thereby summarizing each segment into a single dominant state. In contrast, the proposed clustering approach offers a more detailed and quantitative characterization by estimating the distribution of multiple modality combinations within each segment. As such, this complementary perspective may provide additional granularity beyond conventional reporting, and comparisons based solely on human annotations may not fully reflect the added value of the method. Registration between the four patient volumes was also challenging. Using supervoxels reduces the impact of small registration errors by considering only the median values of each group of voxels and excluding supervoxels with more than 80% outside the registration area. However, poor image quality or movement meant that a few patients still had to be excluded, and large registration errors meant that in many cases, some slices or parts of slices had to be removed from the analysis. As expected, misalignments occurred most frequently in highly arrhythmogenic patients. Finally, the limited number of patients presented a challenge in such a clustering workflow. However, using repeated nested cross-validation maximised the number of patients used while evaluating the method and its variability. Testing the method on phantoms enabled more cases to be evaluated and compared to a known gold standard. Although this model is simpler than images of real patients and is subject to fewer sources of error (such as the absence of resampling and registration steps), it enabled the method to be validated with balanced accuracy greater than 0.8. However, despite being evaluated on controlled phantom data, the model did not achieve perfect balanced accuracy. This is probably due to the Gaussian smoothing simulating partial volume effects. However, this step was necessary, as excluding the filter resulted in clustering instabilities due to numerous supervoxels with identical intensity values. ALVC is a recently recognised pathogenic variant of ACM, which was first acknowledged with specific diagnostic criteria in the Padua criteria of 2020 [1]. To our knowledge, no risk score is yet defined for this specific variant, and the current diagnostic criteria still need validation on larger and more diverse cohorts [2]. Although studies showed the interest of PET imaging [4, 5, 6], as well as T1 [3] and T2 mapping for the characterization of ALVC and fibrosis, such imaging markers are still not in the current guidelines [2]. In this context, the proposed automatic framework enables the generation of quantitative PET/MRI reports through the identification of multimodal

myocardial habitats. This approach opens perspectives for a more systematic characterization of myocardial heterogeneity in ALVC patients. It provides a basis for future work relating these habitats to distinct clinical and phenotypic patient profiles and identifying prognostic TEP/MRI biomarkers in ALVC. Combining these imaging clusters with clinical, genetic, and follow-up data could provide a better understanding of the associations between imaging and phenotype, while demonstrating the complementarity of imaging and clinical data for clinical diagnosis. 5. Conclusion In conclusion, the current study focused on automating the process of generating quantitative PET/MRI imaging reports and identifying abnormal myocardial regions in ALVC patients. The proportion of high or normal PET, T1, T2, and LGE combinations for each cardiac segment of each patient was reported precisely, paving the way for future identification of connections between imaging combinations and clinical, genetic, and follow-up data. Future developments could include analysing such correlations, as well as exploring new imaging biomarkers and disease mechanisms. Sample CRediT Author Statement Brunnhilde Ponsi: Conceptualization, Formal Analysis, Investigation, Methodology, Software, Validation, Visualization, Writing - Original Draft. Thomas Carlier: Conceptualization, Funding Acquisition, Methodology, Project Administration, Resources, Supervision, Writing - Review and Editing. Lara Marteau: Methodology, Supervision, Writing - Review and Editing. Aurélien Monnet: Methodology, Writing - Review and Editing. Thomas Eugène: Data Curation, Methodology, Supervision, Writing - Review and Editing. Jean-Michel Serfaty: Methodology, Supervision, Writing - Review and Editing. Nicolas Piriou: Data Curation, Methodology, Supervision, Writing - Review and Editing. Hatem Necib: Conceptualization, Methodology, Project Administration, Resources, Supervision, Writing - Review and Editing. Declaration of Competing Interest Aurélien Monnet declares that he is an employee from Siemens Healthineers, which may be considered as potential competing interests with the work reported in this paper. All other authors declare that they have no known conflicts of interest in terms of competing financial interests or personal relationships that could have an influence or are relevant to the work reported in this paper. Ethics Statement This work involved human subjects or animals in its research. Approval of all ethical and experimental procedures and protocols was granted by the local independent ethics committee and the clinical trial is registered under the number NCT05450783. 7

Data Availability Statement

• rotation of 145 degrees around the z-axis

The patient data can be made available on request due to privacy/ethical restrictions.

• rotation of 180 degrees around the y-axis • flipping around the x-axis • flipping around the y-axis

Acknowledgments

• flipping around the z-axis

This work was supported by the French ”Program d’Investissement d’Avenir” (ANR-16-IDEX-0007) and the region ”Pays de la Loire” through their support to I-Site NExT, as well as by Siemens Healthineers, the industrial partner of the NExT research industrial chair IMRAM.

Following the rotation or flipping of the input volumes, supervoxels were extracted by SLIC as described in Step B. The final set of supervoxels associated with each patient was defined as the aggregation of all supervoxels obtained across the nine scenarios.

Appendix A. Choice of the SLIC parameters

Appendix C. Method for validation on the phantoms

Table A.1 describes the parameters chosen for SLIC.

Each phantom’s heart was created following the same process, but variations within the phantoms were brought by the inclusion of some randomness in the definition of their characteristics, as described in the following sections.

Table A.1: Chosen SLIC Parameters

Parameter n_segments compactness max_num_iter sigma min_size_factor max_size_factor

Value 400 · number of slices 0.01 10 0 0.5 3

Appendix C.1. Creation of a mask of the left ventricle shape Each phantom was given dimensions of 100 x 100 x Z, where Z is a randomly chosen integer between 3 and 10, representing the number of slices in which a patient can be imaged. This size was selected based on the approximate X x Y size of the LV VOI considered when applying the intra-patient clustering on the real patients. A LV shape was attributed to each phantom. This LV shape was modeled on each (X, Y) slice as two circles inside each other, which slightly different centers (respectively C1 and C2), as human hearts are not necessarily symmetrical. To introduce controlled variability in the choice of the centers, a range of variation ∆C1 was defined around the image center Cslicex = Cslicey = Cslice . This variation was set as a fraction of the image width X: X ∆C1 = (C.1) 20 The center C1 (x, y) of the first circle was then defined as:    C1x = rand (Cslice − ∆C1 , Cslice + ∆C1 ) (C.2)   C1y = rand (Cslice − ∆C1 , Cslice + ∆C1 )

The number of supervoxels, n_segments, was chosen visually to delineate well-defined regions while minimizing the number of supervoxels. The requested number of supervoxels only serves as a guideline for the algorithm, which adjusts this number as necessary to adequately segment all regions. In addition, we opted to perform SLIC within an ROI narrowed around the left ventricle, an area larger than the LV myocardium, where the supervoxels are selected and exported. Consequently, the clustering algorithm had some autonomy in determining the quantity of supervoxels to be distributed within this ROI. Compactness was intentionally set to a low value (0.01) to enable the supervoxels to adopt various shapes and align closely with potential geometric patterns of disease markers in the images. Other parameters were also explored, but the default settings of the scikit SLIC algorithm were found to yield the best results.

where the function rand(a, b) selects a random value uniformly between a and b. The second circle C2 was positioned near C1 by introducing a small random displacement around C1 . The position of C2 (x, y) was determined as follows:   −∆C ∆C    C2x = C1x + rand  2 1 , 2 1  (C.3)   C2y = C1y + rand −∆C1 , ∆C1 2 2

Appendix B. Data augmentation Standard data augmentation techniques (rotations and flips) were used to reduce the method’s sensitivity to imaging variations (Fig. 2, Steps A+B). In addition to the initial set of images, eight modification scenarios were established, each corresponding to the application of a single alteration to the input multimodal volumes among:

To ensure a physiologically plausible shape variation across slices, the radii of the circles were defined to decrease progressively along the slice index. Two random scaling coefficients were defined:    ascaling = rand(0.9, 1.2) (C.4)   bscaling = rand(0.35, 0.45)

• rotation of 10 degrees around the y-axis • rotation of 23 degrees around the x-axis • rotation of 90 degrees around the x-axis 8

The coefficient ascaling introduced randomness between phantoms, as a global scaling factor, while bscaling was used to control the rate at which the radius decreases along the slices (as the number of slices varies). The maximum initial radius R1,max of the first circle was defined as: R1,max = 5.2 ∗ ∆C1

(c) 20% probability of inclusion in the T2 abnormality mask. (d) 40% probability of inclusion in the PET abnormality mask. (e) 50% probability of inclusion in the ’PET - large’ abnormality mask.

(C.5)

Appendix C.3. Creation of the MRI and PET volumes, and inclusion of noise

For each slice z, the radii R1 and R2 of the circles C1 and C2 decreased according to the following formula:  z   cscaling (z) = ascaling − bscaling (Z−1)      (C.6) R1 (z) = cscaling · rand 4.6 · ∆C1 , R1,max        R2 (z) = R1 (z) − rand 4·∆C1 , 2 · ∆C1 3

T1, T2, LGE, and PET volumes were then created by differentiating three regions using the mask of the LV previously created: the ventricular cavity, the myocardium, and the area outside the LV. The inclusion of various levels of noise was tested. Each of these regions was given voxels values based on a region "central" value Xmod . The expression of these voxel values can be divided into two components:

where Z is the total number of slices. A mask of the LV was finally defined as the voxels in between the two circles.

1. First, a uniform value was given to the region, equal to a region region random number within [(1 − η) · Xmod , (1 + η) · Xmod ], where η is the noise percentage. This value allowed variation between phantoms. region region 2. Secondly, random uniform noise in [−ηXmod , ηXmod ] was added to each voxel of the zone to simulate variability within voxels.

Appendix C.2. Random creation of unhealthy LGE, T1, T2 and PET zones Five distinct masks were defined to represent the localization of randomly generated abnormal regions in LGE, T1, T2, PET, and ’PET - large’ imaging. The large PET mask was introduced to account for two different types of abnormal PET signal variations, as PET abnormalities tend to exhibit greater variability in human subjects compared to MRI-based markers. For each phantom, a random number of abnormal zones were created, uniformly chosen between 0 and 30. Each region was constructed following the procedure below:

The values chosen for the central values per region are given in Table C.2. region

Table C.2: Central Values Xmod Attributed to Each Region for Each Modality in the Creation of the Phantom Volumes.

1. A random point (x, y, z) within the LV is selected. 2. A circular mask centered at the selected point (x, y, z) is defined on the corresponding slice z, representing the locality of a lesion in the myocardium. The radius of this mask is 1 randomly selected from the interval [∆C1 , 5·∆C 2 ]. For large 1 PET regions, a larger radius is applied in 2.5 · [∆C1 , 5·∆C 2 ], reflecting the typically greater extent of PET abnormalities compared to MRI-detected lesions. 3. A second ring-shaped mask is then created to simulate a radial distribution of abnormalities around the LV. It is composed of a (X, Y)-plane ring of center C1 , width comprised between 0 and 5 voxels, and passing through (x, y, z). 4. The product of these two masks is computed to retain only the overlapping region. This ensures that the anomaly is both spatially constrained around the selected point and aligned with a concentric pattern around the LV, mimicking physiologically plausible distributions of myocardial lesions, such as fibrosis or infarction. 5. Finally, this constructed abnormal zone is attributed to the five abnormal LGE, T1, T2, PET, and large PET masks with various probabilities: (a) 70% probability of inclusion in the LGE abnormality mask. (b) 40% probability of inclusion in the T1 abnormality mask.

Sequence PET LGE T1 T2

Ventricular Cavity 2500 2222 1800 70

Myocardium 3000 2075 1250 43

Outside the LV 2000 2222 1000 130

Finally, the abnormal regions identified in the MRI and PET masks were incorporated into the modality volumes by modifying the previously defined voxel values. For each voxel within an abnormal area, a random value was added, sampled from the range [(1 − η) · Ymod , (1 + η) · Ymod ], where Ymod represents a predefined abnormality-related additive term specific to each modality, and is given in Table C.3. Table C.3: Abnormality-Related Additive Terms Ymod used in the Creation of the Phantom Volumes.

Sequence PET LGE T1 T2 9

Ymod 750 200 100 6

When affinity was "rbf" or "laplacian": • gamma ∈ {0.15, 0.25, 0.35, 0.50, 0.60, 0.75, 1, 5, 10} • thresh_high_LGE ∈ {0.15, 0.20, 0.25} • thresh_high_PET ∈ {0.15, 0.20, 0.25} with thresh_high_LGE and thresh_high_PET the thresholds used to distinguish between normal and abnormally high score values in a cluster. For visualization purposes and tests on the phantoms, hyperparameters had to be selected. The Laplacian affinity was selected, as it was chosen in 90% of cases during cross-validation. When this kernel was selected, the thresholds for high PET and high MRI scores were both equal to 0.25 in more than 50% of cases (79% and 53%, respectively), and were therefore retained as the optimal hyperparameters. In contrast, the gamma parameter, corresponding to the kernel coefficient, showed greater variability, most frequently ranging between 0.15 and 0.35, as shown in Table D.4. It was thus fixed to 0.25. When this kernel was fixed, n_neighbors was not a parameter. Table D.4: Distribution of the Percentage of Time Each Gamma Value was Identified as the Best Hyperparameter in the Nested Cross-Validation, When Considering Cases with Laplacian Kernel.

Figure C.6: Example of one randomly generated phantom with 5% noise.

Gamma 0.15 0.25 0.35 0.5 0.6 0.75 1

As two different masks represented abnormal areas in PET, abnormal values from both masks were successively added to the PET volume of each phantom. Appendix C.4. Inclusion of the volume partial effect A 3D Gaussian filter was finally applied to each of the created volumes to simulate the volume partial effect. The filter was applied in the x, y, and z directions by adjusting the kernel size to account for anisotropic voxel spacing. The width of the filter was chosen to be 2 mm for MRI volumes and 4 mm for PET volumes. Appendix C.5. Example of one numerical phantom Fig. C.6 gives an example of one randomly generated phantom with 5% noise.

Percentage of time 33% 20% 26% 11% 5% 3% 3%

References [1] D. Corrado, M. Perazzolo Marra, A. Zorzi, G. Beffagna, A. Cipriani, M. D. Lazzari et al., Diagnosis of arrhythmogenic cardiomyopathy: The Padua criteria, Int. J. Cardiol. 319 (2020) 106–14. doi:10.1016/j.ijcard.2020.06. 005.

Appendix D. Choice of hyperparameters The nested cross-validation relied on an inner grid search with conditional hyperparameter sets, defined as follows. When affinity was "nearest_neighbors":

[2] D. Corrado, A. Anastasakis, C. Basso, B. Bauce, C. Blomström-Lundqvist, C. Bucciarelli-Ducci et al., Proposed diagnostic criteria for arrhythmogenic cardiomyopathy: European Task Force consensus report, Int. J. Cardiol. 395 (2024) 131447. doi:10.1016/j.ijcard. 2023.131447.

• n_neighbors ∈ {5, 10, 12, 15, 20, 30} • thresh_high_LGE ∈ {0.15, 0.20, 0.25} • thresh_high_PET ∈ {0.15, 0.20, 0.25} When affinity was "polynomial":

[3] R. Everett, C. Stirrat, S. Semple, D. Newby, M. Dweck S. Mirsadraee, Assessment of myocardial fibrosis with T1 mapping MRI, Clin. Radiol. 71 (2016) 768–78. doi:10. 1016/j.crad.2016.02.013.

• thresh_high_LGE ∈ {0.15, 0.20, 0.25} • thresh_high_PET ∈ {0.15, 0.20, 0.25} 10

[4] R. Tessier, L. Marteau, M. Vivien, B. Guyomarch, A. Thollet, I. Fellah et al., 18F-Fluorodeoxyglucose Positron Emission Tomography for the Detection of Myocardial Inflammation in Arrhythmogenic Left Ventricular Cardiomyopathy, Circ. Cardiovasc. Imaging 15 (2022) e014065. doi:10.1161/CIRCIMAGING.122.014065.

[15] R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua S. Süsstrunk, SLIC superpixels compared to stateof-the-art superpixel methods, IEEE Trans. Pattern Anal. Mach. Intell. 34 (2012) 2274–2282. doi:10.1109/ TPAMI.2012.120. [16] F. von Knobelsdorff-Brenkenhoff, M. Prothmann, M. A. Dieringer, R. Wassmuth, A. Greiser, C. Schwenke et al., Myocardial T1 and T2 mapping at 3T: reference values, influencing factors and implications, J. Cardiovasc. Magn. Reson. 15 (2013) 53. doi:10.1186/1532-429X-15-53.

[5] R. Neves, A. S. Tseng, R. Garmany, A. L. Fink, C. J. McLeod, L. T. Cooper et al., Cardiac fludeoxyglucose18 positron emission tomography in genotype-positive arrhythmogenic cardiomyopathy, Int. J. Cardiol. 389 (2023) 131173. doi:10.1016/j.ijcard.2023.131173.

[17] M. A. Stoto J. D. Emerson, Power Transformations for Data Analysis, Sociol. Methodol. 14 (1983) 126. doi:10. 2307/270905.

[6] A. Protonotarios E. Wicks, The role of FDG-PET imaging in arrhythmogenic cardiomyopathy, Int. J. Cardiol. 391 (2023) 131275. doi:10.1016/j.ijcard.2023.131275.

[18] G. E. P. Box D. R. Cox, An Analysis of Transformations, J. R. Stat. Soc. Ser. B (Stat. Methodol.) 26 (1964) 211–243. doi:10.1111/j.2517-6161.1964.tb00553. x.

[7] R. Inano, N. Oishi, T. Kunieda, Y. Arakawa, T. Kikuchi, H. Fukuyama et al., Visualization of heterogeneity and regional grading of gliomas by multiple features using magnetic resonance-based clustered images, Sci. Rep. 6 (2016) 30344. doi:10.1038/srep30344. [8] H. Tatekawa, A. Hagiwara, H. Uetani, S. Bahri, C. Raymond, A. Lai et al., Differentiating IDH status in human gliomas using machine learning and multiparametric MR/PET, Cancer Imaging 21 (2021) 27. doi:10.1186/ s40644-021-00396-5. [9] A. J. Even, B. Reymen, M. D. La Fontaine, M. Das, F. M. Mottaghy, J. S. Belderbos et al., Clustering of multiparametric functional imaging to identify high-risk subvolumes in non-small cell lung cancer, Radiother. Oncol. 125 (2017) 379–84. doi:10.1016/j.radonc.2017.09. 041. [10] S. Hansen, S. Kuttner, M. Kampffmeyer, T.-V. Markussen, R. Sundset, S. K. Øen et al., Unsupervised supervoxelbased lung tumor segmentation across patient scans in hybrid PET/MRI, Expert Syst. Appl. 167 (2021) 114244. doi:10.1016/j.eswa.2020.114244. [11] 3D Slicer image computing platform, https://www. slicer.org/, 2012. [accessed 14 April 2026]. [12] A. Fedorov, R. Beichel, J. Kalpathy-Cramer, J. Finet, J. C. Fillion-Robin, S. Pujol et al., 3D Slicer as an Image Computing Platform for the Quantitative Imaging Network, Magn. Reson. Imaging 30 (2012) 1323–41. doi:https: //doi.org/10.1016/j.mri.2012.05.001. [13] CVI, Circle Cardiovascular Imaging Inc., 2026. URL: https://www.circlecvi.com/cardiac-mr. [14] S. Klein, M. Staring, K. Murphy, M. A. Viergever J. P. W. Pluim, Elastix: a toolbox for intensity-based medical image registration, IEEE Trans. Med. Imaging 29 (2010) 196–205. doi:10.1109/TMI.2009.2035616. 11

Record · ID 370326 · SHA-256 3e2ee57980271d7a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.