SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Accepted at the 43rd International Conference on Machine Learning (ICML 2026).
Luke James Miller 1 Yugyung Lee 1
arXiv:2605.12389v1 [cs.CV] 12 May 2026
Abstract
More broadly, SEMIR establishes a framework for learning task-adapted, topology-preserving latent representations with exact decoding for highresolution structured visual data.
Segmenting small and sparse structures in largescale images is fundamentally constrained by voxel-level, lattice-bound computation and extreme class imbalance–dense, full-resolution inference scales poorly and forces most pipelines to rely on fixed regionization or downsampling, coupling computational cost to image resolution and attenuating boundary evidence precisely where minority structures are most informative. We introduce SEMIR (Semantic Minor-Induced Representation Learning), a representation framework that decouples inference from the native grid by learning a task-adapted, topology-preserving latent graph representation with exact decoding. SEMIR transforms the underlying grid graph into a compact, boundary-aligned graph minor through parameterized edge contraction, node deletion, and edge deletion, while preserving an exact lifting map from minor predictions to lattice labels. Minor construction is formalized as a fewshot structure learning problem that replaces handtuned preprocessing with a boundary-alignment objective: minor parameters are learned by maximizing agreement between predicted boundary elements and target-specific semantic edges under a boundary Dice criterion, and the induced minor is annotated with scale- and rotation-robust geometric and intensity descriptors and supports efficient region-level inference via message passing on a graph neural network (GNN) with relational edge features. We benchmark SEMIR on three tumor segmentation datasets—BraTS 2021, KiTS23, and LiTS—where targets exhibit high structural variability and distributional uncertainty. SEMIR yields consistent improvements in minority-structure Dice at practical runtime.
1. Introduction Semantic segmentation of medical images presents a fundamental tension between computational scalability and structural fidelity. Volumetric modalities such as CT and MRI routinely produce images exceeding 108 voxels, yet clinically relevant structures—tumors, lesions, fine anatomical boundaries—often occupy fewer than 1% of the volume (Ghankot et al., 2025). Dense, voxel-wise inference scales poorly to such regimes: computational cost grows with image resolution rather than structural complexity, and extreme class imbalance attenuates the gradient signal from minority structures precisely where accurate delineation is most critical (Gao et al., 2025). Existing approaches address scalability through bottom-up compression: patch-based processing, fixed-grid downsampling, or hand-crafted superpixel extraction. These methods couple the inference resolution to the native lattice or predetermined regionization, sacrificing boundary precision for tractability. The resulting representations are task-agnostic—they do not adapt to the semantic structure of the target domain—and lose boundary evidence before models can leverage it. Multi-class formulations that jointly segment all structures must balance competing objectives across classes with vastly different spatial extent—a tradeoff that loss reweighting only partially addresses. We propose an alternative, top-down formulation: rather than compressing the image into a fixed representation, we learn a task-adapted inference space that preserves semantically relevant structure while dramatically reducing computational burden. Our approach, SEMIR (Semantic Minor-Induced Representation Learning), constructs a compact graph minor H ⪯ G from the native voxel lattice G through parameterized edge contraction, node deletion, and edge deletion. The minor H defines a sparse set of supernodes aligned to image boundaries, with node count typi-
1
Department of Computing, Analytics and Mathematics; University of Missouri-Kansas City, Kansas City, United States. Correspondence to: Luke James Miller <[email protected]>. Preprint. May 13, 2026.
1
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation Table 1. Comparison of region reduction approaches. SEMIR provides topology-preserving, boundary-aligned reduction with taskadapted parameters, unlike token merging or fixed pooling methods. 1 (Ying et al., 2018) 2 (Achanta et al., 2012) 3 (Felzenszwalb & Huttenlocher, 2004)
p cally on the order of |V (G)|—reducing a 2563 volume from ∼107 voxels to ∼103 supernodes while maintaining an exact bijective lifting map for voxel-level prediction. Unlike multi-class pipelines that segment all structures jointly, SEMIR constructs a target-specific inference space: for each structure of interest, we learn a boundary-aligned minor optimized to delineate that structure, reducing the problem to binary classification on the induced graph. The minor construction is governed by a small parameter set Θ, optimized via black-box few-shot optimization on boundary alignment objective. This few-shot optimization replaces manual threshold tuning with a principled, datadriven procedure: given a handful of labeled examples, the optimizer recovers parameters Θopt such that the induced minor boundaries maximize agreement with ground-truth semantic edges, independent of the downstream label set. The resulting representation generalizes across segmentation tasks sharing similar boundary statistics. Downstream inference operates entirely on the minor H: supernodes are annotated with scale- and rotation-invariant geometric and intensity descriptors, and a graph neural network propagates information across the reduced topology to produce supernode-level predictions. The lifting operator then maps these predictions back to the original lattice, yielding voxel-wise segmentation at a fraction of the computational cost of dense inference.
DiffPool1
SLIC2
Felzenswalb3
SEMIR
Preserves Topology
∼
✓
✓
✓
Task Aware
✓
✗
✗
✓
Boundary Aligned
✗
∼
∼
✓
Exact Lifting
✗
✓
✓
✓
Decoding Guarantee
✗
✗
✗
✓
2. Related Work 2.1. Medical Image Segmentation U-Net and its variants—incorporating attention, residual connections, and transformer encoders—dominate medical image segmentation (Jiangtao et al., 2025; Luo et al., 2025; Peng et al., 2025). However, computational cost scales with image resolution rather than structural complexity, necessitating patch-based inference or downsampling that discards fine boundary information before the model observes it. The challenge is acute for minority structures: tumors and lesions occupy vanishing fractions of image volume, creating class imbalance that specialized losses only partially mitigate (Yeung et al., 2022; Hosseini, 2025). Boundary-aware and topological losses improve fidelity but remain tied to dense, lattice-bound inference (Zheng et al., 2025).
We evaluate SEMIR on three challenging tumor segmentation benchmarks—BraTS 2021, KiTS23, and LiTS—where targets exhibit high structural variability, extreme class imbalance, and ambiguous boundaries. Across all three datasets, SEMIR yields consistent improvements in minority-structure Dice while maintaining competitive performance on aggregate metrics, demonstrating that structure-adaptive representations offer a principled path toward scalable, boundary-aware segmentation in highresolution medical imaging.
2.2. Superpixel and Oversegmentation Methods
Contributions (1) A learned graph minor framework decoupling inference complexity from image resolution with exact voxel-level lifting—no interpolation, no boundary artifacts, no approximation. (2) Few-shot, black-box optimization for boundary alignment, replacing manual parameter tuning with a data-driven procedure requiring only 5–20 labeled examples. (3) Scale- and rotation-invariant node/edge descriptors enabling effective message passing on anisotropic medical volumes. (4) Consistent improvements on minoritystructure Dice for tumor segmentation—the regime where multi-class methods systematically underperform.
Superpixel algorithms—SLIC (Achanta et al., 2012), Felzenszwalb-Huttenlocher (Felzenszwalb & Huttenlocher, 2004), watershed (Neubert & Protzel, 2014)—reduce complexity by grouping pixels into perceptually coherent regions. However, their regionization is task-agnostic: boundaries follow low-level gradients rather than semantic structure, and parameters require manual tuning. Learned superpixel methods improve boundary adherence but remain constrained by fixed grids or differentiable relaxations (Stutz et al., 2018). In medical imaging, superpixels serve primarily as preprocessing for classical classifiers (Liu et al., 2020). A fundamental limitation is the absence of a principled framework relating induced regions to the original image; mapping predictions back requires heuristics that introduce boundary artifacts. Differentiable pooling methods (MinCutPool, DMoN) learn soft cluster assignments but
Conflict of Interest Disclosure. The authors declare no financial conflicts of interest related to this work.
2
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
lack lifting guarantees; learned superpixel networks (SSN, SEAL) improve boundary adherence but remain grid-bound.
predictions are mapped to the voxel grid via the bijection encoded in T and a lifting function. The minor typically reduces inference from ∼107 voxels to ∼103 supernodes while guaranteeing exact voxel-level recovery.
2.3. Graph Neural Networks in Medical Imaging GNNs have been applied to anatomical structure modeling, histopathology cell graphs (Guan et al., 2025), and region adjacency graphs over superpixels (Brussee et al., 2025; Mienye & Viriri, 2025). However, graphs are typically constructed from fixed regionizations—anatomical templates, regular grids, or classical superpixel outputs—determined independently of the downstream task (Ding et al., 2022).
3.2. Notation We segment volumetric images I ∈ RH×W ×D×C into K classes, producing predictions Ŷ ∈ {0, . . . , K−1}H×W ×D . A dataset D = {(I (n) , Y (n) )}N n=1 pairs images with groundtruth label maps. The N-connected voxel grid defines graph G = (V (G), E(G)) with connectivity N . The graph minor H = (V (H), E(H), X(H), F (H)) comprises supernodes with feature matrix X(H) ∈ R|V (H)|×dx and edge features F (H) ∈ R|E(H)|×df . Few-shot optimization yields Θopt = R(Dfew , Θ) by aligning predicted boundaries ŶB = SB (T, Θ) with ground-truth boundaries YB relevant to target structures. Full notation table in Appendix A.
2.4. Graph Minors Graph minors provide a rigorous framework for relating a derived graph H to a parent G: H ⪯ G if H can be obtained through edge contractions, edge deletions, and node deletions. The Robertson-Seymour theorem establishes that minor-closed properties admit finite forbidden minor characterizations, with polynomial-time testing for fixed H (Robertson & Seymour, 2004; Lovász, 2006). Minors have seen limited vision application. The key insight we exploit is that edge contraction defines a surjection from parent to minor nodes, inducing an exact partition of the original vertex set (Demaine et al., 2005). This partition provides a bijective lifting map for transferring predictions to the original lattice without approximation.
3.3. Expanded Tensor Representation The tensor T ∈ {0, · · · , 255}(2H−1)×(2W −1)×(2D−1) interleaves node and edge states along each dimension by encoding a bitflag with the state of the element. Positions with all even indices (2j, 2k, 2l) encode nodes corresponding to voxel (j, k, l). Positions with at least one odd index encode edges: an odd index along a single axis indicates a face-adjacent edge (e.g., (2j + 1, 2k, 2l) encodes the edge between voxels (j, k, l) and (j +1, k, l)), while odd indices along multiple axes indicate diagonal edges under higher connectivity. The exact set of valid edge positions depends on the chosen N-connectivity (N ∈ {6, 10, 18, 26}). Each entry is binary:
Our approach differs from superpixel methods in three respects: (1) the minor is constructed through parameterized graph operations with formal guarantees; (2) parameters are optimized for boundary alignment via black-box optimization; and (3) the lifting map is exact by construction.
This representation avoids explicit storage of voxel coordinates and edge lists, reducing memory to O(|V (G)| + |E(G)|) with single-byte entries. Flood-fill visits each edge at most once and terminates immediately upon revisiting assigned voxels, yielding O(|E(G)|−|E(H)|) operations— complexity that scales with contractions performed rather than image resolution. In practice, minor construction on large volumes completes in under one second on CPU using a Rust backend called from Python. Graph-minor construction is performed once after Θopt is fixed, and the GPU receives precomputed graph batches for training and inference; CPU–GPU transfer therefore does not interrupt the forward pass. This makes minor construction negligible compared to downstream inference.
3. Method 3.1. Overview SEMIR constructs a compact, boundary-aligned graph minor from the native voxel lattice and performs segmentation via node classification on this reduced representation (Figure 1). The input volume is encoded as a binary tensor T representing the N-connected grid graph G. A graph minor is then derived through three parameterized operations: edge contraction which merges similar voxels into supernodes, node deletion which removes supernodes violating size constraints, and edge deletion which severs connections across intensity gradients. The parameter set Θ defines a family of feasible partitions; few-shot optimization selects from this hypothesis class by maximizing boundary alignment with target-specific semantic edges. Supernodes are annotated with geometric and intensity descriptors; edge features encode scale- and rotation-invariant relative differences. A GNN performs node classification on the graph minor, and
3.4. Graph Minor Construction The initial conversion of an image I to a graph represents each voxel as a node and connects neighboring voxels according to a configurable N-connectivity in 3D, where N ∈ {6, 10, 18, 26} denotes the number of adjacent neigh3
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Figure 1. SEMIR pipeline visualization on a KiTS23 case. (a) Input contrast-enhanced CT. (b) Zoomed region showing the native voxel grid; each voxel corresponds to a node in G. (c) Boundary-aligned graph minor H ⪯ G constructed via parameterized edge contraction, node deletion, and edge deletion with few-shot optimized Θopt ; supernodes colored by ID. (d) Node-level predictions ŶH from GNN inference on H. (e) Final voxel segmentation Ŷ obtained via bijective lifting. (f) Ground truth Y . Kidney shown in red, tumor in blue.
bors (face-, face+edge-, face+edge+corner-adjacent, etc.). The choice of N is modality-dependent and ablated in Section 4.4; we default to N =6. The resulting grid graph G = (V (G), E(G)) is defined as follows:
After contraction, node deletion removes supernodes violating retention criteria. Let β = (βmin , βmax , mmin , mmax ), where βmin , βmax control area thresholds and mmin , mmax control intensity thresholds. Vdel := v ∈ V (H) : av < βmin ∨ av > βmax ∨I¯v < mmin ∨ I¯v > mmax , (3)
V (G) = {(j, k, l) : 1 ≤ j ≤ H, 1 ≤ k ≤ W, 1 ≤ l ≤ D}, E(G) = (u = (j, k, l), v = (j ′ , k ′ , l′ )) ∈ V (G) × V (G)
V (H) ← V (H) \ Vdel .
s.t. u and v are N-connected neighbors. (1)
This step removes small noisy regions (low area), large background regions (high area), and supernodes with intensities outside the desired range. The lower area bound is specifically intended to suppress acquisition noise: isolated voxels that fail the contraction criterion form singleton or very small supernodes and are pruned before downstream inference. Deleted nodes receive background by default during lifting, so this mechanism is conservative rather than a source of artificial foreground improvement. Finally, edge deletion defines segmentation boundaries by removing edges across strong intensity differences: Edel := (vi , vj ) ∈ E(H) : ∥I¯vi − I¯vj ∥n > α , (4) E(H) ← E(H) \ Edel ,
The core data-representation method of this work treats the graph derived from the image (via its expanded tensor T ) as the parent graph G and constructs a graph minor H ⪯ G through a sequence of operations: edge contraction, node deletion, and edge deletion. These operations are performed during a pseudo-random coprime traversal of the nodes to avoid directional bias (e.g., rightward, downward, or backward smearing artifacts in the resulting superpixels). Edge contraction grows each supernode from a seed voxel. A neighboring voxel p is merged into the current supernode seeded at s when it is connected by an edge in the current graph and satisfies ∥Ip − Is ∥n ≤ ψ,
where α is the edge deletion threshold from Θ. Removed edges separate supernodes across significant intensity gradients, effectively partitioning H into distinct components aligned with image boundaries.
(2)
where Ip , Is ∈ RC are voxel intensity vectors, ∥ · ∥n is a configurable Ln norm (e.g., Euclidean n = 2, Manhattan n = 1, Chebyshev n = ∞), and ψ is the contraction threshold parameter from Θ. Contraction is therefore anchored to the seed voxel rather than to a running supernode mean. This prevents gradual low-contrast transitions from collapsing into a single merged region and instead tends to preserve them as chains of adjacent supernodes.
3.4.1. F EW- SHOT R EPRESENTATIONAL PARAMETER L EARNING Manual specification of the parameter set Θ for the graph minor generation function S(·) is labor-intensive and imprecise. Instead, we frame parameter selection as black-box 4
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
During the contraction phase of S(·), voxel counts, boundary exposures, coordinates, and intensities are aggregated per supernode. Spatial coordinates are used transiently to compute the covariance matrix of voxel positions, from which the dominant axis and elongation are derived. Intensity statistics (per-channel standard deviation and covariance) are computed directly from the voxel intensities in each supernode. For each supernode u ∈ V (H) with associated voxels Pu = {p = (j, k, l) : p belongs to u} (where p is the voxel coordinate and I(p) ∈ RC its intensity vector), we define the following stored features:
au := |Pu | s 1 X σu := (I(p) − I¯u ) ⊙ (I(p) − I¯u ), au p∈Pu
Figure 2. Boundary alignment visualization on a LiTS case. (top left) Ground-truth semantic boundary YB . (top right) Supernode boundaries induced by naive parameters Θinit . (bottom left) Supernode boundaries induced by optimized Θopt . (bottom right) Zoomed comparison. Θopt boundaries align with the semantic edge, while Θinit boundaries cut through the tumor boundary.
p∈Pu
au is the supernode area/voxel count, σu is the per-channel standard deviation of intensity, and P Σu is the intensity covariance matrix where I¯u = a1u p∈Pu I(p) ∈ RC is the mean intensity vector (computed transiently for variance and covariance). The spatial covariance Σcoord ∈ R3×3 is u computed transiently from voxel coordinates. Let λu,1 ≥ λu,2 ≥ λu,3 ≥ 0 be the eigenvalues of Σcoord with correu sponding eigenvectors vu,1 , vu,2 , vu,3 . We extract:
optimization over the discrete parameter space with a treebased surrogate R(·) to minimize boundary misalignment on a small held-out subset Dfew ⊂ D. The optimized parameters are obtained via: h i Θopt = arg min E(I,Y )∼Dfew L SB (T, Θ), YB , (5) Θ
du := vu,1 ∈ R3 , q elongu := (λu,1 + ε)/(λu,2 + ε),
where T is the expanded tensor representation of I, YB is the ground-truth binary boundary map derived from the corresponding label map Y (target-specific), and the loss is the inverse Dice-Sørensen coefficient (DSC): L(ŶB , YB ) = 1 − DSC(ŶB , YB ) = 1 −
2|ŶB ∩ YB | |ŶB | + |YB |
(7)
1 X Σu := (I(p) − I¯u )(I(p) − I¯u )⊤ ∈ RC×C , au
(8)
p⋆u := argmin(j,k,l)∈Pu ; lexicographic order Where du is the unit eigenvector of the largest eigenvalue, elongu is an elongation ratio with ε > 0 for stability, and p⋆u is the supernode’s canonical voxel. Boundary length, bu , and a 3D compactness proxy (inverse spikiness), compu , are computed from the original grid graph G with ε > 0 for numerical stability.
. (6)
This boundary-focused loss encourages the minor generation process to produce superpixels whose boundaries match the semantic edges in the ground truth, independent of the specific class identities in {0, . . . , K −1}. Crucially, Θ does not configure a fixed segmentation model—it parameterizes a family of graph homomorphisms πΘ : G → HΘ , each inducing a distinct partition of the voxel lattice. Few-shot optimization over Θ thus constitutes representation learning over structured latent spaces: the optimizer selects a partition structure that minimizes semantic boundary risk, rather than tuning hyperparameters within a fixed architecture.
bu := {(p, q) ∈ E(G) : p ∈ Pu , q ∈ / Pu } . compu :=
36πa2u ∈ (0, 1] b3u + ε
(9)
While node features capture absolute properties of individual supernodes, edge features encode relative differences between adjacent supernodes, promoting scale- and rotationinvariant representations suitable for the graph neural network. For each edge e = (u, v) ∈ E(H), we first order the incident supernodes by area, and for any scalar node feature we compute a scale-invariant log-ratio. We calculate relative geometric and orientation edge features for each scalar node feature. Full descriptions can be found in Appendix C.
3.4.2. G RAPH M INOR F EATURES Given the optimized parameter set Θopt , we execute the graph minor generation H = S(T, Θopt ), to extract a compact set of node/edge features for downstream prediction. 5
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
seed-anchored contraction, α controls boundary separation, and β controls size and intensity retention. Thus, the induced family of partitions is much smaller than the class of arbitrary voxel labelings. Few-shot optimization therefore searches over stable boundary statistics rather than classconditional appearance, which explains why 5–20 labeled examples are sufficient in our experiments while still allowing Θopt to vary across modality, resolution, and target structure.
3.5. Downstream Prediction Given the learned graph minor H, we apply a graph neural network for supernode classification. The final voxel-wise segmentation is obtained by lifting these predictions back to the image grid via the bijective mapping in T : ŶH = GNN(H), Ŷ = Lift(H, ŶH , T ).
(10)
where ŶH ∈ {0, . . . , K − 1}|V (H)| denotes the predicted class for each supernode u ∈ V (H). This lifting operation assigns the predicted class of each supernode to all voxels belonging to that supernode, ensuring consistent labeling within each superpixel. During training, the GNN loss and performance metrics are computed by comparing the lifted voxel predictions Ŷ against the ground-truth Y (n) .
Lemma 3.1 (Supernode Connectivity). For any supernode u ∈ V (H), the pre-image π −1 (u) ⊆ V (G) induces a connected subgraph of the N-connected grid G. Proof. Supernodes are constructed via flood-fill from a seed voxel (Algorithm 3). At each iteration, only N-adjacent voxels satisfying the contraction threshold ψ are merged. Since N-adjacency in the grid graph G implies an edge in E(G), and flood-fill proceeds by traversing such edges, the set of merged voxels π −1 (u) forms a connected subgraph of G by construction.
Composable multi-class inference. While we train separate binary models per target structure, SEMIR supports compositional multi-class inference by constructing multiple task-adapted minors and resolving overlapping predictions through confidence-weighted voting or energy minimization over the induced partition lattice. This compositional design is intentional: rather than forcing a single representation to balance competing objectives across structures with vastly different spatial extent and boundary statistics, SEMIR constructs a representation optimized for each target. For K target structures, compositional inference requires K minor-construction and GNN passes. The scaling is therefore linear in the number of targets, but each pass operates on the induced graph rather than the full voxel lattice, and the final lifting and overlap-resolution step adds negligible overhead relative to graph construction and GNN inference.
Theorem 3.2 (Lifting Exactness). Let H = S(G, Θ) be the induced minor over retained supernodes, with predictions ŶH : V (H) → {0, . . . , K − 1}. The lifting operator Lift(H, ŶH , T ) assigns labels to voxels exactly: for any u ∈ V (H) with prediction ŷu , every voxel v ∈ π −1 (u) receives label ŷu with no interpolation or approximation. Proof. The bijection tensor T encodes the supernode membership of each voxel via the flood-fill assignment. For retained supernodes, T provides a surjective map from voxels to supernodes; lifting inverts this by assigning each voxel the label of its unique containing supernode. No interpolation is required because supernode boundaries are defined exactly by the contraction process, and each voxel belongs to exactly one supernode.
3.6. Theoretical Properties Superpixel segmentation is inherently ill-posed, requiring trade-offs between boundary adherence, compactness, size uniformity, and tractability (Stutz et al., 2018; Wang et al., 2017). Our graph-minor formulation makes these tradeoffs explicit and learnable rather than implicit and manual. The parameter set Θ = {ψ, α, β} explicitly parameterizes these trade-offs: edge contraction (ψ) promotes compact supernodes by merging similar voxels, edge deletion (α) enforces boundary separation along intensity gradients, and node deletion (β) prunes outliers violating size constraints. Crucially, Θ does not tune a fixed model; it defines the hypothesis class of feasible partitions. Few-shot optimization over Θ thus constitutes learning the structure of the inference space itself, not merely selecting hyperparameters within a fixed architecture.
Structure-complexity duality. SEMIR formalizes a fundamental duality: inference cost scales with the topological complexity of semantic boundaries rather than the resolution of the underlying lattice. Dense methods pay for every voxel; SEMIR pays only for structure. Additional properties. Beyond the formal results above, the construction provides: (i) Adaptivity—few-shot optimization tunes Θ to domain-specific boundary statistics without retraining the downstream predictor; (ii) Invariance—relative edge features (log-ratios of geometric descriptors, normalized intensity differences) reduce sensitivity to absolute scale and orientation, valuable for anisotropic medical volumes; (iii) Tractability—near-linear complexity in voxel count (see Section 3.3).
Few-shot generalization of Θ. The few-shot claim concerns the optimization process, not transfer of a single fixed parameter vector across datasets. The parameter set Θ is low-dimensional and physically constrained: ψ controls
Limitations. The construction lacks the global optimality guarantees of energy-minimization approaches (Kostrykin 6
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation Table 2. Dataset characteristics.
Source Focus Classes Trn/Val Slice size Res. (mm)
BraTS
KiTS
LiTS
MRI Brain Glioma 1,251 / – 240x240 1x1x1
CT Kidney Tumor/cyst 489 / 110 512x512 variable
CT Liver Tumor 131 / 70 512x512 variable
(foreground = target, background = all other voxels), reducing each segmentation task to structure-specific inference rather than joint multi-class prediction. This formulation sidesteps class imbalance by construction. See Appendix F for full reproducibility details. Architecture and training. We use a 3-layer GINE (Xu et al., 2019; Hu* et al., 2020) with hidden dimension 128 for supernode prediction. Training uses Adam (Kingma & Ba, 2017) (lr = 10−3 ) for up to 200 epochs with early stopping (patience 10) on validation Dice. Parameter optimization. Minor parameters Θ are optimized via SMBO with an ExtraTrees surrogate (Bergstra et al., 2011; Berrouachedi et al., 2019) on a few-shot subset (|Dfew | ∈ {5, 10, 20} depending on dataset). Optimized parameters are fixed for all subsequent training; no manual tuning is performed.
& Rohr, 2022). Pseudo-random traversal order introduces stochastic variation across runs. Empirically, variation across runs with different traversals was negligible. Boundary adherence may degrade in low-contrast regions if α is poorly calibrated, and excessive contraction (ψ too permissive) can oversmooth fine structures. Higher N-connectivity (18- or 26-connected) better preserves anisotropic structures at the cost of increased memory. In our implementation, increasing connectivity increases the expanded tensor and edge-state storage, with limited marginal value for volumetric medical images; N =6 therefore remains the recommended default for CT and MRI. For thin-structure 2D settings where diagonal adjacency is more important, higher connectivity can be preferable. Ablation studies (Section 4.4) empirically quantify these trade-offs.
4.3. Results We restrict comparisons to methods reporting per-class Dice, as aggregate metrics mask failures on minority structures (cf. Swin UNETR on KiTS: 0.762 aggregate vs. 0.343 tumor). Table 3 provides contextual comparison to published methods under their reported protocols, while Table 4 provides the controlled binary target-vs-rest comparison under identical splits. Across all three benchmarks, SEMIR consistently improves performance on the clinically relevant minority targets while remaining competitive on the corresponding organ-plustarget aggregates. On BraTS (Table 3), SEMIR attains the strongest performance on TC and ET, indicating improved delineation of tumor core structures that are typically the most fragile under class imbalance and heterogeneous appearance. Notably, gains are concentrated on the more difficult tumor-derived regions rather than being driven solely by the largest whole-tumor mask.
4. Experiments 4.1. Datasets We evaluate on three volumetric medical imaging benchmarks spanning MRI and CT modalities with varying resolution, anisotropy, and class imbalance characteristics as listed in Table 2. KiTS23 and LiTS use their documented benchmark splits. BraTS 2021 does not provide an official train/validation split; we therefore use an 80/20 train/validation partition, and BraTS comparisons to published methods should be interpreted as contextual rather than split-controlled.
On KiTS and LiTS (Table 3), the same pattern holds: SEMIR yields substantial improvements on the tumor-only class while remaining within the competitive range on organplus-tumor aggregates. This is by design: SEMIR does not optimize for aggregate metrics across all classes. Instead, each model targets a single structure, and we report the metric that matters for that target. SEMIR’s structure-targeted approach avoids this failure mode. Qualitative examples and ablations in the following sections further characterize where these improvements arise and how they depend on graph construction choices.
BraTS 2021 provides multi-parametric MRI at isotropic 1mm resolution for glioma segmentation (WT/TC/ET) (Baid et al., 2021; Bonato et al., 2025). KiTS23 includes portal venous and nephrogenic CT phases with substantial inter-case variability in slice count and spacing (Heller et al., 2023). LiTS exhibits intentionally heterogeneous acquisition parameters, with slice thickness ranging from 0.45–6.0mm (Bilic et al., 2023). Throughout the paper, we abbreviate BraTS, KiTS, and LiTS to indicate these datasets.
SEMIR reduces inference from ∼ 107 to ∼ 103 supernodes—a 104 × reduction in node count compared to dense methods (Appendix E). This complexity scales with image structure rather than resolution, enabling efficient processing of high-resolution volumes.
4.2. Implementation Details Experiments use an NVIDIA Tesla T4 GPU (16GB) with PyTorch. We train a separate binary model per target structure 7
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation Table 3. Comparison across BraTS 2021, KiTS23, and LiTS benchmarks. Key: Best, Second. BraTS regions: TC (Tumor Core), WT (Whole Tumor), ET (Enhancing Tumor). KiTS regions: KTC (Kidney+Tumor+Cyst), TC (Tumor+Cyst), T (Tumor). LiTS regions: LT (Liver+Tumor), T (Tumor). 1 (El Badaoui et al., 2025) 2 (Tobias et al., 2025) 3 (Alwadee et al., 2025) 4 (Bonato et al., 2025) 5 (Nowakowski & Patel, 2025) 6 (Zhang et al., 2025a) 7 (Alonso-Monsalve et al., 2025) 8 (Uhm et al., 2023) 9 (Liu et al., 2025) 10 (Yang et al., 2025) 11 (Ren et al., 2025) 12 (Zheng et al., 2025) 13 (Zhang et al., 2025b) 14 (Saifullah & Dreżewski, 2025) 15 (Perera et al., 2024) 16 (Lilhore et al., 2025) 17 (Lei et al., 2022) 18 (Li et al., 2018) 19 (Myronenko et al., 2024) 20 (Kaczmarska & Majek, 2024) 21 (Pandey et al., 2024) 22 (Liao et al., 2024) 23 (xmed lab, 2024)
BraTS 2021 TC WT
KiTS23 KTC
TC
T
0.943 0.781 0.958 0.948 0.762 0.831
0.725 —– 0.856 0.776 0.439 0.130
0.693 0.613 0.803 0.738 0.343 0.260
Challenge Winners Auto3DSeg19 0.926 Uhm et al.8 0.948
0.784 0.776
0.751 0.738
SEMIR
0.861
0.819
ET
Method
0.823 0.851 0.802 0.915 0.902 0.927
0.748 0.792 0.781 0.835 0.839 0.845
ConvOccNet5 AlignUNet6 Alonso7 3DU-Net Ens.8 Swin UNETR20 Multi-Planner21
Transformer / State-Space GTMamba13 0.940 0.943 PSO-UNet14 —– 0.958 SegFormer15 0.822 0.899 MS3D-CNN16 0.912 0.928
0.884 —– 0.742 0.857
SEMIR (ours)
0.894
Method
Swin UNETR1 3DCATBraTS1 E-CATBraTS1 PAU-Net2 LATUP-Net3 Ext. nnUNet4
0.682 0.784 0.726 0.879 0.895 0.878
0.941
0.920
0.945
Table 4. Controlled binary target-vs-rest comparison against nnUNet under identical splits. Runtime reports end-to-end training time in the practical hardware regime used for each method. †SEMIR completed on a 16GB T4 GPU; nnU-Net required A100-class hardware to complete in practical time.
Dataset
Target
nnU-Net DSC
SEMIR DSC
Time
BraTS BraTS KiTS LiTS
ET TC T T
0.812 0.829 0.720 0.733
0.894±.006 0.941±.002 0.819±.006 0.891±.007
43h/2.5h† 39h/1.6h† 19h/0.8h† 11h/0.6h†
LiTS Method
LT
T
Yang et al.10 MA-UNet10 Lgma-net11 DefED-Net17 Swin-UNet12 Li & Zhao12
0.971 0.960 0.977 0.963 0.967 0.967
0.721 0.729 0.874 0.875 0.857 0.865
Transformer / Hybrid HDenseUNet18 0.961 MAcGAN23 0.970 SwinUNETR23 0.867 MedNeXtL23 0.946
0.722 0.785 0.742 0.779
SEMIR
0.891
0.959
Table 5. Ablation summary. BraTS ET Dice and NWPU VHR-10 IoU. Ablation
BraTS ET
NWPU
Full SEMIR
0.894
0.862
Minor operations
No edge contraction No edge deletion No node deletion
0.441 0.719 0.812
0.408 0.681 0.749
Parameter optimization
Learned (|Dfew |=5) Fixed (best manual)
0.894 0.837
0.789 0.763
Features
No edge features No spatial features
0.725 0.661
0.741 0.629
4.4. Ablation Studies Ablation studies (Section 4.4) confirm that our method transfers to non-medical domains: on NWPU VHR-10 aerial imagery, the learned minor achieves 0.862 small-object IoU, demonstrating that boundary-aligned minor construction is not specific to volumetric medical data.
We ablate key components on BraTS (minority classes: ET, TC) and a 100-sample subset of NWPU VHR-10 (Cheng et al., 2014), an aerial detection benchmark, to verify that SEMIR’s minor construction generalizes beyond medical imaging to 2D RGB imagery with small, sparse targets. Full results in Appendix D.
Figures 1 and 2 visualize SEMIR’s representation learning. The pipeline figure shows how the boundary-aligned minor H reduces a KiTS volume to a sparse supernode graph while preserving structure for accurate tumor delineation. The boundary alignment figure demonstrates the effect of few-shot optimization: Θopt produces supernodes whose boundaries track semantic edges, whereas naive parameters Θinit cut arbitrarily through the tumor boundary.
Key findings. Disabling edge contraction causes severe fragmentation (ET Dice drops 51%). Few-shot optimization outperforms manual tuning even with 5 samples. Spatial features (compactness, elongation) are essential for irregular structures, with 24–27% degradation when removed. Higher N-connectivity (18 vs. 6) improves small-object IoU by 3–5% at 2.5× runtime cost; L∞ norm shows strongest robustness on multi-channel data. 8
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
5. Conclusion
Analysis and Machine Intelligence, 34(11):2274–2282, 2012. doi: 10.1109/TPAMI.2012.120.
We introduced SEMIR, a framework that constructs a learned graph minor over the voxel lattice, enabling segmentation inference on a compact, boundary-aligned representation with exact lifting back to voxel-level predictions. Few-shot optimization of minor construction parameters replaces manual superpixel tuning with a principled, data-driven procedure. Across three tumor segmentation benchmarks, SEMIR yields consistent improvements on structure-specific Dice for clinically critical targets—the regime where multi-class methods suffer from imbalanceinduced gradient attenuation. By reducing each task to binary segmentation on a target-adapted graph, SEMIR sidesteps the balancing problem entirely—while reducing inference to 1–3% of the original node count.
Alonso-Monsalve, S., Whitehead, L. H., Aurisano, A., and Sanchez, L. E. Submanifold sparse convolutional networks for automated 3d segmentation of kidneys and kidney tumours in computed tomography. arXiv preprint arXiv:2511.04334, 2025. Alwadee, E. J., Sun, X., Qin, Y., and Langbein, F. C. Latupnet: A lightweight 3d attention u-net with parallel convolutions for brain tumor segmentation. Computers in Biology and Medicine, 184:109353, 2025. ISSN 0010-4825. doi: https://doi.org/10.1016/j.compbiomed.2024.109353. URL https://www.sciencedirect.com/ science/article/pii/S0010482524014380. Baid, U., Ghodasara, S., Mohan, S., Bilello, M., Calabrese, E., Colak, E., Farahani, K., Kalpathy-Cramer, J., Kitamura, F. C., Pati, S., Prevedello, L. M., Rudie, J. D., Sako, C., Shinohara, R. T., Bergquist, T., Chai, R., Eddy, J., Elliott, J., Reade, W., Schaffter, T., Yu, T., Zheng, J., Moawad, A. W., Coelho, L. O., McDonnell, O., Miller, E., Moron, F. E., Oswood, M. C., Shih, R. Y., Siakallis, L., Bronstein, Y., Mason, J. R., Miller, A. F., Choudhary, G., Agarwal, A., Besada, C. H., Derakhshan, J. J., Diogo, M. C., Do-Dai, D. D., Farage, L., Go, J. L., Hadi, M., Hill, V. B., Iv, M., Joyner, D., Lincoln, C., Lotan, E., Miyakoshi, A., Sanchez-Montano, M., Nath, J., Nguyen, X. V., Nicolas-Jilwan, M., Jimenez, J. O., Ozturk, K., Petrovic, B. D., Shah, C., Shah, L. M., Sharma, M., Simsek, O., Singh, A. K., Soman, S., Statsevych, V., Weinberg, B. D., Young, R. J., Ikuta, I., Agarwal, A. K., Cambron, S. C., Silbergleit, R., Dusoi, A., Postma, A. A., Letourneau-Guillon, L., Perez-Carrillo, G. J. G., Saha, A., Soni, N., Zaharchuk, G., Zohrabian, V. M., Chen, Y., Cekic, M. M., Rahman, A., Small, J. E., Sethi, V., Davatzikos, C., Mongan, J., Hess, C., Cha, S., Villanueva-Meyer, J., Freymann, J. B., Kirby, J. S., Wiestler, B., Crivellaro, P., Colen, R. R., Kotrotsou, A., Marcus, D., Milchenko, M., Nazeri, A., FathallahShaykh, H., Wiest, R., Jakab, A., Weber, M.-A., Mahajan, A., Menze, B., Flanders, A. E., and Bakas, S. The rsnaasnr-miccai brats 2021 benchmark on brain tumor segmentation and radiogenomic classification, 2021. URL https://arxiv.org/abs/2107.02314.
Limitations include sensitivity to boundary statistics in the few-shot set and evaluation restricted to volumetric CT/MRI. The current formulation decouples minor construction from downstream prediction; this modularity means the learned minor can serve as a preprocessing stage for any downstream model (CNNs, transformers, GNNs), though end-toend joint optimization remains an obvious target for future work. Extension to additional modalities (histopathology, ultrasound) and theoretical analysis of minor structure under varying acquisition conditions warrant further exploration.
Impact Statement This work aims to improve segmentation of minority structures in medical images, with potential applications in tumor detection and treatment planning. Two considerations merit attention: (1) the benchmarks used do not represent a global sample; performance on underrepresented populations requires further validation. (2) SEMIR’s node deletion mechanism excludes image regions outside learned thresholds, which improves efficiency but could discard atypical pathologies. Clinical implementations should maintain radiologist oversight and preserve access to full-resolution data. We believe improved minority-structure segmentation offers meaningful clinical benefit when deployed with appropriate validation and safeguards.
Acknowledgments Luke Miller and Yugyung Lee acknowledge support from the National Science Foundation (NSF) under Award No. 2152057.
Bergstra, J., Bardenet, R., Bengio, Y., and Kégl, B. Algorithms for hyper-parameter optimization. In Shawe-Taylor, J., Zemel, R., Bartlett, P., Pereira, F., and Weinberger, K. (eds.), Advances in Neural Information Processing Systems, volume 24. Curran Associates, Inc., 2011. URL https://proceedings.neurips. cc/paper_files/paper/2011/file/ 86e8f7ab32cfd12577bc2619bc635690-Paper. pdf.
References Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P., and Süsstrunk, S. Slic superpixels compared to state-of-theart superpixel methods. IEEE Transactions on Pattern 9
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Berrouachedi, A., Jaziri, R., and Bernard, G. Deep extremely randomized trees. In International Conference on Neural Information Processing, pp. 717–729. Springer, 2019.
002. URL https://www.sciencedirect.com/ science/article/pii/S0924271614002524. Demaine, E. D., Hajiaghayi, M. T., and Kawarabayashi, K.-i. Algorithmic graph minor theory: Decomposition, approximation, and coloring. In 46th Annual IEEE Symposium on Foundations of Computer Science (FOCS’05), pp. 637–646. IEEE, 2005.
Bilic, P., Christ, P., Li, H. B., Vorontsov, E., Ben-Cohen, A., Kaissis, G., Szeskin, A., Jacobs, C., Mamani, G. E. H., Chartrand, G., Lohöfer, F., Holch, J. W., Sommer, W., Hofmann, F., Hostettler, A., Lev-Cohain, N., Drozdzal, M., Amitai, M. M., Vivanti, R., Sosna, J., Ezhov, I., Sekuboyina, A., Navarro, F., Kofler, F., Paetzold, J. C., Shit, S., Hu, X., Lipková, J., Rempfler, M., Piraud, M., Kirschke, J., Wiestler, B., Zhang, Z., Hülsemeyer, C., Beetz, M., Ettlinger, F., Antonelli, M., Bae, W., Bellver, M., Bi, L., Chen, H., Chlebus, G., Dam, E. B., Dou, Q., Fu, C.-W., Georgescu, B., i Nieto, X. G., Gruen, F., Han, X., Heng, P.-A., Hesser, J., Moltz, J. H., Igel, C., Isensee, F., Jäger, P., Jia, F., Kaluva, K. C., Khened, M., Kim, I., Kim, J.-H., Kim, S., Kohl, S., Konopczynski, T., Kori, A., Krishnamurthi, G., Li, F., Li, H., Li, J., Li, X., Lowengrub, J., Ma, J., Maier-Hein, K., Maninis, K.-K., Meine, H., Merhof, D., Pai, A., Perslev, M., Petersen, J., Pont-Tuset, J., Qi, J., Qi, X., Rippel, O., Roth, K., Sarasua, I., Schenk, A., Shen, Z., Torres, J., Wachinger, C., Wang, C., Weninger, L., Wu, J., Xu, D., Yang, X., Yu, S. C.-H., Yuan, Y., Yue, M., Zhang, L., Cardoso, J., Bakas, S., Braren, R., Heinemann, V., Pal, C., Tang, A., Kadoury, S., Soler, L., van Ginneken, B., Greenspan, H., Joskowicz, L., and Menze, B. The liver tumor segmentation benchmark (lits). Medical Image Analysis, 84:102680, 2023. ISSN 1361-8415. doi: https://doi.org/10.1016/j.media.2022.102680. URL https://www.sciencedirect.com/ science/article/pii/S1361841522003085.
Ding, K., Zhou, M., Wang, H., Zhang, S., and Metaxas, D. N. Spatially aware graph neural networks and cross-level molecular profile prediction in colon cancer histopathology: a retrospective multi-cohort study. The Lancet Digital Health, 4(11):e787–e795, 2022. El Badaoui, R., Bonmati Coll, E., Psarrou, A., Asaturyan, H. A., and Villarini, B. Enhanced catbrats for brain tumour semantic segmentation. Journal of Imaging, 11(1), 2025. ISSN 2313-433X. doi: 10.3390/jimaging11010008. URL https://www.mdpi.com/2313-433X/11/ 1/8. Felzenszwalb, P. F. and Huttenlocher, D. P. Efficient graphbased image segmentation. International journal of computer vision, 59(2):167–181, 2004. Gao, Y., Jiang, Y., Peng, Y., Yuan, F., Zhang, X., and Wang, J. Medical image segmentation: A comprehensive review of deep learning-based methods. Tomography, 11(5), 2025. ISSN 2379-139X. doi: 10.3390/ tomography11050052. URL https://www.mdpi. com/2379-139X/11/5/52. Ghankot, R. S., Singh, M., Desroches, S. T., Jester, N., Mahajan, A., Lorr, S., Buono, F. D., Wiznia, D. H., Johnson, M. H., and Tommasini, S. M. Evaluating the effect of voxel size on the accuracy of 3d volumetric analysis measurements of brain tumors. Frontiers in Radiology, 5: 1618261, 2025.
Bonato, B., Nanni, L., and Bertoldo, A. Advancing precision: A comprehensive review of mri segmentation datasets from brats challenges (2012–2025). Sensors, 25 (6), 2025. ISSN 1424-8220. doi: 10.3390/s25061838. URL https://www.mdpi.com/1424-8220/25/ 6/1838.
Guan, B., Chu, G., Wang, Z., Li, J., and Yi, B. Instance-level semantic segmentation of nuclei based on multimodal structure encoding. BMC bioinformatics, 26(1):42, 2025. Heller, N., Isensee, F., Trofimova, D., Tejpaul, R., Zhao, Z., Chen, H., Wang, L., Golts, A., Khapun, D., Shats, D., Shoshan, Y., Gilboa-Solomon, F., George, Y., Yang, X., Zhang, J., Zhang, J., Xia, Y., Wu, M., Liu, Z., Walczak, E., McSweeney, S., Vasdev, R., Hornung, C., Solaiman, R., Schoephoerster, J., Abernathy, B., Wu, D., Abdulkadir, S., Byun, B., Spriggs, J., Struyk, G., Austin, A., Simpson, B., Hagstrom, M., Virnig, S., French, J., Venkatesh, N., Chan, S., Moore, K., Jacobsen, A., Austin, S., Austin, M., Regmi, S., Papanikolopoulos, N., and Weight, C. The kits21 challenge: Automatic segmentation of kidneys, renal tumors, and renal cysts in corticomedullary-phase ct, 2023. URL https://arxiv.org/abs/2307. 01984.
Brussee, S., Buzzanca, G., Schrader, A. M., and Kers, J. Graph neural networks in histopathology: Emerging trends and future directions. Medical Image Analysis, 101:103444, 2025. ISSN 1361-8415. doi: https://doi.org/10.1016/j.media.2024.103444. URL https://www.sciencedirect.com/ science/article/pii/S1361841524003694. Cheng, G., Han, J., Zhou, P., and Guo, L. Multiclass geospatial object detection and geographic image classification based on collection of part detectors. ISPRS Journal of Photogrammetry and Remote Sensing, 98:119–132, 2014. ISSN 09242716. doi: https://doi.org/10.1016/j.isprsjprs.2014.10. 10
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Hosseini, S. M. Pixel-wise modulated dice loss for medical image segmentation, 2025. URL https://arxiv. org/abs/2506.15744. Hu*, W., Liu*, B., Gomes, J., Zitnik, M., Liang, P., Pande, V., and Leskovec, J. Strategies for pre-training graph neural networks. In International Conference on Learning Representations, 2020. URL https://openreview. net/forum?id=HJlWWJSFDH. Jiangtao, W., Ruhaiyem, N. I. R., and Panpan, F. A comprehensive review of u-net and its variants: Advances and applications in medical image segmentation. IET Image Processing, 19(1): e70019, 2025. doi: https://doi.org/10.1049/ipr2.70019. URL https://ietresearch.onlinelibrary. wiley.com/doi/abs/10.1049/ipr2.70019. Kaczmarska, M. and Majek, K. 3d segmentation of kidneys, kidney tumors and cysts on ct images - kits23 challenge. In Heller, N., Wood, A., Isensee, F., Rädsch, T., Teipaul, R., Papanikolopoulos, N., and Weight, C. (eds.), Kidney and Kidney Tumor Segmentation, pp. 149–155, Cham, 2024. Springer Nature Switzerland. ISBN 978-3-03154806-2.
A. M. Ag-ms3d-cnn multiscale attention guided 3d convolutional neural network for robust brain tumor segmentation across mri protocols. Scientific Reports, 15(1): 24306, 2025. Liu, H., Wang, H., Wu, Y., and Xing, L. Superpixel region merging based on deep network for medical image segmentation. ACM Trans. Intell. Syst. Technol., 11(4), May 2020. ISSN 2157-6904. doi: 10.1145/3386090. URL https://doi.org/10.1145/3386090. Liu, Z., Han, K., Ma, S., Zhu, Y., Chen, J., Lyu, C., Qiu, X., Qian, C., Song, Y., Liu, Y., et al. Limt: A multi-task liver image benchmark dataset. arXiv preprint arXiv:2511.19889, 2025. Lovász, L. Graph minor theory. Bulletin of the American Mathematical Society, 43(1):75–86, 2006. Luo, Z., Zhu, X., Zhang, L., and Sun, B. Rethinking u-net: Task-adaptive mixture of skip connections for enhanced medical image segmentation. Proceedings of the AAAI Conference on Artificial Intelligence, 39 (6):5874–5882, Apr. 2025. doi: 10.1609/aaai.v39i6. 32627. URL https://ojs.aaai.org/index. php/AAAI/article/view/32627.
Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization, 2017. URL https://arxiv.org/abs/ 1412.6980.
Mienye, I. D. and Viriri, S. Graph neural networks in medical imaging: Methods, applications and future directions. Information, 16(12), 2025. ISSN 2078-2489. doi: 10.3390/info16121051. URL https://www.mdpi. com/2078-2489/16/12/1051.
Kostrykin, L. and Rohr, K. Superadditivity and convex optimization for globally optimal cell segmentation using deformable shape models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(3):3831– 3847, 2022.
Myronenko, A., Yang, D., He, Y., and Xu, D. Automated 3d segmentation of kidneys and tumors in miccai kits 2023 challenge. In Heller, N., Wood, A., Isensee, F., Rädsch, T., Teipaul, R., Papanikolopoulos, N., and Weight, C. (eds.), Kidney and Kidney Tumor Segmentation, pp. 1–7, Cham, 2024. Springer Nature Switzerland. ISBN 978-3031-54806-2.
Lei, T., Wang, R., Zhang, Y., Wan, Y., Liu, C., and Nandi, A. K. Defed-net: Deformable encoder-decoder network for liver and liver tumor segmentation. IEEE Transactions on Radiation and Plasma Medical Sciences, 6(1):68–78, 2022. doi: 10.1109/TRPMS.2021.3059780.
Neubert, P. and Protzel, P. Compact watershed and preemptive slic: On improving trade-offs of superpixel segmentation algorithms. In 2014 22nd International Conference on Pattern Recognition, pp. 996–1001, 2014. doi: 10.1109/ICPR.2014.181.
Li, X., Chen, H., Qi, X., Dou, Q., Fu, C.-W., and Heng, P.-A. H-denseunet: Hybrid densely connected unet for liver and tumor segmentation from ct volumes. IEEE Transactions on Medical Imaging, 37(12):2663–2674, 2018. doi: 10.1109/TMI.2018.2845918.
Nowakowski, L. and Patel, R. Convolutional occupancy networks for medical imaging with applications to the kits23 challenge. In IEEE-EMBS International Conference on Biomedical and Health Informatics 2025, 2025.
Liao, J., Wang, H., Gu, H., and Cai, Y. Liver tumor segmentation method combining multi-axis attention and conditional generative adversarial networks. PLOS ONE, 19(12):1–24, 12 2024. doi: 10.1371/journal. pone.0312105. URL https://doi.org/10.1371/ journal.pone.0312105.
Pandey, S., Toshali, Perslev, M., and Dam, E. B. Advancing kidney, kidney tumor, cyst segmentation: A multiplanner u-net approach for the kits23 challenge. In Heller, N., Wood, A., Isensee, F., Rädsch, T., Teipaul, R., Papanikolopoulos, N., and Weight, C. (eds.), Kidney and
Lilhore, U. K., Sunder, R., Simaiya, S., Alsafyani, M., Monish Khan, M., Alroobaea, R., Alsufyani, H., and Baqasah, 11
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Kidney Tumor Segmentation, pp. 143–148, Cham, 2024. Springer Nature Switzerland. ISBN 978-3-031-54806-2.
xmed lab. TriALS: MICCAI 2024/2025 segmentation benchmark. https://github.com/xmed-lab/ TriALS, 2024. Accessed: 2026-01-28.
Peng, Y., Chen, D. Z., and Sonka, M. U-net v2: Rethinking the skip connections of u-net for medical image segmentation. In 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI), pp. 1–5, 2025. doi: 10.1109/ISBI60581.2025.10980742.
Xu, K., Hu, W., Leskovec, J., and Jegelka, S. How powerful are graph neural networks? In International Conference on Learning Representations, 2019. URL https:// openreview.net/forum?id=ryGs6iA5Km.
Perera, S., Navard, P., and Yilmaz, A. Segformer3d: An efficient transformer for 3d medical image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 4981–4988, June 2024.
Yang, B., Zhang, J., Lyu, Y., and Zhang, J. Automatic computed tomography image segmentation method for liver tumor based on a modified tokenized multilayer perceptron and attention mechanism. Quantitative Imaging in Medicine and Surgery, 15(3):2385, 2025.
Ren, W., Li, B., Peng, H., and Wang, J. Lgma-net: liver and tumor segmentation methods based on local–global feature mergence and attention mechanisms. Signal, Image and Video Processing, 19(1):43, 2025.
Yeung, M., Sala, E., Schönlieb, C.-B., and Rundo, L. Unified focal loss: Generalising dice and cross entropy-based losses to handle class imbalanced medical image segmentation. Computerized Medical Imaging and Graphics, 95:102026, 2022. ISSN 0895-6111. doi: https://doi.org/10.1016/j.compmedimag.2021.102026. URL https://www.sciencedirect.com/ science/article/pii/S0895611121001750.
Robertson, N. and Seymour, P. D. Graph minors. xx. wagner’s conjecture. Journal of Combinatorial Theory, Series B, 92(2):325–357, 2004. Saifullah, S. and Dreżewski, R. Particle swarm-optimized u-net framework for precise multimodal brain tumor segmentation. In Proceedings of the Genetic and Evolutionary Computation Conference Companion, GECCO ’25 Companion, pp. 323–326. ACM, July 2025. doi: 10.1145/3712255.3726561. URL http://dx.doi. org/10.1145/3712255.3726561. Stutz, D., Hermans, A., and Leibe, B. Superpixels: An evaluation of the state-of-the-art. Computer Vision and Image Understanding, 166:1–27, 2018. ISSN 1077-3142. doi: https://doi.org/10.1016/j.cviu.2017.03. 007. URL https://www.sciencedirect.com/ science/article/pii/S1077314217300589.
Ying, R., You, J., Morris, C., Ren, X., Hamilton, W. L., and Leskovec, J. Hierarchical graph representation learning with differentiable pooling. In NeurIPS, 2018. Zhang, H., Bi, Y., Lu, Y., Qi, T., Jia, X., Zhao, Z., Yu, N., and Li, K. Exploring semi-supervised domain adaptation for precise kidney tumor segmentation. In 2025 IEEE International Conference on Real-time Computing and Robotics (RCAR), pp. 515–520. IEEE, 2025a. Zhang, Y., Zhang, M., Zhang, J., Shen, Y., and Niu, D. Gtmamba: Graph tri-orientated mamba network for 3d brain tumor segmentation. International Journal of Imaging Systems and Technology, 35 (3):e70111, 2025b. doi: https://doi.org/10.1002/ima. 70111. URL https://onlinelibrary.wiley. com/doi/abs/10.1002/ima.70111. e70111 IMA-25-181.R1.
Tobias, S., Alexander, B., Simon, W., Maria, T. A., and Elmar W., L. Segmenting the non-enhancing compartment of brain tumor mris. In 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pp. 1–5, 2025. doi: 10.1109/EMBC58623.2025.11254161.
Zheng, Y., Tian, B., Yu, S., Yang, X., Yu, Q., Zhou, J., Jiang, G., Zheng, Q., Pu, J., and Wang, L. Adaptive boundary-enhanced dice loss for image segmentation. Biomedical Signal Processing and Control, 106:107741, 2025. ISSN 1746-8094. doi: https://doi.org/10.1016/j.bspc.2025.107741. URL https://www.sciencedirect.com/ science/article/pii/S1746809425002526.
Uhm, K.-H. et al. Exploring 3d u-net training configurations and post-processing strategies for the miccai 2023 kidney and tumor segmentation challenge. arXiv preprint arXiv:2312.05528, 2023. URL https:// arxiv.org/abs/2312.05528. Wang, M., Liu, X., Gao, Y., Ma, X., and Soomro, N. Q. Superpixel segmentation: A benchmark. Signal Processing: Image Communication, 56:28–39, 2017. ISSN 0923-5965. doi: https://doi.org/10.1016/j.image.2017.04. 007. URL https://www.sciencedirect.com/ science/article/pii/S0923596517300735. 12
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
A. Notation Symbol
Description
I ∈ RH×W ×D×C Ij,k,l ∈ RC Y ∈ {0, . . . , K − 1}H×W ×D YB ∈ {0, 1}H×W ×D Ŷ ∈ {0, . . . , K − 1}H×W ×D T ∈ {0, 1}(2H−1)×(2W −1)×(2D−1) G = (V (G), E(G)) H = (V (H), E(H), X(H), F (H)) X(H) ∈ Rdx ×|V (H)| au bu
Input medical image volume (height × width × depth × channels) Intensity vector at voxel coordinate (j, k, l) Ground-truth voxel class labels Binary ground-truth boundary map Predicted voxel class map Expanded binary tensor encoding graph nodes and edges Original N-connected grid graph of I (N ∈ {6, 10, 18, 26}) Graph minor of G Node feature matrix Area: voxel count of supernode u Boundary length: number of exposed edges of supernode u
36πa2
compu = b3 +εu u
3
du ∈ R elongu p⋆u σu ∈ RC Σu ∈ RC×C F (H) ∈ Rdf ×|E(H)| Θ = {ψ, α, β} S(T, Θ) 7→ (T, H) SB (T, Θ) 7→ ŶB R(Dfew , Θ) Dfew ⊂ D D = {(I (n) , Y (n) )}N n=1 Lift(H, ŶH , T ) 7→ Ŷ
3D compactness of supernode u (ε > 0 for stability) Dominant axis: unit eigenvector of largest eigenvalue of spatial covariance Elongation ratio: derived from eigenvalues of spatial covariance Canonical voxel: lexicographically smallest coordinate in supernode u Per-channel standard deviation of intensities in supernode u Covariance matrix of intensity vectors in supernode u Edge feature matrix (log-ratios of corresponding node features) Parameters for graph minor generation S(T, Θ) Graph minor generation function Boundary voxel map extractor Few-shot parameter optimization Few-shot tuning subset of dataset Full dataset Mapping of supernode predictions ŶH back to voxel grid
B. Algorithms Algorithm 1 G ET C OPRIME: Find coprime step for pseudo-random traversal Require: Dimension size n, divisor d Ensure: Step size s coprime to n s ← ⌊n/ max(d, 1)⌋ s ← clamp(s, 1, n − 1) for i = s to n − 1 do if gcd(n, i) = 1 then Return i end if end for for i = s − 1 to 1 do if gcd(n, i) = 1 then Return i end if end for Return 1 13
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Image Volume I
Voxel Labels Y
Input
Boundary Voxels YB
Grid Graph G
Minor Gen S(·)
Graph Minor H
Graph Neural Net GN N (·)
Rep Learner R(·)
Opt. Params Θ
Bijection Tensor T
Node7→ Voxel Lift(·)
Representation Learning
Node Labels ŶH
Voxel Predictions Ŷ
Loss L(·)
Segmentation
Figure 3. SEMIR pipeline overview. From the input volume I, a boundary-aware graph minor H is constructed using few-shot-optimized parameters Θ. A graph neural network yields supernode predictions ŶH , which are bijectively lifted via tensor T to produce the final voxel segmentation Ŷ . Ground-truth voxel labels Y provide supervision through the loss and few-shot boundary alignment.
Algorithm 2 M INOR C ONSTRUCTION: Graph minor construction S(I, Θ) Require: Image volume I ∈ RH×W ×D×C , parameters Θ = {ψ, α, β} Ensure: Tensor T , graph minor H = (V, E, X, F ) {Initialize expanded tensor} T ← 1(2H−1)×(2W −1)×(2D−1) {Compute coprime steps for pseudo-random traversal} (sh , sw , sd ) ← (G ET C OPRIME(H, d), G ET C OPRIME(W, d), G ET C OPRIME(D, d)) (r0 , c0 , l0 ) ← random start position V ← ∅, X ← ∅ {Phase 1: Edge contraction via flood-fill with node deletion} for i = 0 to H − 1 do for j = 0 to W − 1 do for k = 0 to D − 1 do (r, c, l) ← ((r0 + i · sh ) mod H, (c0 + j · sw ) mod W, (l0 + k · sd ) mod D) if T [2r, 2c, 2l] is visited then continue end if (T, xu , Pu ) ← F LOOD F ILL C ONTRACT(I, T, (r, c, l), ψ, α) {Node deletion check} if βmin < |Pu | < βmax then V ← V ∪ {u}, X ← X ∪ {xu } else Mark all voxels in Pu as deleted in T end if end for end for end for {Phase 2: Build edge list and edge features from surviving nodes} (E, F ) ← E XTRACT E DGES(T, V, X) Return (T, H = (V, E, X, F )) 14
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Algorithm 3 F LOOD F ILL C ONTRACT: Region growing with contraction and edge deletion Require: Image I, tensor T , seed voxel (r, c, l), merge threshold ψ, cut threshold α Ensure: Updated T , node features xu , voxel set Pu vseed ← I[r, c, l] {Seed intensity (canonical voxel)} Initialize stack S ← {(2r, 2c, 2l)}, voxel set Pu ← ∅ Initialize running statistics: area, sums for coordinates, intensities, products while S ̸= ∅ do Pop (y, x, z) from S if T [y, x, z] is visited then continue end if Mark T [y, x, z] as visited (r′ , c′ , l′ ) ← (y/2, x/2, z/2) {Original voxel coordinates} Pu ← Pu ∪ {(r′ , c′ , l′ )} Update running statistics with (r′ , c′ , l′ ) and I[r′ , c′ , l′ ] Update canonical voxel if (r′ , c′ , l′ ) is lexicographically smaller for each N-connected neighbor direction δ do (y ′ , x′ , z ′ ) ← (y, x, z) + 2δ {Neighbor node position} if (y ′ , x′ , z ′ ) out of bounds or T [y ′ , x′ , z ′ ] is merged then continue end if diff ← ∥vseed − I[y ′ /2, x′ /2, z ′ /2]∥n if diff ≤ ψ then Mark T [y ′ , x′ , z ′ ] as merged Push (y ′ , x′ , z ′ ) onto S else if diff ≥ α then (ey , ex , ez ) ← (y, x, z) + δ {Edge position} Mark T [ey , ex , ez ] as edge-deleted end if Mark T [y, x, z] as boundary-adjacent end if end for if T [y, x, z] is boundary-adjacent then Increment boundary length counter end if end while xu ← C OMPUTE N ODE F EATURES(running statistics, Pu ) Return (T, xu , Pu )
15
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Algorithm 4 E XTRACT E DGES: Build edge list and features from tensor Require: Tensor T , node set V , node features X Ensure: Edge set E, edge features F E ← ∅, F ← ∅ for each node u ∈ V do (r, c, l) ← canonical voxel of u for each N-connected neighbor direction δ do (ey , ex , ez ) ← (2r, 2c, 2l) + δ {Edge position} if T [ey , ex , ez ] is not edge-deleted then (y ′ , x′ , z ′ ) ← (2r, 2c, 2l) + 2δ {Neighbor node position} v ← node containing voxel (y ′ /2, x′ /2, z ′ /2) if v ∈ V and v ̸= u and (u, v) ∈ / E then E ← E ∪ {(u, v)} fuv ← C OMPUTE E DGE F EATURES(Xu , Xv ) F ← F ∪ {fuv } end if end if end for end for Return (E, F )
Algorithm 5 F EW S HOT O PTIMIZATION: Parameter optimization R(Dfew , Θ) Require: Few-shot dataset Dfew , parameter bounds Θ, iterations Niter , initial samples Ninit Ensure: Optimized parameters Θopt {Initialize surrogate with random samples} H ← ∅ {History of (θ, loss) pairs} for i = 1 to Ninit do θ ∼ Uniform(Θ) L ← E VALUATE B OUNDARY L OSS(θ, Dfew ) H ← H ∪ {(θ, L)} end for {Sequential model-based optimization} for i = 1 to Niter do Fit ExtraTrees surrogate fˆ on H θnext ← arg maxθ EI(θ; fˆ, H) {Expected Improvement} L ← E VALUATE B OUNDARY L OSS(θnext , Dfew ) H ← H ∪ {(θnext , L)} end for Θopt ← arg min(θ,L)∈H L Return Θopt 16
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Algorithm 6 E VALUATE B OUNDARY L OSS: Compute boundary alignment loss Require: Parameters θ, few-shot dataset Dfew Ensure: Mean boundary loss L̄ Ltotal ← 0 for each (I, Y ) ∈ Dfew do YB ← E XTRACT G ROUND T RUTH B OUNDARY(Y ) {Voxels adjacent to different class} (T, ) ← M INOR C ONSTRUCTION(I, θ) {Features not needed} ŶB ← E XTRACT M INOR B OUNDARY(T ) {Boundary-adjacent voxels in T } Ltotal ← Ltotal + (1 − DSC(ŶB , YB )) end for Return L̄ ← Ltotal /|Dfew |
Algorithm 7 L IFT: Map supernode predictions to voxel grid Require: Tensor T , node predictions ŶH , node set V with canonical voxels Ensure: Voxel predictions Ŷ ∈ {0, . . . , K − 1}H×W ×D Ŷ ← 0H×W ×D {Background default for deleted nodes} for each node u ∈ V do (r, c, l) ← canonical voxel of u Initialize stack S ← {(2r, 2c, 2l)}, visited set V ← ∅ while S ̸= ∅ do Pop (y, x, z) from S if (y, x, z) ∈ V then continue end if V ← V ∪ {(y, x, z)} Ŷ [y/2, x/2, z/2] ← ŶH [u] for each N-connected neighbor direction δ do (ey , ex , ez ) ← (y, x, z) + δ {Edge position} (y ′ , x′ , z ′ ) ← (y, x, z) + 2δ {Neighbor node position} if in bounds and T [ey , ex , ez ] not edge-deleted and T [y ′ , x′ , z ′ ] is merged then Push (y ′ , x′ , z ′ ) onto S end if end for end while end for Return Ŷ
C. Full Description of Features While node features capture absolute properties of individual supernodes, edge features encode relative differences between adjacent supernodes, promoting scale- and rotation-invariant representations suitable for the graph neural network. For each edge e = (u, v) ∈ E(H), we first order the incident supernodes by area: u− := arg min{au , av }.
u+ := arg max{au , av },
(11)
For any scalar node feature g(·) (e.g., au , bu , compu , elongu ), we compute a scale-invariant log-ratio: rg (e) := log
g(u+ ) + ε , g(u− ) + ε
with ε > 0 for numerical stability. 17
(12)
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
We calculate relative geometric and orientation edge features for each scalar node feature. ∆µ (e) := p
µu − µv ∈ R3 , λu,1 + λv,1 + ε
cos θ(e) := |d⊤ u dv | ∈ [0, 1],
(13) (14)
where µu and µv are the (transiently computed) centroids of supernodes u and v, λu,1 and λv,1 are their respective largest spatial eigenvalues, and du , dv are the dominant axes. For intensity-based features, we use normalized per-channel differences: ∆I (e) := q
I¯u − I¯v s2I,u + s2I,v + ε
∈ RC ,
(15)
where I¯u , I¯v are the mean intensity vectors and s2I,u , s2I,v are the per-channel variance vectors (i.e., element-wise squares of σu , σv ). These relative features (rg (e) for relevant scalars, ∆µ (e), cos θ(e), ∆I (e)) are stored in the edge feature matrix F (H).
D. Extended Ablation Studies D.1. Graph-Minor Construction Operations Table 6 shows the impact of disabling individual operations. Edge contraction prevents severe fragmentation; edge deletion separates minority structures; node deletion prunes noise and oversized background supernodes. Uniform grid downsampling (intensity-agnostic) serves as a weak baseline. Table 6. Ablation of graph-minor construction operations.
Variant
BraTS ET TC
NWPU VHR-10
Full SEMIR (learned Θ) No edge contraction No edge deletion No node deletion Grid downsampling
0.894 0.441 0.719 0.812 0.393
0.862 0.408 0.681 0.749 0.252
0.941 0.492 0.774 0.837 0.351
D.2. Few-Shot Parameter Optimization Table 7 compares learned Θ against fixed thresholds across few-shot set sizes. Performance saturates around 5 samples for BraTS and 20 samples for multi-channel NWPU. Table 7. Few-shot optimization strategy and set size |Dfew |. Θ strategy
|Dfew |
BraTS ET
NWPU
Learned Learned Learned Learned Fixed (loose) Fixed (tight) Fixed (visual)
1 5 10 20 – – –
0.766 0.894 0.891 0.890 0.499 0.837 0.711
0.641 0.789 0.854 0.862 0.544 0.749 0.763
D.3. Distance Norms and Connectivity Table 8 evaluates distance norms and N-connectivity. Norms are equivalent on single-channel BraTS. On multi-channel NWPU, L∞ shows strongest robustness to channel-specific outliers; higher connectivity improves small-object IoU at moderate runtime cost. 18
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation Table 8. Distance norm and N-connectivity ablation.
Norm
N
BraTS ET DSC Time
NWPU IoU Time
L1 L1 L1 L2 L2 L2 L∞ L∞ L∞
6 10 18 6 10 18 6 10 18
0.894 0.889 0.890 — — — — — —
0.812 0.819 0.825 0.796 0.804 0.813 0.809 0.836 0.849
1.0× 1.6× 2.7× — — — — — —
1.0× 1.8× 2.3× 1.4× 2.5× 3.8× 1.1× 1.9× 2.5×
D.4. Feature Design Removing relative edge features causes ET Dice to drop by 11–19% on BraTS and small-object IoU by 7–14% on NWPU, underscoring their importance for scale- and rotation-invariant relational modeling. Disabling spatial features (compactness, elongation, dominant axis) reduces performance by 24–27% in both domains, with greater impact on irregularly shaped objects.
E. Complexity Analysis Minor construction. Each voxel is visited exactly once during flood-fill traversal, and each edge is examined at most twice. Construction is therefore O(N · HW D) where N ∈ {6, 10, 18, 26} is the grid connectivity—linear in voxel count regardless of the induced minor size. GNN inference. Message passing operates on the minor H, with cost O(L(|V (H)| + |E(H)|)) for L layers. The critical quantity is |V (H)|, which depends on image content rather than resolution: |V (H)| reflects the number of intensityhomogeneous regions satisfying the learned thresholds Θopt . In the degenerate case (all adjacent voxels differ by more than ψ), |V (H)| = O(HW D). In practice, medical images contain large homogeneous regions (background, parenchyma) and |V (H)| ≪ HW D. Table 9 reports empirical supernode counts. Table 9. Empirical supernode counts |V (H)| across benchmarks.
Volume size Voxels |V (G)| Supernodes |V (H)| Reduction factor
BraTS
KiTS
LiTS
240 × 240 × 155 ∼8.9 × 106 2034 ± 187 ∼4400×
variable ∼107 –108 1429 ± 456 ∼104 ×
variable ∼107 –108 1075 ± 297 ∼104 ×
Table 10 compares the number of inference units across segmentation paradigms. Dense methods (U-Net variants, transformers) perform per-voxel prediction, scaling directly with image resolution regardless of structural complexity. Patchbased inference reduces peak memory but does not reduce total computation, as overlapping windows collectively cover all voxels. Classical superpixel methods (SLIC, Felzenszwalb) reduce inference units but use task-agnostic regionization with manually tuned parameters. SEMIR achieves 3–4 orders of magnitude reduction in inference nodes compared to dense methods by learning a boundaryaligned graph minor adapted to the target structure. Empirical supernode counts from Table 9 confirm this reduction: BraTS volumes (∼8.9 × 106 voxels) yield 2034 ± 187 supernodes; KiTS and LiTS volumes (∼107 –108 voxels) yield 1075–1429 supernodes. The downstream GNN operates entirely on this reduced representation, with voxel-level predictions recovered via exact lifting (Theorem 3.2). 19
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation Table 10. Inference complexity across segmentation paradigms for a typical medical volume (256 × 256 × 128 voxels, ∼8.4 × 106 total). SEMIR operates on 3–4 orders of magnitude fewer nodes than dense methods while maintaining exact voxel-level output via lifting. Method
†
Paradigm
Inference Units 6
Unit Type
3D U-Net nnU-Net Swin UNETR TransUNet
Dense Dense Dense + Attention Patch + Dense
8.4 × 10 8.4 × 106 8.4 × 106 8.4 × 106
voxels voxels voxels voxels
Patch-based (64³) Patch-based (128³)
Sliding window Sliding window
8.4 × 106 8.4 × 106
voxels† voxels†
SLIC superpixels Felzenszwalb
Fixed regions Fixed regions
104 –105 104 –105
superpixels superpixels
SEMIR (ours)
Learned minor
103 –104
supernodes
Patch-based methods reduce memory per forward pass but process equivalent total voxels with overlapping windows.
F. Implementation Details Hardware and software. All experiments were conducted in Google Colab on an NVIDIA Tesla T4 GPU (16GB). Preprocessing and learning are implemented in Python using NumPy and PyTorch. GNN architecture. We use a 3-layer GIN variant with edge features (GINE) (Xu et al., 2019; Hu* et al., 2020). Each layer applies message passing with edge-conditioned messages; node and edge features are embedded by MLPs with ReLU activations and hidden dimension d = 128. A final per-node linear layer outputs logits for each supernode. Training details. We train with Adam (Kingma & Ba, 2017) using learning rate 10−3 and no weight decay for up to 200 epochs. Early stopping monitors validation Dice on the binarized foreground target with patience 10. Batch size is 1–4 volumes depending on memory constraints from variable graph sizes. Few-shot dataset sizes.
We use |Dfew | = 5 for BraTS, 10 for KiTS, and 20 for LiTS, drawn from the training split.
Graph construction settings. We evaluate N-connectivity N ∈ {6, 10, 18} in the expanded tensor T , using N = 6 by default. For intensity-distance computations, we evaluate L1 (Manhattan), L2 (Euclidean), and L∞ (Chebyshev) norms. Numerical stabilizers use ε = 10−6 . Node-deletion bounds are initialized with βmin = 1 and βmax = ⌊(HW D)/3⌋, where HW D is the voxel count of the input volume. Reproducibility. Random seeds are fixed to 42 for Python, NumPy, and PyTorch. torch.cuda.manual seed all(42).
20
CUDA seeds are set via