Adaptive Routing for Efficient Diffusion Transformer-Based PNI Prediction Youngung Han1,3 , Dohyun Kweon2,3 , Kyeonghun Kim3 , Hyunsu Go1 , Jina Jeong1 , Suah Park1 , Induk Um4 , Junga Kim1 , Anna Jung1 , Yului Jeong1 , Sungha Park1,5 , Jinyong Jun1 Pa Hong6 , Woo Kyoung Jeong7 , Won Jae Lee6 , Ken Ying-Kai Liao8 , Hyuk-Jae Lee1 , Nam-Joon Kim1,† Seoul National University, Seoul, Republic of Korea 2 Kyung Hee University, Seoul, Republic of Korea 3 OUTTA, Seoul, Republic of Korea 4 Chung-Ang University, Seoul, Republic of Korea 5 Seoul National University School of Medicine, Seoul, Republic of Korea 6 Samsung Changwon Hospital, Changwon, Republic of Korea 7 Samsung Medical Center, Seoul, Republic of Korea 8 NVIDIA AI Technology Center, Taipei, Taiwan † Corresponding author: [email protected]
arXiv:2607.11533v1 [cs.CV] 13 Jul 2026
1
Abstract—Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. However, its preoperative prediction from magnetic resonance imaging (MRI) remains challenging due to subtle imaging features that extend beyond tumor boundaries into surrounding regions. Conventional convolutional neural networks are limited in capturing long-range spatial dependencies. Transformer-based architectures improve global modeling of volumetric MRI by aggregating spatially distributed contextual cues, yet capturing subtle and noise-sensitive patterns in peritumoral regions remains challenging. Diffusion-based classifiers offer an alternative formulation by leveraging denoising-based class scoring to better capture such subtle patterns. However, these approaches introduce substantial computational overhead due to the combination of transformer-based modeling and iterative denoising processes. To address these challenges, we formulate PNI prediction as a diffusion-based classification problem and implement the denoising network using a transformer-based representation. To improve computational efficiency, we introduce adaptive routing across attention heads, spatial tokens, and MLP width. Experimental results demonstrate that the proposed approach achieves an AUC of 0.731 with 257.57 GFLOPs. Index Terms—diffusion transformer, adaptive routing, token selection, computational efficiency, 3D MRI analysis, perineural invasion
I. I NTRODUCTION Perineural invasion (PNI) is a critical prognostic factor closely associated with tumor aggressiveness and poor clinical outcomes [1]. Preoperative prediction of PNI from MRI may assist in surgical planning and inform neoadjuvant treatment strategies; however, it remains challenging due to subtle and spatially diffuse imaging features that often extend beyond tumor boundaries [2]–[4]. PNI involves tumor spread along and around nerve structures, resulting in weak and spatially distributed patterns that are difficult to detect [1], [5]–[8]. CNN-based models primarily rely on local receptive fields [9], which may limit their ability to capture long-range spatial dependencies across slices and surrounding peritumoral regions. Subtle peritumoral cues are often weak and spatially distributed beyond localized regions, making them difficult to capture with purely local representations.
Transformer-based architectures model global interactions via self-attention [10]–[12], aggregating dispersed contextual cues. However, subtle peritumoral cues remain sensitive to noise and ambiguity, especially in medical imaging with low signal-to-noise ratios and inter-patient variability [13]. Diffusion-based classifiers have therefore been explored to improve robustness to subtle and noisy patterns [14], [15]. Diffusion models probabilistically model data distributions [16], [17] and can be adapted for classification by comparing class-conditional denoising or reconstruction errors [15]. To combine denoising-based classification with longrange volumetric MRI modeling, we adopt a transformer-based denoising backbone [18]–[20]. Despite these advantages, applying diffusion transformers to 3D MRI introduces substantial computational overhead due to the large number of tokens and repeated denoising steps, posing practical challenges for hardware-efficient deployment in clinical settings. While dynamic token selection has been explored for efficient vision transformers [21], identifying relevant tokens in volumetric medical imaging requires awareness of local anatomical context around tumor boundaries. To address this, we propose Diffusion Transformer with Routing for Classification (DiT-RC), an efficient diffusion transformer classifier with adaptive routing across attention heads, spatial tokens, and MLP width for volumetric MRI. The proposed design dynamically allocates computation based on input complexity, enabling more effective identification of clinically relevant regions around tumor boundaries. Our contributions are summarized as follows: • We propose DiT-RC, a diffusion transformer-based framework for PNI prediction from 3D MRI, leveraging denoising-based classification. • We introduce dynamic routing across attention heads, spatial tokens, and MLP width to enable adaptive computation. • We incorporate a lightweight convolutional module for improved token importance estimation around tumor boundaries.
II. P ROPOSED M ETHOD A. Diffusion Transformer for Classification We employ a diffusion-based transformer [22] classifier to predict perineural invasion (PNI) from tumor-centered 3D MRI volumes. Given a clean input volume x0 , a forward diffusion process generates a noisy sample xt at timestep t: √ √ xt = ᾱt x0 + 1 − ᾱt ϵ (1) where ϵ ∼ N (0, I) and ᾱt denotes the cumulative noise schedule. The noisy volume xt is partitioned into non-overlapping 3D patches and embedded into a token sequence Xt , which is processed by a transformer backbone conditioned on both the timestep and a class hypothesis. Specifically, the model estimates the noise ϵ̂θ (xt , t, y)
(2)
where y ∈ {0, 1} denotes the class hypothesis for PNI status. During inference, the same noisy input is evaluated under each class condition, and the prediction is determined by comparing the corresponding noise reconstruction errors my = ∥ϵ̂θ (xt , t, y) − ϵ∥22 , with ŷ = arg miny my . Compared with a conventional discriminative classifier, this formulation leverages noise-perturbed inputs across multiple timesteps and encourages the model to learn representations that remain informative under stochastic corruption. B. Adaptive Computation To improve efficiency, we introduce adaptive routing within each transformer block, controlling attention-head activation, spatial token selection, and MLP width for the token sequence X ∈ RN ×C . The routing condition PN is defined as r = [temb ; xglobal ], where xglobal = N1 i=1 Xi . It is used for attention-head and MLP-width routing, while token routing uses token features with local 3D context. AdaLN [23], [24] modulates each block with c = temb + yemb . 1) Attention Head Routing: Multi-head self-attention models diverse feature interactions through multiple attention heads. Given routing condition r, a lightweight router predicts head-wise scores shead ∈ RH from the routing condition r. These scores are converted into a binary routing mask mhead ∈ {0, 1}H using a Gumbel–Sigmoid estimator [25]. The routed attention output is defined as MHSA(X) =
H X
mhead · Attnh (X) h
(3)
h=1
where H denotes the number of attention heads and Attnh (·) is the attention computation over head h. 2) Spatial Token Routing: Volumetric MRI inputs yield a large number of spatial tokens, many of which contain limited information for PNI prediction; we therefore introduce spatial token routing to focus computation on informative regions using local spatial context. Specifically, token scores stok ∈ RN are computed by combining each token’s own
features with local spatial context extracted via a lightweight 3D convolution (kernel size 3×3×3) applied over the spatial token grid, which are then converted into a binary mask mtok ∈ {0, 1}N . The routed token representation is X ′ = mtok ⊙ X
(4)
where ⊙ denotes token-wise masking. For notational simplicity, token routing is written as masking, while the actual implementation performs sparse token selection before the MLP block so that only selected tokens participate in the feedforward computation. 3) MLP Width Routing: The transformer MLP expands token features into a high-dimensional space, making it a major source of computational cost. Given the routing condition r, a lightweight router predicts width scores smlp ∈ RG , which are converted into a binary width mask mmlp ∈ {0, 1}G . Let Z = ϕ(W1 X ′ ) denote the hidden activation of the MLP, where W1 is the first projection matrix and ϕ(·) is the nonlinear activation function. Width routing applies channel-group masking to the hidden representation, Z ′ = mmlp ⊙ Z
(5)
and the final MLP output is computed from Z ′ . Since the MLP is evaluated only for tokens selected by spatial token routing, token routing and width routing jointly reduce the feed-forward computation. These routing mechanisms focus computation on clinically relevant regions while reducing overall cost by dynamically adapting the computation pathway to the noisy input and diffusion timestep. C. Training Objective The proposed model is trained with a joint objective that combines diffusion-based classification with a routing budget constraint: Ltotal = Lcls + λbudget Lbudget (6) where Lcls denotes the diffusion classification loss and Lbudget regulates the adaptive routing behavior. The budget loss is only applied when routing is enabled, as it requires controllable computation through routing decisions. For a noisy input xt , we compute reconstruction errors my for y ∈ {0, 1} as defined in Sec. II-B and convert them into classification logits l = τ ·[−m0 , −m1 ], where τ is a learnable temperature parameter. To control computational cost, the effective computation ratio reff is defined as: reff = wattn kattn + wmlp (ktok kmlp )
(7)
where reff approximates the relative computational cost in terms of FLOPs, providing a hardware-agnostic proxy for efficiency [21], [26], and kattn , ktok , kmlp denote the average keep ratios of attention heads, spatial tokens, and MLP hidden channels, respectively.
AdaLN
ROI Extraction Noise Addition
x
C
Router Condition
Head Router 1 0 1 ... 0
V
... Σ
O
t
Timestep Embedding
G
Condition c Token Router Local Conv
y
0 1 0 ... 1
Class Embedding
C G
activated deactivated Concat Gating
Token Mask
AdaLN Token Selection G ...
Width Router 1 1 0 ... 0
Width Mask
Linear
SiLU
Linear
Σ
+
MLP
Gated MLP
Encoding
×N blocks
+ +
Noisy Volume
K
Q
Head Mask
Gated MHSA
3D Tumor MRI
ε0 ε1
PNI Prediction
Fig. 1: Overview of DiT-RC. The architecture consists of a stack of N transformer blocks (N = 6 in our implementation). TABLE I: Classification performance (AUC) under varying noise levels for PNI prediction.
The classification and budget losses are then given by: Lcls = L(y) mse + αLCE ,
Lbudget = (reff − λtarget )2
(8)
The coefficient λtarget is gradually increased during early training to stabilize optimization. This objective jointly learns diffusion classification and dynamic routing policies, enabling the model to balance predictive performance and computational efficiency. III. E XPERIMENT R ESULTS A. Dataset and Implementation Details The dataset consists of 155 patients (61 PNI-positive and 94 PNI-negative) collected over a 10-year period at Samsung Medical Center. T2-weighted MRI scans were used for all experiments. Each patient volume was cropped to a tumorcentered region of size 96 × 96 × 48 and divided into nonoverlapping 3D patches of size 6×6×6, forming a 16×16×8 token grid. Tumor masks were manually delineated using 3D Slicer and used only for tumor-centered ROI extraction, not as model inputs. We performed 5-fold stratified cross-validation at the patient level to avoid data leakage across scans from the same patient. The model was trained using AdamW with a learning rate of 8 × 10−5 , weight decay of 0.05, and a batch size of 4. All experiments were conducted on a single NVIDIA RTX 3090 GPU using mixed-precision training. During training, diffusion timesteps were randomly sampled from a 1000-step schedule. During inference, predictions were obtained by averaging logits over multiple timesteps (t ∈ {200, 350, 500, 650, 800}) and random seeds, following standard diffusion-based classification protocols to reduce stochastic variance. The 95% confidence interval was estimated by bootstrapping patient-level predictions. FLOPs and latency were reported per patient prediction under the fivetimestep and two-class protocol, excluding repeated noise ensembling. For routed models, FLOPs were computed using the activated heads, selected tokens, and MLP channel groups. For
Model ResNet-18 [27] DenseNet-121 [28] EfficientNet [29] Vision Transformer [10] Swin Transformer [30] Diffusion Classifier [17] DiT-RC
Clean 0.675 0.687 0.680 0.700 0.710 0.692 0.731
NL 0.630 0.643 0.645 0.654 0.700 0.692 0.730
NM 0.573 0.573 0.612 0.633 0.649 0.680 0.723
NH 0.475 0.527 0.500 0.610 0.615 0.672 0.710
robustness evaluation, additive Gaussian noise with standard deviations σ ∈ {0.1, 0.3, 0.5} was applied after per-volume min–max normalization, using pre-generated corrupted test sets shared across all models. B. Quantitative and Qualitative Results Table I summarizes the AUC performance under clean and noisy inputs. CNN-based models show relatively lower performance, while transformer-based models improve clean AUC but degrade under stronger noise. Diffusion-based models are more robust to noise, and DiT-RC achieves the best AUC across all settings. DiT-RC achieved an AUC of 0.731 on clean inputs (bootstrap 95% CI: 0.66–0.80) and showed five-fold standard deviations of 0.036/0.038/0.041/0.046 for Clean/NL/NM/NH. Table II compares our method with the same diffusion classifier implemented using a U-Net backbone based on MONAI and DDPM [17], which represents a standard CNNbased diffusion approach. While the U-Net-based diffusion classifier achieves lower computational cost in terms of FLOPs and latency, our method attains a higher AUC under the same per-prediction evaluation setting. This suggests that transformer-based denoising better captures PNI-related volumetric context, while introducing additional computational cost compared with the CNN-based diffusion baseline. Figure 2 shows Grad-CAM [31] visualizations across axial, coronal, and sagittal views. The red regions indicate areas associated with PNI-positive predictions. The proposed model
TABLE II: Comparison with Diffusion Classifier. Model Diffusion Classifier [17] DiT-RC
FLOPs (G) 151.75 257.57
Lat. (ms) 127.10 140.47
AUC 0.692 0.731
TABLE III: Efficiency comparison between DiT-RC and its non-routing variant (DiT-C). Model DiT-C DiT-RC
FLOPs (G) 418.79 257.57
Lat. (ms) 319.32 140.47
AUC 0.733 0.731
TABLE IV: Ablation study on routing components. Ablation FLOPs (G) — Module Contribution — w/o Head routing 338.28 w/o Width routing 294.58 w/o Token selection 273.30 — Token Selection Design — w/o Local conv 245.07 DiT-RC 257.57
DiT-C
DiT-RC
Axial
Coronal
Fig. 2: Grad-CAM visualization for PNI prediction.
(A) Head/Width routing
AUC
222.62 176.23 155.26
0.725 0.719 0.698
129.14 140.47
0.700 0.731
TABLE V: Effect of the budget loss coefficient λtarget .
Sagittal
ROI slice
Lat. (ms)
Compute map
(B) Token routing
Fig. 3: Routing analysis. (A) Average head and width routing ratios across timesteps. (B) Token-wise computation map. produces activation patterns similar to those of its non-routing variant (DiT-C), while reducing computational cost. C. Efficiency Analysis Transformer-based diffusion classifiers are computationally expensive compared to CNN-based approaches. To address this, we introduce a routing mechanism that dynamically reduces redundant computation. Table III shows that routing significantly reduces FLOPs from 418.79G to 257.57G and latency from 319.32 ms to 140.47 ms, while maintaining comparable performance. As shown in Figure 3, the attention head utilization decreases from 0.72 to 0.27, and the MLP width utilization from 0.69 to 0.27 as the timestep increases, indicating reduced computational demand at higher timesteps. Furthermore, tokenwise computation is concentrated on specific regions, showing that the model selectively allocates computation to informative areas. D. Ablation Study Table IV shows that head routing has the largest impact on performance, while token selection is essential for spatially adaptive computation. Removing local convolution slightly degrades performance, indicating the importance of local context.
λtarget FLOPs (G) 0.3 114.51 0.5(Ours) 257.57 0.7 341.23 1.0 418.80
Lat. (ms) AUC 95.71 0.650 140.47 0.731 225.91 0.733 321.23 0.740
Overall, combining global and local routing yields the best performance–efficiency balance. Table V shows that the routing budget controls the trade-off between performance and efficiency. Smaller budgets reduce computation but degrade performance, while larger budgets improve performance at higher cost. Our setting achieves a balanced trade-off. Notably, λtarget = 1.0 achieves higher AUC than DiT-C at the same FLOPs. As routing is effectively inactive in this setting, both models share the same inference structure. The improvement is attributed to the regularization effect of the budget loss. IV. C ONCLUSION This work presented an efficient diffusion transformer classifier for preoperative PNI prediction from volumetric MRI. By introducing adaptive routing across attention heads, spatial tokens, and MLP width, along with a local context-aware convolution module, the proposed method dynamically allocates computation according to input complexity. Experimental results demonstrate that the model achieves 0.731 AUC at 257.57 GFLOPs, while maintaining stable predictive performance. These findings demonstrate the potential of diffusion transformer-based classification with adaptive computation for efficient and robust medical image analysis, although external multi-center validation remains necessary to assess generalization across institutions, scanners, and imaging protocols. ACKNOWLEDGMENT This work was supported by the IITP grant (IITP-2023-RS2023-00256081) funded by MSIT, Korea, and the ANCHOR program (2026-ANCHOR-01-110) funded by the Ministry of Education and the Seoul Metropolitan Government, Republic of Korea.
R EFERENCES [1] T. Wei, X.-F. Zhang, J. He, I. Popescu, H. P. Marques, L. Aldrighetti, S. K. Maithel, C. Pulitano, T. W. Bauer, F. Shen, et al., “Prognostic impact of perineural invasion in intrahepatic cholangiocarcinoma: multicentre study,” British Journal of Surgery, vol. 109, no. 7, pp. 610–616, 2022. [2] Y. Han, H. Go, K. Kim, I. Um, J. Kim, J. Jung, N.-J. Kim, W. K. Jeong, W. J. Lee, K. Y.-K. Liao, et al., “Losa-net: A localized and scale-adaptive network for boundary-sensitive prediction of perineural invasion in 3d mri,” in 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI), pp. 1–5, IEEE, 2026. [3] Y. Han, I. Um, K. Kim, J. Kim, H. Go, J. Jung, N.-J. Kim, W. K. Jeong, W. J. Lee, P. Hong, et al., “Mma-former: Multi-window mixtureof-head attention transformer for adaptive pni prediction in 3d mri,” in 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI), pp. 1–5, IEEE, 2026. [4] Y. Han, M. Cha, K. Kim, I. Um, M. Sho, J. Y. Bae, J. Jung, J. H. Park, S. Lee, N.-J. Kim, et al., “Neonet: An end-to-end 3d mri-based deep learning framework for non-invasive prediction of perineural invasion via generation-driven classification,” arXiv preprint arXiv:2603.29449, 2026. [5] Z. Liu, C. Luo, X. Chen, Y. Feng, J. Feng, R. Zhang, F. Ouyang, X. Li, Z. Tan, L. Deng, et al., “Noninvasive prediction of perineural invasion in intrahepatic cholangiocarcinoma by clinicoradiological features and computed tomography radiomics based on interpretable machine learning: a multicenter cohort study,” International Journal of Surgery, vol. 110, no. 2, pp. 1039–1051, 2024. [6] C. Liebig, G. Ayala, J. A. Wilks, D. H. Berger, and D. Albo, “Perineural invasion in cancer: a review of the literature,” Cancer: Interdisciplinary International Journal of the American Cancer Society, vol. 115, no. 15, pp. 3379–3391, 2009. [7] S. Conti, N. S. Tissera, F. Castet, M. Basagaña-Farrés, M. T. Salcedo, E. Pando, C. Dopazo, L. Caritá, A. Turpin, V. N. Garcés, et al., “Perineural invasion is a prognostic factor in cholangiocarcinoma, regardless of anatomical location: a systematic review and meta-analysis: Perineural invasion in cholangiocarcinoma prognosis,” JHEP Reports, p. 101770, 2026. [8] C.-G. Li, Z.-P. Zhou, X.-L. Tan, and Z.-M. Zhao, “Perineural invasion of hilar cholangiocarcinoma in chinese population: One center’s experience,” World Journal of Gastrointestinal Oncology, vol. 12, no. 4, p. 457, 2020. [9] K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014. [10] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020. [11] A. Hatamizadeh, Y. Tang, V. Nath, D. Yang, A. Myronenko, B. Landman, H. R. Roth, and D. Xu, “Unetr: Transformers for 3d medical image segmentation,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 574–584, 2022. [12] J. Chen, Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, L. Lu, A. L. Yuille, and Y. Zhou, “Transunet: Transformers make strong encoders for medical image segmentation,” arXiv preprint arXiv:2102.04306, 2021. [13] F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier-Hein, “nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,” Nature methods, vol. 18, no. 2, pp. 203–211, 2021. [14] H. Chen, Y. Dong, Z. Wang, X. Yang, C. Duan, H. Su, and J. Zhu, “Robust classification via a single diffusion model,” arXiv preprint arXiv:2305.15241, 2023. [15] K. Clark and P. Jaini, “Text-to-image diffusion models are zero shot classifiers,” Advances in Neural Information Processing Systems, vol. 36, pp. 58921–58937, 2023. [16] J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in International Conference on Learning Representations (ICLR), 2021. [17] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems, vol. 33, pp. 6840– 6851, 2020. [18] W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF international conference on computer vision, pp. 4195–4205, 2023. [19] Y. Han, K. Kim, S. Ju, Y. Jean, M. Cha, S. Park, H. Jung, N.-J. Kim, W. K. Jeong, K. Y.-K. Liao, et al., “Foscu: Feasibility of synthetic mri
generation via duo-diffusion models for enhancement of 3d u-nets in hepatic segmentation,” in 2025 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS), pp. 1–5, IEEE, 2025. [20] K. Kim, J. Bae, Y. Han, J. Y. Bae, S. Ju, J. Lim, G. Kim, N.-J. Kim, W. K. Jeong, K. Y.-K. Liao, et al., “3d-lldm: Label-guided 3d latent diffusion model for improving high-resolution synthetic mr imaging in hepatic structure segmentation,” in 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI), pp. 1–5, IEEE, 2026. [21] Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “Dynamicvit: Efficient vision transformers with dynamic token sparsification,” Advances in neural information processing systems, vol. 34, pp. 13937– 13949, 2021. [22] J. Li, Y. Xu, T. Lv, L. Cui, C. Zhang, and F. Wei, “Dit: Self-supervised pre-training for document image transformer,” in Proceedings of the 30th ACM international conference on multimedia, pp. 3530–3539, 2022. [23] Y. Guo, C. Wang, S. X. Yu, F. McKenna, and K. H. Law, “Adaln: A vision transformer for multidomain learning and predisaster building information extraction from images,” Journal of Computing in Civil Engineering, vol. 36, no. 5, p. 04022024, 2022. [24] E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, 2018. [25] E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” arXiv preprint arXiv:1611.01144, 2016. [26] H. Cai, C. Gan, T. Wang, Z. Zhang, and S. Han, “Once-for-all: Train one network and specialize it for efficient deployment,” arXiv preprint arXiv:1908.09791, 2019. [27] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016. [28] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700–4708, 2017. [29] M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International conference on machine learning, pp. 6105–6114, PMLR, 2019. [30] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision, pp. 10012–10022, 2021. [31] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision, pp. 618–626, 2017.