Conceptio › Archive › arXiv CS
arXiv CSopen access

Evidential Reasoning Advances Interpretable Real-World Disease Screening

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Evidential Reasoning Advances Interpretable Real-World Disease Screening

Chenyu Lian 1 2 Hong-Yu Zhou 3 Jing Qin 1 2

arXiv:2605.15171v1 [cs.CV] 14 May 2026

Abstract

(a) Deviation-based prediction

Disease screening is critical for early detection and timely intervention in clinical practice. However, most current screening models for medical images suffer from limited interpretability and suboptimal performance. They often lack effective mechanisms to reference historical cases or provide transparent reasoning pathways. To address these challenges, we introduce EviScreen, an evidential reasoning framework for disease screening that leverages region-level evidence from historical cases. The proposed EviScreen offers retrospection interpretability through regional evidence retrieved from dual knowledge banks. Using this evidential mechanism, the subsequent evidence-aware reasoning module makes predictions using both the current case and evidence from historical cases, thereby enhancing disease screening performance. Furthermore, rather than relying on post-hoc saliency maps, EviScreen enhances localization interpretability by leveraging abnormality maps derived from contrastive retrieval. Our method achieves superior performance on our carefully established benchmarks for real-world disease screening, yielding notably higher specificity at clinical-level recall. Code is publicly available at https://github.com/DopamineLcy/ EviScreen.

Interpretability Prediction

Features from Deviation from normal samples normalcy

(b) Direct prediction

Deviation map

Interpretability Prediction

Neural network

(c) Evidential reasoning (ours)

Relying on post-hoc saliency map

Interpretability (1) Retrospection interpretability

Referencing Current case Retrieval Dual knowledge banks Evidence from historical cases Evidence

Evidential reasoning

Prediction

(2) Localization interpretability

Dual distance maps Contrast

Abnormality map

Figure 1. Comparison of three paradigms for disease screening. (a) Deviation-based prediction methods only generate deviation maps to provide localization interpretability. (b) Direct prediction methods rely on post-hoc saliency maps, such as Grad-CAM, to achieve localization interpretability. (c) We propose evidential reasoning, which retrieves regional evidence from historical cases. Our EviScreen not only provides retrospection interpretability, mirroring the decision-making process of clinicians, but also produces better localization interpretability than deviation-based approaches through more focused abnormality maps.

1. Introduction Medical imaging serves as a crucial tool for disease screening, as it provides visual clues that help clinicians locate

potential anomalies (Adams et al., 2023; Zhang et al., 2024; Aggarwal et al., 2021). In clinical practice, specialists usually rely on a combination of their professional experience and evidence from historical cases to make clinical judgments (Fanaroff et al., 2019; Sackett et al., 1996). However, mainstream disease-screening pipelines lack this critical ability to trace back to historical cases, reducing their interpretability, trustworthiness, and performance. Although saliency map-based methods such as Grad-CAM (Selvaraju et al., 2017) can offer some localization interpretability,

1

The Center for Smart Health, School of Nursing, the Hong Kong Polytechnic University, Hong Kong, China 2 Research Institute for Smart Ageing, the Hong Kong Polytechnic University, Hong Kong, China 3 School of Biomedical Engineering, Tsinghua Medicine, Tsinghua University, Beijing, China. Correspondence to: Hong-Yu Zhou, Jing Qin <[email protected], [email protected]>. Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).

1

Evidential Reasoning Advances Interpretable Real-World Disease Screening

Lack of a clinically oriented evaluation framework. While disease screening is often framed as an anomaly detection problem, existing evaluation protocols fall short of real-world clinical needs. Firstly, standard performance metrics, such as the area under the ROC curve (AUROC), are not aligned with clinical requirements. Secondly, current benchmarks often fail to assess model generalizability, as they typically lack real-world testing on external test sets.

their quality is still considered unsatisfactory (Saporta et al., 2022). More importantly, these approaches fail to provide an evidence-based reasoning process for their predictions, leading to a lack of retrospection interpretability and potentially constraining their performance. There are two major paradigms for disease screening. The first paradigm, deviation-based prediction (Figure 1a), aims to solve the problem by one-class classification (also known as unsupervised anomaly detection), which learns a mode of normalcy using normal cases (Roth et al., 2022; Li et al., 2025; Guo et al., 2023; Lian et al., 2025). While this allows for the generation of deviation maps by identifying deviations of the current case from normalcy, a key limitation is that it cannot fully utilize information from positive cases. This omission can impair performance, particularly with complex modalities such as chest X-rays and dermoscopic images. Another intuitive approach is direct prediction (Figure 1b) by fully supervised classification, which frames the task as a binary problem, training on both normal and pathological images (Zhang et al., 2023; Cai et al., 2024b; Adams et al., 2023). The interpretability of these models typically relies on post-hoc saliency maps that visualize regions deemed important for the prediction (Selvaraju et al., 2017; Marjanovic et al., 2024).

Limited evidence awareness in mainstream screening pipelines. The interpretability of many models relies on post-hoc saliency maps. While these maps highlight regions deemed important, they lack case-based evidence explaining why a region appears pathological, limiting their alignment with expert reasoning. Prototype-based interpretable methods partially address this issue by comparing image regions with learned prototypes or historical examples. However, their representational capacity is bounded by a fixed number of prototypes specified before training, which may be insufficient to cover the diverse pathological appearances encountered in real-world screening. Insufficient mining of granular evidence from pathological cases. Deviation-based anomaly detection methods can effectively identify what deviates from normal, but cannot leverage the rich, granular information contained within pathological cases. The core challenge lies in the fact that a pathological image comprises both normal and pathological regions, which makes it difficult to directly model the “pathological” class at the patch level.

Prototype-based interpretable methods have shown the value of visual evidence from historical cases (Kim et al., 2021; Wang et al., 2025). Nevertheless, their applicability to clinically oriented disease screening remains limited. Existing methods are typically designed for diagnosis or multi-label recognition, where predictions are explained by a fixed set of learned prototypes associated with predefined disease classes. In contrast, clinically oriented disease screening requires maintaining scalable evidence repositories of both normal and pathological historical cases, retrieving queryspecific region-level evidence, and integrating such evidence into the prediction process. How to achieve these capabilities remains underexplored.

1.2. Contributions This paper addresses the aforementioned limitations by making the following key contributions: 1. We propose a new, clinically relevant evaluation framework for real-world disease screening. It uses ten public datasets across three critical medical modalities, prioritizing the proposed clinically oriented metrics and tests on external datasets. This setting mirrors real-world clinical scenarios, providing convincing benchmarks. 2. We introduce EviScreen, an evidential reasoning framework powered by the dual knowledge banks to facilitate real-world disease screening. The framework enhances prediction transparency through a synergy of retrospection interpretability and localization interpretability. 3. The proposed dual knowledge banks provide a scalable alternative to fixed prototype-based representations, capturing a broader spectrum of regional features from both normal and pathological cases. This enables our model to learn a rich representation of both normal regions and diverse pathological patterns, facilitating a more precise evidence-aware reasoning process. 4. Comprehensive experiments validate that the proposed EviScreen outperforms different types of comparative

To address these gaps, we introduce EviScreen, an evidential reasoning framework for disease screening with medical images (Figure 1c). Our framework provides retrospection interpretability with similar image patches from dual knowledge banks of normal and pathological historical cases, making its reasoning process more transparent and reliable. Unlike previous methods that rely on saliency maps or deviation maps, EviScreen advances localization interpretability with abnormality maps generated by contrastive retrieval from dual knowledge banks. 1.1. Current Limitations We identify three primary limitations in developing interpretable real-world disease screening methods:

2

Evidential Reasoning Advances Interpretable Real-World Disease Screening

� Dual knowledge bank construction

Step 1: Evidence retrieval Step 2: Evidence-aware reasoning

Normal historical cases Pathological historical cases

Cross-attention Z0N = Z0P = Z

Z

Each regional feature

XNB

XPB

Fθ

Fθ

Regional feature extraction

SN

EN (i, j)

ZN

Z(i, j)

ZN (i, j)

ZP

EP (i, j)

ZP (i, j)

EP (i, j)

key, value

query

Z+1 N

SN

Inter-patch refinement

TP (i, j)

Z+1 P

SP

×L

Cross-attention

evidential reasoning (main method)

×N

TN (i, j)

SA

SP

×N

EN (i, j)

Regional Each regional Evidence awareness Feature patch features feature

KP

Evidence vectors

SA: Self-attention

key, value

SA

Evidence retrieval

KN

query

Evidential reasoning

MLP

ZN

Prediction

ZP

KN

KP

Feature maps

� Disease screening by

Dual knowledge banks

Input

x

Fθ

Z Foundation model Regional features

Pooling

contrastive retrieval (training-free variant) Contrastive retrieval

Prediction

M

Abnormality map

Normal & pathological regional feature

KN

Feauture maps of normal & pathological branch Normal & pathological knowledge bank

Each regional feature Z

EN (i, j) MN (i, j)

MN

Reference retrieval k nearest neighbors Distance patch Distance map with distances

Z(i, j)

KP

EP (i, j) MP (i, j)

Contrast M Abnormality map

MP

Figure 2. Overview of our framework that consists of two main stages: (1) Dual knowledge bank construction, where patch-level features from historical normal and pathological cases are extracted by a foundation model to construct two distinct knowledge banks, KN and KP . (2) Disease screening by evidential reasoning (main method) or contrastive retrieval (training-free variant). • Evidential reasoning includes two steps, the first step is evidence retrieval, where regional features from the current image input x are extracted to serve as queries. These queries retrieve the k-nearest neighbors from both knowledge banks, which serve as the evidence for subsequent reasoning (EN , EP ). During the second step of evidence-aware reasoning, evidence awareness is realized by cross-attention between the current case and retrieved evidence, followed by inter-patch refinement via self-attention. • During the training-free variant with contrastive retrieval, distance maps (MN , MP ) are generated based on the distance to k-nearest neighbors retrieved from the dual knowledge banks. The final abnormality map M is obtained by contrasting the dual distance maps.

methods for real-world disease screening, particularly with respect to the clinically oriented metrics.

fixed-prototype design can limit capacity when real-world screening involves highly diverse appearances.

2. Related work

2.2. Disease Detection by Medical Anomaly Detection

2.1. Interpretability in AI for Medical Imaging

Medical anomaly detection aims to distinguish abnormal cases from normal ones based on their deviation from normality (Cai et al., 2025; Li et al., 2025; Guo et al., 2023; Fernando et al., 2021). BMAD established benchmarks for medical anomaly detection, but it was not designed for clinically oriented disease screening (Bao et al., 2024). BenchReAD focused on retinal imaging to enhance the systematicity for medical anomaly detection (Lian et al., 2025). In this paper, we regard medical anomaly detection as one category of deviation-based prediction methods.

Interpretability is critical for building trust in medical imaging applications. Conventional models that rely on post-hoc saliency maps (Selvaraju et al., 2017; Chattopadhay et al., 2018) to highlight regions of interest are often considered to provide interpretations of unsatisfactory quality (Saporta et al., 2022). Several newer methods move beyond mere feature attribution by employing techniques such as counterfactual intervention (Pan et al., 2025), graph networks (Hu et al., 2024), and natural language descriptions (Cai et al., 2024a). Nevertheless, these methods still do not provide a rationale by referring to historical cases in the way human experts do. While prototype-based interpretable models (Kim et al., 2021; Wang et al., 2025) improve transparency by associating predictions with representative visual prototypes, their

2.3. Coreset-Based Memory Bank Coresets have been widely used for k-NN and k-Means approaches (Har-Peled & Kushal, 2005), representation learning (Roth et al., 2020), and deviation-based anomaly detec-

3

Evidential Reasoning Advances Interpretable Real-World Disease Screening

where Gagg represents a locally aware patch-feature aggregation function (Roth et al., 2022). In addition, to reduce redundancy and enhance efficiency, we apply greedy coreset subsampling (Agarwal et al., 2005; Roth et al., 2022) to SN and SP , producing compact knowledge banks KN and KP . The optimization objective is to find a subset that best represents the entire feature set:

tion methods (Roth et al., 2022; Jiang et al., 2022). Typically, a coreset-based memory bank is constructed by extracting features from intermediate layers of ImageNet-pretrained models (Deng et al., 2009) and subsampling to reduce redundancy. In this paper, we construct dual knowledge banks using coreset-based memory banks to store regional features from both normal and pathological cases.

∗ KN = arg min

2.4. Vision Foundation Models for Medical Imaging Vision foundation models, including general ones (Oquab et al., 2024; He et al., 2022) and those tailored for medical images (Zhou et al., 2023b;a; Yan et al., 2025; Yang et al., 2025), have shown promising performance on medical imaging (Zhang & Metaxas, 2024; Moor et al., 2023; Chen et al., 2022; Zhang et al., 2024). However, the adoption of foundation models for disease screening remains understudied. In this work, we adopt vision foundation models to extract regional features for dual knowledge bank construction.

Given an input image x ∈ RH×W ×C , we first extract its intermediate regional features using the same frozen foundation model Fθ and aggregation function Gagg : Z = Gagg (Fθ (x)) ,

[ xi ∈XPB

(3)

where Z ∈ Rh×w×d is the feature map with spatial dimensions h × w and feature dimension d. Each regional feature vector Z (i, j) ∈ Rd in Z serves as a query. 3.2.1. E VIDENCE R ETRIEVAL For each regional feature vector Z (i, j) (query), we retrieve the k-nearest neighbors from both the normal knowledge bank KN and the pathological knowledge bank KP . This process yields two sets of evidence vectors EN (i, j) and EP (i, j) for each spatial location (i, j):

Let XN and XP denote the training sets of normal and pathological cases, respectively, where ∀x ∈ XN : yx = 0 and ∀x ∈ XP : yx = 1. We partition each set into two disjoint subsets: one for constructing the dual knowledge banks (XNB , XPB ) and the other for training the evidential reasoning module (XNR , XPR ), i.e., XN = XNB ∪ XNR and XP = XPB ∪ XPR . Using a frozen foundation model Fθ , we extract intermediate regional features from all images in XNB and XPB , yielding two sets of regional features:

SP =

(2b)

3.2. Disease Screening by Evidential Reasoning

3.1. Dual Knowledge Bank Construction

B xi ∈XN

KP∗ = arg min max min ∥m − n∥2 ,

where ∥ · ∥2 denotes the Euclidean distance. Since this problem is NP-hard, an iterative greedy approximation is employed (Sener & Savarese, 2018; Wolsey & Nemhauser, 2014; Roth et al., 2022). The resulting dual knowledge banks provide the foundational evidence for the subsequent reasoning stage. In addition to their role as evidence providers in the reasoning module, the banks themselves enhance training-free disease screening through the proposed contrastive retrieval, which we discuss in Section 3.3.

Overview. As illustrated in Figure 2, our proposed framework comprises two primary stages: dual knowledge bank construction (Section 3.1) and evidential reasoning (Section 3.2). In the first stage, we extract intermediate regional features from historical normal and pathological cases using a pretrained foundation model. These features are subsequently subsampled to form compact knowledge banks that represent normal and pathological patterns. In the second stage, evidence is retrieved from the dual knowledge banks to enable interpretable disease screening (Section 3.2.1). Evidence-aware reasoning is performed by first using crossattention for evidence awareness, followed by self-attention for inter-patch refinement (Section 3.2.2). Additionally, our framework enhances training-free disease screening through the proposed contrastive retrieval (Section 3.3).

[

(2a)

KP ⊂SP m∈SP n∈KP

3. Method

SN =

max min ∥m − n∥2 ,

KN ⊂SN m∈SN n∈KN

EN (i, j) = NN(Z (i, j) , k; KN ), ∀i ∈ {1, · · · , h},

j ∈ {1, · · · , w}, (4a) EP (i, j) = NN(Z (i, j) , k; KP ), ∀i ∈ {1, · · · , h}, j ∈ {1, · · · , w}, (4b)

where the function NN (z, k; K) returns a set of k closest vectors in knowledge bank K along with the Euclidean distances to them. Therefore, EN , EP ∈ Rh×w×k×d .

Gagg (Fθ (xi )) ,

(1a)

3.2.2. E VIDENCE -AWARE R EASONING

Gagg (Fθ (xi )) ,

(1b)

The retrieved evidence above is then used to conduct evidence-aware reasoning and generate evidence-aware feature maps ZN and ZP . We denote the initial query as 4

Evidential Reasoning Advances Interpretable Real-World Disease Screening (a) Training data distribution

(b) Distribution of abnormality categories in test sets

46.2 % (13,264)

53.8 % (15,424)

53.1 % (97,243)

46.9 % (85,882) 38.6 % (7,990)

61.4 % (12,717)

Figure 3. (a) Training data distribution. (b) Distribution of abnormality categories in test sets (JSIEC, RIADD, CheXpert, and Derm12345).

Z0N (i, j) = Z (i, j). For each layer l = 0, . . . , L − 1, we first perform evidence-aware cross-attention. The evidence vectors from EN (i, j) and EP (i, j) are linearly projected to serve as keys and values. For the branch with the normal knowledge bank (similarly for the pathological branch): ! ⊤ ZℓN (i, j) (EN (i, j)) ℓ √ TN (i, j) = softmax EN (i, j) , d ∀i ∈ {1, · · · , h}, j ∈ {1, · · · , w}.

the distance maps produced by retrieving from the dual knowledge banks to generate the abnormality map. Contrastive retrieval. Following Equation Equation (3), we obtain the feature map Z ∈ Rh×w×d of the input image x ∈ RH×W ×C . For each regional feature vector Z(i, j), we calculate the average distance to the top k-nearest neighbors from both the normal knowledge bank KN and the pathological knowledge bank KP . This process yields two distance maps, where for each spatial location (i, j):

(5)

MN (i, j) = NN-Dis (Z(i, j), k; KN ) ,

To achieve inter-patch refinement, we reshape TℓN (i, j) ∈ Rh×w×d to SℓN (i, j) ∈ R(h·w)×d and apply self-attention to SℓN (i, j), generating the feature map of the next layer: ! ⊤ SℓN (SℓN ) l+1 √ ZN = softmax SℓN . (6) d

∀i ∈ {1, · · · , h}, j ∈ {1, · · · , w},

(8a)

∀i ∈ {1, · · · , h}, j ∈ {1, · · · , w},

(8b)

MP (i, j) = NN-Dis (Z(i, j), k; KP ) ,

where MN , MP ∈ Rh×w , and NN-Dis(z, k; K) outputs the average Euclidean distances to the nearest k vectors in the knowledge bank K. We generate the abnormality map by contrasting the two distance maps:

Finally, we obtain evidence-aware feature map ZN = ZL−1 N , and similarly, ZP = ZL−1 . The prediction is obtained from P the aggregated evidence-aware features through a multilayer perceptron (MLP):   CLS ŷ = MLP ZCLS , (7) N ; ZP

M(i, j) = ReLU (MN (i, j) − MP (i, j)) ,

∀i ∈ {1, · · · , h}, j ∈ {1, · · · , w}.

(9)

where ZCLS refers to the [CLS] token of Z.

The final prediction score is calculated by pooling the pointwise scores of the abnormality map M.

3.3. Training-Free Variant with Contrastive Retrieval

4. Experiments and Results

The proposed dual knowledge banks not only provide the evidence for evidential reasoning but can also enhance trainingfree disease screening. Leveraging the pathological knowledge bank is nontrivial since the regional features extracted from pathological cases are not “clean” (containing both normal and pathological image regions). To address this challenge, we propose contrastive retrieval, which contrasts

4.1. Evaluation Framework Construction To the best of our knowledge, there is no clinically oriented evaluation framework for real-world disease screening. We construct the evaluation framework to facilitate this and future research, focusing on clinically oriented metrics and external tests that simulate clinical scenarios. 5

Evidential Reasoning Advances Interpretable Real-World Disease Screening Table 1. Results for disease screening in four testing sets (%). The best results are in bold, and the second-best results are underlined. * A variant of PatchCore that uses the same foundation models as ours. Dataset

Metric AUROC AP Spe@95%R JSIEC Spe@99%R Spe@100%R CSR AUROC AP Spe@95%R RIADD Spe@99%R Spe@100%R CSR AUROC AP Spe@95%R CheXpert Spe@99%R Spe@100%R CSR AUROC AP Spe@95%R Derm12345 Spe@99%R Spe@100%R CSR

Ours 98.06 96.10 94.74 91.62 91.27 88.95 91.32 60.35 72.92 59.39 55.35 54.38 96.72 94.71 84.04 74.26 68.72 70.59 97.43 34.23 90.29 85.78 78.49 77.92

Ours-TF 96.76 94.20 91.48 87.74 87.33 85.07 90.42 59.96 69.96 55.26 51.25 50.73 92.30 82.22 66.60 60.00 55.74 48.91 97.21 37.23 88.27 80.27 70.35 69.94

FM 95.84 94.24 87.95 79.29 78.39 80.02 87.88 38.16 65.28 55.73 51.75 48.91 95.60 91.81 79.79 67.66 60.00 57.41 95.93 28.41 84.37 73.57 55.41 55.09

PatchCore* PatchCore NFM-DRA DRA 94.96 92.12 95.53 92.53 89.61 86.62 93.23 89.53 87.26 81.09 90.37 80.12 83.31 74.31 84.07 71.88 82.34 73.75 82.96 71.12 68.70 68.67 81.61 69.81 87.49 79.56 84.63 81.62 59.73 32.46 40.97 50.45 61.64 48.70 56.50 51.62 48.23 33.62 42.68 36.71 41.63 29.36 37.64 29.74 42.36 28.25 36.27 30.38 67.41 57.92 63.94 85.63 38.15 30.61 34.23 70.64 23.19 22.13 25.11 45.96 16.81 11.49 16.17 23.62 15.53 11.06 13.62 18.72 12.61 8.57 10.54 17.63 84.10 81.59 85.27 94.97 3.24 4.49 5.69 19.70 57.24 51.56 57.85 77.19 50.29 43.50 49.21 66.16 46.90 36.82 41.80 56.91 46.62 36.67 41.63 56.66

EDC 79.12 71.44 51.45 43.01 42.11 37.16 57.23 11.67 16.44 9.51 9.08 8.94 65.26 33.88 25.11 16.81 14.89 12.52 75.71 2.79 36.49 25.10 22.15 22.04

SimpleNet 73.73 57.66 53.81 49.31 49.10 38.27 60.21 13.44 24.58 15.86 14.26 13.94 55.24 29.90 11.49 9.15 8.94 7.58 75.33 4.22 38.10 25.95 20.94 20.84

CIPL 94.83 91.36 87.33 79.57 78.60 73.41 87.09 59.08 64.92 51.40 43.99 44.85 92.57 84.60 68.09 48.72 49.15 47.29 94.00 13.37 82.04 68.08 54.95 54.74

ficity achieved by any threshold in the set T≥X% :

4.1.1. B ENCHMARKS AND DATA Benchmarks are established across three medical domains (ophthalmology, radiology, and dermatology) using ten public datasets spanning three modalities. The datasets include color fundus photography (JSIEC (Cen et al., 2021), RIADD (Pachade et al., 2021), EDDFS (Xia et al., 2024), BRSET (Nakayama et al., 2024)); chest X-rays (CheXpert (Irvin et al., 2019), MIMIC-CXR (Johnson et al., 2019a)); and dermoscopic images (Derm12345 (Yilmaz et al., 2024), HAM10000 (Tschandl et al., 2018), BCN20000 (Hernández-Pérez et al., 2024), PAD-UFES20 (Pacheco et al., 2020)). Figure 3 illustrates the distribution of training data and abnormality categories in test sets. Refer to Section A of Appendix for more details.

Spe@X%R = max {Specificity(τ )} . τ ∈T≥X%

(11)

Clear Separation Rate. Another key goal for AI in disease screening is to reduce the manual review workload for clinicians to double-check ambiguous cases. To this end, we introduce a metric termed Clear Separation Rate (CSR), which quantifies the proportion of cases that fall outside the overlapping prediction score region between the two classes, i.e., all negative cases score below any positive case and vice versa. Let si be the predicted score for the i-th case and y ∈ {0, 1} be its true label, CSR is calculated as:

4.1.2. C LINICALLY O RIENTED M ETRICS P ROPOSED CSR =

Specificity at X% Recall. In clinical practice, disease screening demands high recall to minimize missed diagnoses. Concurrently, maximizing specificity is a key objective to reduce unnecessary follow-up tests. Based on this clinical need, we introduce the metric Specificity at X% Recall (Spe@X%R), formally calculated as follows:

1 N

N X

I(si < min sj )

i=1

+

N X i=1

j:yj =1

!

(12)

I(si > max sj ) , j:yj =0

where N is the sample size and I(·) is the indicator function.

First, we define the set of decision thresholds, T≥X% , to include all thresholds τ for which the recall is at least X%: T≥X% = {τ | Recall(τ ) ≥ X/100} .

SCRD4AD 94.88 89.85 88.50 78.74 78.46 70.72 83.83 39.94 57.92 47.73 43.98 42.64 51.50 27.74 12.55 5.53 5.53 4.61 74.30 3.80 45.73 38.03 32.86 32.69

In our experiments, we conduct a comprehensive evaluation using multiple metrics, including the average AUROC, Average Precision (AP), Spe@95%R, Spe@99%R, Spe@100%R, and CSR.

(10)

The value of Spe@X%R is defined as the maximum speci6

Evidential Reasoning Advances Interpretable Real-World Disease Screening 1 Query

2

3

4

Evidence from normal knowledge bank

5

6

7

Evidence from pathological knowledge bank

8

9

PatchCore*

Abnormality map (Ours)

JSIEC Large optic cup

Normal

Normal

Normal

Increased cup disc Increased cup disc Increased cup disc

CRVO

Normal

Normal

Normal

RVO

RVO

Vascular occlusion

Pleural Effusion

Normal

Normal

Normal

Pleural Effusion

Pleural Effusion

Pleural Effusion

JSIEC

CheX

CheX Atelectasis

Normal

Normal

Normal

Atelectasis

Atelectasis

Atelectasis

Derm Basal Cell Carcinoma (BCC)

Non-malignant

Non-malignant

Non-malignant

BCC

BCC

BCC

Derm Melanoma

Non-malignant

Non-malignant

Non-malignant

Melanoma

Melanoma

Melanoma

Figure 4. Visualization showing our method provides interpretable evidence when making predictions. Columns 2-7 depict the regional evidence for each representative query patch retrieved from historical cases, providing retrospection interpretation. In addition, the last column represents abnormality maps, providing localization interpretation. Specifically, the first column includes cases for testing, with representative query patches highlighted in blue squares. The 2-4 columns depict the retrieved evidence from the normal knowledge bank (patches in green squares), and the whole source images for the patches are shown for reference. Similarly, the 5-7 columns depict the retrieved evidence from the pathological knowledge bank (patches in red squares). The final two columns compare the localization interpretability, showing our method provides more focused abnormality maps than PatchCore while using the same foundation models.

4.2. Results on Real-World Disease Screening

able dual knowledge banks are better suited to real-world screening than fixed learned prototypes. To provide a more comprehensive evaluation, we also include a training-free variant of our approach (Ours-TF). Further discussions are provided in Section 4.3.

EviScreen outperforms various types of comparative approaches for real-world disease screening. We evaluate EviScreen on our proposed evaluation framework across three clinical domains, comparing against various types of state-of-the-art methods. The compared methods include direct prediction by fine-tuning foundation models (FM), deviation-based anomaly detection approaches (PatchCore (Roth et al., 2022), SCRD4AD (Li et al., 2025), EDC (Guo et al., 2023), and SimpleNet (Liu et al., 2023)), supervised variants (NFM-DRA (Lian et al., 2025), DRA (Ding et al., 2022)), and a prototype-based interpretable method (CIPL (Wang et al., 2025)). As shown in Table 1, EviScreen outperforms various approaches for real-world disease screening, consistently achieving the highest AUROC, AP, Spe@X%R, and CSR, across all datasets. Notably, the consistent gains of EviScreen over the recent prototype-based CIPL model suggest that scal-

EviScreen exhibits a notable improvement with regard to clinically oriented metrics. AUROC is widely used in binary classification and anomaly detection to evaluate the ability of models to distinguish between positive and negative cases. However, we observe that models with small differences in AUROC may exhibit significant differences in clinically oriented metrics, as shown in Table 1. In terms of the Spe@X%R metric, EviScreen outperforms other approaches by notable margins, even when improvements in AUROC are relatively modest. For instance, on the JSIEC dataset, EviScreen produces relative improvements over FM of 2.3% in AUROC, but by

7

Evidential Reasoning Advances Interpretable Real-World Disease Screening Table 2. Ablation analysis results (%) on components of evidential reasoning: evidence retrieval and evidence-aware reasoning. Evidence retrieval

Evidence-aware reasoning

AUROC

Spe@95R

Spe@99R

Spe@100R

CSR

AUROC

Spe@95R

Spe@99R

Spe@100R

CSR

✓ ✓

96.76 97.78 98.06

91.48 92.73 94.74

87.74 88.02 91.62

87.33 86.22 91.27

85.07 86.99 88.95

90.42 86.12 91.32

69.96 61.86 72.92

55.26 47.97 59.39

51.25 41.93 55.35

50.73 41.53 54.38

✓ ✓

JSIEC

RIADD

(a)

ing predictions. As shown in Figure 4, the interpretability of EviScreen is two-fold. Columns 2-7 depict retrospection interpretability by tracing historical cases, generating regional evidence for each representative query patch from both the normal and pathological knowledge bank. Column 9 illustrates the localization interpretability presented by abnormality maps. By comparing our abnormality maps and the deviation maps produced by another training-free method (PatchCore*), it can be observed that our method generates more focused maps, providing clearer interpretation. More visualizations are in Section C of Appendix. 4.3. Dual Knowledge Banks Enhance Training-Free Disease Screening via Contrastive Retrieval Beyond providing evidence for evidential reasoning, the dual knowledge banks have the potential to enable trainingfree disease screening. To achieve this, we propose contrastive retrieval to capture the discrepancies between the retrieved neighbors from the dual knowledge banks. The pipeline of the training-free variant of our method is illustrated in Figure 2 (refer to Section 3.3 for more details).

(b)

Performance compared to other training-free methods. Figure 5a highlights the advantage of our dual knowledge banks with contrastive retrieval over previous training-free methods. As illustrated by the radar charts, the foundation model-enhanced PatchCore (PatchCore*) outperforms the original version using ImageNet-pretrained feature extractors. Crucially, even when adopting the same foundation models, our method consistently outperforms PatchCore across all benchmarks and various evaluation metrics.

Figure 5. (a) Results (%) of ours (training-free variant) compared to other training-free methods. (b) Performance improves when the number of samples increases, especially for Spe@X%R.

Scaling the dual knowledge banks. As Figure 5b shows, the performance of our training-free methods with the dual knowledge banks improves as the number of samples increases, in particular for clinically oriented metrics.

7.7%, 15.6%, and 16.43% in Spe@95%R, Spe@99%R, and Spe@100%R, respectively. These advantages could meaningfully reduce clinical costs by minimizing the need for re-examinations. Similar trends are observed for the metric of CSR that measures the separation between predictions of negative and positive cases, where EviScreen surpasses the best runner-up (excluding Ours-TF) by 9.0%, 11.2%, 23.0%, and 37.5%, on the four datasets, respectively. These results show that EviScreen achieves clearer separation between negative and positive cases, indicating its potential to reduce the workload of clinicians in double-checking AI predictions.

4.4. Ablation Study Ablation analysis on evidential reasoning. The proposed evidential reasoning consists of two key components: evidence retrieval and evidence-aware reasoning. The former supplies query-specific visual evidence from dual knowledge banks of normal and pathological cases, whereas the latter enables the model to incorporate such evidence into the prediction process. As shown in Table 2, performance

EviScreen provides interpretable evidence when mak8

Evidential Reasoning Advances Interpretable Real-World Disease Screening

ter interpretability than current deviation-based prediction methods and direct prediction methods. Extensive experiments consistently validate the improvements our method brings, in particular for clinically oriented metrics. We hope that this work can inspire more clinically oriented, interpretable algorithms for real-world disease screening. Future work will address broader modalities, 3D medical images, and finer-grained screening tasks.

Figure 6. Hyperparameter analysis on the number of nearest neighbors k, ratio of subsampling, and normalization before calculating Euclidean distances. Default options are underlined.

Impact Statement The paper presents work that aims to advance the field of medical image analysis by introducing an interpretable framework for disease screening, mimicking clinical decision-making through evidential reasoning. By prioritizing high specificity at clinical-level recall and providing transparent visual evidence from historical cases, our method has the potential to significantly reduce unnecessary follow-up examinations and the associated psychological and economic burdens on patients. While the evidencebased mechanism enhances trust between AI and clinicians, potential ethical considerations, such as strict privacy safeguards when storing patient data in real-world clinical deployments, must be addressed.

decreases when any of the components is absent, showing the necessity of these mechanisms. Hyperparameter analysis on dual knowledge bank construction. We conduct hyperparameter analysis on the number of nearest neighbors k (Section 3.2.1), ratio of subsampling (Section 3.1), and the normalization before calculating Euclidean distance. Figure 6 illustrates the performance in the validation set of the ophthalmology benchmark, using the dual knowledge banks with contrastive retrieval. Hyperparameters selected in our default settings are underlined. 4.5. Implementation Details Our code is implemented using PyTorch 2.4.1 (Paszke et al., 2019). All experiments are carried out with Nvidia GeForce RTX 3090 GPUs. We employ state-of-the-art ViTbased (Dosovitskiy et al., 2020) foundation models for each modality: RETFound-Dinov2 (Zhou et al., 2025) for color fundus photography (CFP), CheXFound (Yang et al., 2025) for chest X-rays, and PanDerm (Yan et al., 2025) for dermoscopic images. Faiss 1.8.0 (Johnson et al., 2019b) serves as the engine for dual knowledge banks. For evidential reasoning, AdamW (Loshchilov & Hutter, 2017) is adopted as the default optimizer, with a weight decay of 0.05, β1 of 0.9, and β2 of 0.95. We employ a “warm-up” strategy by linearly increasing the learning rate (selected by the performance in the validation set from 1.25e-4, 2e-4, and 2.5e-4) to the desired value and then decreasing it using a cosine decay schedule. Batch size is 32 for CFP (224 × 224) and dermoscopic images, and 8 (512 × 512) for chest X-rays. Refer to Section B of Appendix for detailed descriptions.

Acknowledgments This work was supported in part by a Shenzhen-Hong KongMacao Science and Technology Plan Project (Category C Project) under Shenzhen Municipal Science and Technology Innovation Commission (project no. SGDX20230821092 359002) and a project under Innovation and Technology Support Programme of Hong Kong Innovation and Technology Commission (project no. ITS/202/23). C. Lian is supported in part by the Research Institute for Smart Ageing of the Hong Kong Polytechnic University.

References Adams, S. J., Stone, E., Baldwin, D. R., Vliegenthart, R., Lee, P., and Fintelmann, F. J. Lung cancer screening. The Lancet, 401(10374):390–408, 2023. Agarwal, P. K., Har-Peled, S., Varadarajan, K. R., et al. Geometric approximation via coresets. Combinatorial and computational geometry, 52(1):1–30, 2005.

5. Discussion and Conclusion In this paper, we systematically address the challenge of interpretable real-world disease screening by introducing evidential reasoning to enhance performance and interpretability simultaneously. Specifically, we first establish a comprehensive evaluation framework featuring novel clinically oriented metrics and extensive datasets for three clinical domains. The proposed EviScreen provides both retrospection interpretability and localization interpretability alongside its predictions, showing higher performance and bet-

Aggarwal, R., Sounderajah, V., Martin, G., Ting, D. S., Karthikesalingam, A., King, D., Ashrafian, H., and Darzi, A. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. NPJ digital medicine, 4(1):65, 2021. Bao, J., Sun, H., Deng, H., He, Y., Zhang, Z., and Li, X. Bmad: Benchmarks for medical anomaly detection. In 9

Evidential Reasoning Advances Interpretable Real-World Disease Screening

Fernando, T., Gammulle, H., Denman, S., Sridharan, S., and Fookes, C. Deep learning for medical anomaly detection– a survey. ACM Computing Surveys (CSUR), 54(7):1–37, 2021.

Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4042–4053, 2024. Cai, L., Fang, H., Xu, N., and Ren, B. Counterfactual causal-effect intervention for interpretable medical visual question answering. IEEE Transactions on Medical Imaging, 2024a.

Guo, J., Lu, S., Jia, L., Zhang, W., and Li, H. Encoderdecoder contrast for unsupervised anomaly detection in medical images. IEEE Transactions on Medical Imaging, 2023.

Cai, Y., Cai, Y.-Q., Tang, L.-Y., Wang, Y.-H., Gong, M., Jing, T.-C., Li, H.-J., Li-Ling, J., Hu, W., Yin, Z., et al. Artificial intelligence in the risk prediction models of cardiovascular disease and development of an independent validation screening tool: a systematic review. BMC medicine, 22(1):56, 2024b.

Har-Peled, S. and Kushal, A. Smaller coresets for k-median and k-means clustering. In Proceedings of the twentyfirst annual symposium on Computational geometry, pp. 126–134, 2005.

Cai, Y., Zhang, W., Chen, H., and Cheng, K.-T. Medianomaly: A comparative study of anomaly detection in medical images. Medical Image Analysis, 102:103500, 2025.

He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 16000–16009, 2022.

Cen, L.-P., Ji, J., Lin, J.-W., Ju, S.-T., Lin, H.-J., Li, T.-P., Wang, Y., Yang, J.-F., Liu, Y.-F., Tan, S., et al. Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks. Nature communications, 12(1):4828, 2021.

Hernández-Pérez, C., Combalia, M., Podlipnik, S., Codella, N. C., Rotemberg, V., Halpern, A. C., Reiter, O., Carrera, C., Barreiro, A., Helba, B., et al. Bcn20000: Dermoscopic lesions in the wild. Scientific data, 11(1):641, 2024. Hu, X., Gu, L., Kobayashi, K., Liu, L., Zhang, M., Harada, T., Summers, R. M., and Zhu, Y. Interpretable medical image visual question answering via multi-modal relationship graph learning. Medical Image Analysis, 97: 103279, 2024.

Chattopadhay, A., Sarkar, A., Howlader, P., and Balasubramanian, V. N. Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks. In 2018 IEEE winter conference on applications of computer vision (WACV), pp. 839–847. IEEE, 2018. Chen, Z., Du, Y., Hu, J., Liu, Y., Li, G., Wan, X., and Chang, T.-H. Multi-modal masked autoencoders for medical vision-and-language pre-training. In International Conference on Medical Image Computing and ComputerAssisted Intervention, pp. 679–689. Springer, 2022.

Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., et al. Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pp. 590–597, 2019.

Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.

Jiang, X., Liu, J., Wang, J., Nie, Q., Wu, K., Liu, Y., Wang, C., and Zheng, F. Softpatch: Unsupervised anomaly detection with noisy data. Advances in Neural Information Processing Systems, 35:15433–15445, 2022.

Ding, C., Pang, G., and Shen, C. Catching both gray and black swans: Open-set supervised anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 7388–7398, 2022.

Johnson, A. E., Pollard, T. J., Greenbaum, N. R., Lungren, M. P., Deng, C.-y., Peng, Y., Lu, Z., Mark, R. G., Berkowitz, S. J., and Horng, S. Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs. arXiv preprint arXiv:1901.07042, 2019a.

Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020.

Johnson, J., Douze, M., and Jégou, H. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 7 (3):535–547, 2019b. Kim, E., Kim, S., Seo, M., and Yoon, S. Xprotonet: diagnosis in chest radiography with global and local explanations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 15719– 15728, 2021.

Fanaroff, A. C., Califf, R. M., and Lopes, R. D. Highquality evidence to inform clinical practice. The Lancet, 394(10199):633–634, 2019. 10

Evidential Reasoning Advances Interpretable Real-World Disease Screening

Li, C., Shi, Y., Hu, J., Zhu, X. X., and Mou, L. Scale-aware contrastive reverse distillation for unsupervised medical anomaly detection. In The Thirteenth International Conference on Learning Representations, 2025.

Pan, J., Liu, C., Wu, J., Liu, F., Zhu, J., Li, H. B., Chen, C., Ouyang, C., and Rueckert, D. Medvlm-r1: Incentivizing medical reasoning capability of vision-language models (vlms) via reinforcement learning. In International Conference on Medical Image Computing and ComputerAssisted Intervention, pp. 337–347. Springer, 2025.

Lian, C., Zhou, H.-Y., Hu, Z., and Qin, J. BenchReAD: A systematic benchmark for retinal anomaly detection . In Medical Image Computing and Computer Assisted Intervention – MICCAI 2025, volume LNCS 15961, pp. 35 – 45. Springer Nature Switzerland, October 2025.

Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.

Liu, Z., Zhou, Y., Xu, Y., and Wang, Z. Simplenet: A simple network for image anomaly detection and localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20402–20411, 2023.

Roth, K., Milbich, T., Sinha, S., Gupta, P., Ommer, B., and Cohen, J. P. Revisiting training strategies and generalization performance in deep metric learning. In International Conference on Machine Learning, pp. 8242–8252. PMLR, 2020.

Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.

Roth, K., Pemula, L., Zepeda, J., Schölkopf, B., Brox, T., and Gehler, P. Towards total recall in industrial anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14318– 14328, 2022.

Marjanovic, S., Page, A., Stone, E., Currie, D. J., Rankin, N. M., Myers, R., Brims, F., Navani, N., and McBride, K. A. Systems mapping: a novel approach to national lung cancer screening implementation in australia. Translational Lung Cancer Research, 13(10):2466, 2024.

Sackett, D. L., Rosenberg, W. M., Gray, J. M., Haynes, R. B., and Richardson, W. S. Evidence based medicine: what it is and what it isn’t, 1996.

Moor, M., Banerjee, O., Abad, Z. S. H., Krumholz, H. M., Leskovec, J., Topol, E. J., and Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature, 616(7956):259–265, 2023. Nakayama, L. F., Restrepo, D., Matos, J., Ribeiro, L. Z., Malerbi, F. K., Celi, L. A., et al. Brset: A brazilian multilabel ophthalmological dataset of retina fundus photos. PLOS Digital Health, 3(7):e0000454, 2024. doi: 10.1371/journal.pdig.0000454. URL https://doi. org/10.1371/journal.pdig.0000454. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., ElNouby, A., et al. Dinov2: Learning robust visual features without supervision. Transactions on Machine Learning Research Journal, pp. 1–31, 2024.

Saporta, A., Gui, X., Agrawal, A., Pareek, A., Truong, S. Q., Nguyen, C. D., Ngo, V.-D., Seekins, J., Blankenberg, F. G., Ng, A. Y., et al. Benchmarking saliency methods for chest x-ray interpretation. Nature Machine Intelligence, 4(10):867–878, 2022. Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, pp. 618–626, 2017. Sener, O. and Savarese, S. Active learning for convolutional neural networks: A core-set approach. In International Conference on Learning Representations, 2018. Tschandl, P., Rosendahl, C., and Kittler, H. The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific data, 5(1):1–9, 2018.

Pachade, S., Porwal, P., Thulkar, D., Kokare, M., Deshmukh, G., Sahasrabuddhe, V., Giancardo, L., Quellec, G., and Mériaudeau, F. Retinal fundus multi-disease image dataset (rfmid): a dataset for multi-disease detection research. Data, 6(2):14, 2021.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.

Pacheco, A. G., Lima, G. R., Salomao, A. S., Krohling, B., Biral, I. P., De Angelo, G. G., Alves Jr, F. C., Esgario, J. G., Simora, A. C., Castro, P. B., et al. Pad-ufes-20: A skin lesion dataset composed of patient data and clinical images collected from smartphones. Data in brief, 32: 106221, 2020.

Wang, C., Liu, F., Chen, Y., Frazer, H., and Carneiro, G. Cross- and intra-image prototypical learning for multilabel disease diagnosis and interpretation. IEEE Transactions on Medical Imaging, 44(6):2568–2580, 2025. 11

Evidential Reasoning Advances Interpretable Real-World Disease Screening

Wolsey, L. A. and Nemhauser, G. L. Integer and Combinatorial Optimization. John Wiley & Sons, 2014. Xia, X., Li, Y., Xiao, G., Zhan, K., Yan, J., Cai, C., Fang, Y., and Huang, G. Benchmarking deep models on retinal fundus disease diagnosis and a large-scale dataset. Signal Processing: Image Communication, 127:117151, 2024. ISSN 0923-5965. doi: https://doi.org/10.1016/j.image.2024.117151. URL https://www.sciencedirect.com/ science/article/pii/S0923596524000523. Yan, S., Yu, Z., Primiero, C., Vico-Alonso, C., Wang, Z., Yang, L., Tschandl, P., Hu, M., Ju, L., Tan, G., et al. A multimodal vision foundation model for clinical dermatology. Nature Medicine, pp. 1–12, 2025. Yang, Z., Xu, X., Zhang, J., Wang, G., Kalra, M. K., and Yan, P. Chest x-ray foundation model with global and local representations integration. IEEE Transactions on Medical Imaging, 2025. Yilmaz, A., Yasar, S. P., Gencoglan, G., and Temelkuran, B. Derm12345: A large, multisource dermatoscopic skin lesion dataset with 40 subclasses. Scientific Data, 11(1): 1302, 2024. Zhang, J., Lin, S., Cheng, T., Xu, Y., Lu, L., He, J., Yu, T., Peng, Y., Zhang, Y., Zou, H., et al. Retfoundenhanced community-based fundus disease screening: real-world evidence and decision curve analysis. NPJ digital medicine, 7(1):108, 2024. Zhang, S. and Metaxas, D. On the challenges and perspectives of foundation models for medical image analysis. Medical image analysis, 91:102996, 2024. Zhang, Y., Luo, L., Dou, Q., and Heng, P.-A. Triplet attention and dual-pool contrastive learning for clinic-driven multi-label medical image classification. Medical image analysis, 86:102772, 2023. Zhou, H.-Y., Lian, C., Wang, L., and Yu, Y. Advancing radiograph representation learning with masked record modeling. In The Eleventh International Conference on Learning Representations, 2023a. Zhou, Y., Chia, M. A., Wagner, S. K., Ayhan, M. S., Williamson, D. J., Struyven, R. R., Liu, T., Xu, M., Lozano, M. G., Woodward-Court, P., et al. A foundation model for generalizable disease detection from retinal images. Nature, 622(7981):156–163, 2023b. Zhou, Y., Wang, Z., Wu, Y., Ong, A. Y., Wagner, S., Ruffell, E., Chia, M., Guan, Z., Ju, L., Engelmann, J., et al. Revealing the impact of pre-training data on medical foundation models. PREPRINT (Version 1) available at Research Square, 2025. 12

Evidential Reasoning Advances Interpretable Real-World Disease Screening

A. More Details of Benchmarks and Data To establish a comprehensive evaluation framework, we develop benchmarks across three critical medical domains: ophthalmology, radiology, and dermatology. Figure 3 in the main paper illustrates the distribution of abnormality categories and the composition of the training set. This section provides detailed specifications to facilitate understanding and reproducibility. A.1. Ophthalmology Our ophthalmology benchmark utilizes color fundus photography (CFP) with training and validation sets derived from EDDFS (Xia et al., 2024) and BRSET (Nakayama et al., 2024). We categorize samples based on abnormality presence: cases without detected abnormalities are classified as normal, while those exhibiting any abnormality are classified as pathological. To simulate real-world disease screening scenarios, we employ two external datasets: JSIEC (Cen et al., 2021) and RIADD (Pachade et al., 2021) as test sets. A.2. Radiology For the radiology benchmark using chest X-rays, we utilize MIMIC-CXR (Johnson et al., 2019a) for training and validation. Our preprocessing pipeline retains only samples with “AP” or “PA” view positions and excludes samples containing uncertain labels. Normal cases are defined as those annotated with “No Finding”, while pathological cases include samples with any abnormal findings except “Support Devices”. We use the official CheXpert (Irvin et al., 2019) evaluation set as our test dataset, chosen for its board-certified radiologist annotations ensuring label reliability. Following the original CheXpert recommendations, our evaluation focuses on five primary pathological categories: atelectasis, cardiomegaly, consolidation, edema, and pleural effusion. A.3. Dermatology The dermatology benchmark incorporates dermoscopic images with training and validation sets constructed from HAM10000 (Tschandl et al., 2018), BCN20000 (Hernández-Pérez et al., 2024), and PAD-UFES-20 (Pacheco et al., 2020). Since dermoscopic examination inherently includes few completely normal samples because clinicians typically examine only areas that appear different from normal skin, we adapt our classification scheme accordingly. Samples annotated as “benign” are treated as “normal” cases, while “malignant” samples constitute the “pathological” cases for detection. We employ Derm12345 (Yilmaz et al., 2024) as our test set for evaluation.

B. More Comprehensive Implementation Details Our code is implemented using PyTorch 2.4.1 (Paszke et al., 2019). All experiments are carried out with Nvidia GeForce RTX 3090 GPUs. We employ state-of-the-art foundation models for each modality: RETFound-Dinov2 (Zhou et al., 2025) for CFP, CheXFound (Yang et al., 2025) for chest X-rays, and PanDerm (Yan et al., 2025) for dermoscopic images. The foundation models chosen as regional feature extractors are based on ViT-L (Dosovitskiy et al., 2020; Vaswani et al., 2017), which consists of 24 transformer blocks. Features in the layers of 7 and 17 are selected and aggregated by adaptive average pooling to generate desired regional features. RETFound-Dinov2 (Zhou et al., 2025) for CFP receives input resolution of 224×224 with the patch size of 14. CheXFound (Yang et al., 2025) for chest X-rays receives input resolution of 512×512 with the patch size of 16. PanDerm (Yan et al., 2025) for dermoscopic images receives input resolution of 224×224 with the patch size of 16. The dimension of evidence vectors is 1,024. Following previous related work (Roth et al., 2022; Lian et al., 2025), no data augmentation is applied to avoid including new abnormality or losing original abnormality in the images. During the construction and usage of dual knowledge banks, Faiss 1.8.0 (Johnson et al., 2019b) is adopted for the nearest neighbor retrieval and distance computations. For evidential reasoning, AdamW (Loshchilov & Hutter, 2017) is adopted as the default optimizer, with a weight decay of 0.05, β1 of 0.9, and β2 of 0.95. We employ a “warm-up” strategy by linearly increasing the learning rate (selected based on validation performance from 1.25e-4, 2e-4, and 2.5e-4) to the desired value and then decreasing it using a cosine decay schedule. Batch size is 32 for CFP and dermoscopic images, and 8 for chest X-rays. The network for evidential reasoning is based on Transformers (Dosovitskiy et al., 2020; Vaswani et al., 2017), with an embedding dimension of 1,024, 8 attention heads, 256 patches, and a depth of 4 layers. Each transformer block comprises a cross-attention module for evidence awareness, followed by a self-attention module for inter-patch refinement. We adopt 2D unlearnable sin-cos positional 13

Evidential Reasoning Advances Interpretable Real-World Disease Screening

embeddings to introduce the positional information.

C. More Examples of Interpretability Visualization Input

Abnormality map

JSIEC 65

Query

Evidence from normal knowledge bank

Evidence from pathological knowledge bank

1

Large optic cup

Normal

Normal

Normal

Increased cup disc Increased cup disc Increased cup disc

Normal

Normal

Normal

Increased cup disc Increased cup disc Increased cup disc

Normal

Normal

Normal

Increased cup disc Increased cup disc Increased cup disc

Normal

Normal

Normal

Diabetic retinopathy Diabetic retinopathy Diabetic retinopathy

Normal

Normal

Normal

Diabetic retinopathy Diabetic retinopathy Diabetic retinopathy

2

3

JSIEC 129

1

Diabetic retinopathy

2

& macular edema

& macular edema

3

JSIEC 420

Normal

Normal

Normal

Diabetic retinopathy Diabetic retinopathy Diabetic retinopathy

Normal

Normal

Normal

Vascular occlusion

RVO

RVO

Normal

Normal

Normal

RVO

RVO

Vascular occlusion

Normal

Normal

Normal

RVO

1

CRVO

2

3

Vascular occlusion Vascular occlusion & hemorrhage

Figure 7. More examples of interpretability visualization for ophthalmology.

14

Evidential Reasoning Advances Interpretable Real-World Disease Screening Input

Xpert 118

Abnormality map

Query

Evidence from normal knowledge bank

Evidence from pathological knowledge bank

1

Pleural Effusion

Normal

Normal

Normal

Pleural Effusion

Pleural Effusion

Pleural Effusion

Normal

Normal

Normal

Pleural Effusion

Pleural Effusion

Pleural Effusion

Normal

Normal

Normal

Pleural Effusion

Pleural Effusion

Pleural Effusion

Normal

Normal

Normal

Atelectasis

Atelectasis

Atelectasis

Normal

Normal

Normal

Atelectasis

Atelectasis

Atelectasis

Normal

Normal

Normal

Atelectasis

Atelectasis

Atelectasis

Normal

Normal

Normal

Pleural Effusion

Pleural Effusion

Pleural Effusion

Normal

Normal

Normal

Pleural Effusion

Pleural Effusion

Pleural Effusion

Normal

Normal

Normal

Pleural Effusion

Pleural Effusion

Pleural Effusion

2

3

pert 419

1

Atelectasis

2

3

Xpert 413

1

Pleural Effusion

2

3

Figure 8. More examples of interpretability visualization for radiology.

15

Evidential Reasoning Advances Interpretable Real-World Disease Screening Input

Abnormality map

45 11281

Query

Evidence from normal knowledge bank

Evidence from pathological knowledge bank

1

Basal Cell Carcinoma (BCC)

Non-malignant

Non-malignant

Non-malignant

BCC

BCC

BCC

Non-malignant

Non-malignant

Non-malignant

BCC

BCC

BCC

Non-malignant

Non-malignant

Non-malignant

BCC

BCC

BCC

Non-malignant

Non-malignant

Non-malignant

Melanoma

Melanoma

Melanoma

Non-malignant

Non-malignant

Non-malignant

Melanoma

Melanoma

Melanoma

Non-malignant

Non-malignant

Non-malignant

Melanoma

Melanoma

Melanoma

Non-malignant

Non-malignant

Non-malignant

AKIEC

AKIEC

AKIEC

Non-malignant

Non-malignant

Non-malignant

AKIEC

AKIEC

AKIEC

Non-malignant

Non-malignant

Non-malignant

AKIEC

AKIEC

AKIEC

2

3

45 12285

1

Melanoma

2

3

345 11157

1

AK

2

3

Figure 9. More examples of interpretability visualization for dermatology.

16

Evidential Reasoning Advances Interpretable Real-World Disease Screening

D. More Experimental Results In this section, we provide additional experimental results that are omitted from the main body due to space constraints. D.1. ROC Curves for Test Sets Here, we present ROC curves for distinguishing normal and pathological samples on the entire test sets.

(a)

(b)

(c)

(d)

Figure 10. ROC curves for test sets: (a) JSIEC, (b) RIADD, (c) CheXpert, and (d) Derm12345.

17

Evidential Reasoning Advances Interpretable Real-World Disease Screening

D.2. Category-Level Performance of Disease Screening The category-level performance for disease screening is presented, based on evaluation with the test sets. D.2.1. C ATEGORY-L EVEL P ERFORMANCE ON JSIEC Table 3. Category-level results for disease screening on JSIEC regarding Spe@100%R (%). The best results are in bold, and the second-best results are underlined. * A variant of PatchCore that uses the same foundation models as ours. Category Mean Tessellated fundus Large optic cup DR1 DR2 DR3 Possible glaucoma Optic atrophy Severe hypertensive retinopathy Disc swelling and elevation Dragged Disc Congenital disc abnormality Retinitis pigmentosa Bietti crystalline dystrophy Peripheral retinal degeneration and break Myelinated nerve fiber Vitreous particles Fundus neoplasm BRVO CRVO Massive hard exudates Yellow-white spots-flecks Cotton-wool spots Vessel tortuosity Chorioretinal atrophy-coloboma Preretinal hemorrhage Fibrosis Laser Spots Silicon oil in eye Blur fundus without PDR Blur fundus with suspected PDR RAO Rhegmatogenous RD CSCR VKH disease Maculopathy ERM MH Pathological myopia

Ours 91.27 44.74 44.74 50.00 100.00 100.00 100.00 97.37 100.00 78.95 100.00 100.00 100.00 100.00 100.00 100.00 100.00 100.00 100.00 100.00 100.00 92.11 100.00 44.74 100.00 100.00 100.00 100.00 100.00 84.21 100.00 86.84 92.11 86.84 86.84 100.00 86.84 92.11 100.00

FM 78.39 15.79 15.79 57.89 100.00 100.00 100.00 36.84 100.00 76.32 92.11 92.11 76.32 100.00 10.53 86.84 100.00 92.11 100.00 100.00 100.00 100.00 100.00 7.89 100.00 100.00 92.11 100.00 36.84 36.84 100.00 100.00 57.89 76.32 60.53 100.00 86.84 71.05 100.00

PatchCore* PatchCore NFM-DRA DRA 82.34 73.75 82.96 71.12 71.05 36.84 36.84 63.16 7.89 52.63 50.00 5.26 10.53 2.63 2.63 13.16 50.00 31.58 73.68 100.00 92.11 50.00 94.74 100.00 92.11 73.68 78.95 63.16 76.32 71.05 73.68 50.00 89.47 97.37 100.00 100.00 92.11 76.32 73.68 5.26 97.37 100.00 100.00 78.95 97.37 100.00 100.00 97.37 97.37 65.79 86.84 100.00 97.37 73.68 97.37 100.00 97.37 94.74 100.00 50.00 97.37 100.00 100.00 50.00 97.37 94.74 100.00 100.00 97.37 76.32 92.11 55.26 100.00 76.32 94.74 97.37 92.11 73.68 92.11 100.00 97.37 100.00 100.00 100.00 44.74 73.68 100.00 100.00 71.05 78.95 92.11 89.47 68.42 76.32 73.68 50.00 97.37 94.74 100.00 100.00 97.37 100.00 100.00 89.47 97.37 81.58 92.11 92.11 97.37 97.37 100.00 100.00 94.74 100.00 100.00 92.11 92.11 81.58 81.58 0.00 97.37 63.16 97.37 63.16 97.37 73.68 71.05 0.00 97.37 73.68 76.32 34.21 71.05 55.26 71.05 21.05 86.84 65.79 73.68 89.47 52.63 65.79 78.95 86.84 71.05 34.21 50.00 81.58 47.37 47.37 47.37 84.21 97.37 92.11 100.00 100.00

18

SCRD4AD 78.46 60.53 47.37 13.16 13.16 63.16 86.84 89.47 92.11 76.32 97.37 97.37 94.74 100.00 97.37 84.21 100.00 94.74 65.79 94.74 100.00 73.68 65.79 47.37 97.37 94.74 97.37 100.00 100.00 86.84 47.37 47.37 92.11 73.68 84.21 92.11 78.95 39.47 94.74

EDC 42.11 13.16 50.00 5.26 0.00 5.26 5.26 47.37 47.37 55.26 47.37 47.37 42.11 47.37 60.53 52.63 63.16 47.37 28.95 60.53 50.00 47.37 47.37 47.37 47.37 47.37 60.53 50.00 47.37 0.00 23.68 52.63 50.00 60.53 52.63 47.37 47.37 47.37 47.37

SimpleNet 49.10 28.95 50.00 2.63 28.95 47.37 47.37 47.37 42.11 52.63 50.00 52.63 55.26 55.26 55.26 52.63 55.26 55.26 42.11 52.63 68.42 52.63 55.26 42.11 50.00 57.89 57.89 63.16 52.63 47.37 44.74 57.89 50.00 52.63 52.63 52.63 44.74 42.11 47.37

CIPL 78.60 0.00 60.53 2.63 100.00 100.00 94.74 92.11 100.00 39.47 86.84 97.37 81.58 97.37 73.68 86.84 100.00 81.58 100.00 97.37 100.00 97.37 97.37 68.42 97.37 57.89 94.74 97.37 97.37 28.95 39.47 76.32 50.00 73.68 76.32 86.84 81.58 76.32 97.37

Evidential Reasoning Advances Interpretable Real-World Disease Screening

D.2.2. C ATEGORY-L EVEL P ERFORMANCE ON RIADD Table 4. Category-level results for disease screening on RIADD regarding Spe@100%R (%). The best results are in bold, and the second-best results are underlined. * A variant of PatchCore that uses the same foundation models as ours. Category Mean DR ARMD MH DN MYA BRVO TSLN ERM LS MS CSR ODC CRVO TV AH ODP ODE ST AION PT RT RS CRS EDN RPEC MHL RP OTHER

Ours 55.35 69.81 77.13 11.36 37.22 69.81 91.18 3.29 27.06 97.91 90.28 20.78 7.47 84.45 82.81 98.80 25.11 22.12 59.79 63.08 48.88 99.40 69.96 35.28 98.95 24.66 29.75 96.86 6.58

FM 51.75 58.30 60.99 3.14 37.67 44.54 76.38 13.45 30.34 94.32 91.33 47.09 4.48 87.44 90.28 88.79 7.77 24.96 8.97 47.68 62.48 90.28 64.87 41.85 98.21 50.22 45.59 41.41 36.17

PatchCore* 41.63 10.76 23.32 31.69 0.60 18.39 48.88 5.08 3.59 43.65 41.70 22.72 0.15 88.94 99.40 98.65 5.38 7.47 73.99 98.21 19.58 100.00 61.88 26.91 48.13 17.79 12.41 95.96 60.54

PatchCore 29.36 2.24 28.10 11.21 0.00 42.30 27.20 16.14 6.58 14.20 50.82 1.05 0.15 23.92 77.28 40.36 0.75 13.15 54.41 69.21 5.98 94.17 60.84 50.07 30.04 0.00 6.73 77.43 17.64

NFM-DRA 37.64 1.94 47.09 15.10 1.35 57.55 23.77 13.30 8.22 19.13 74.14 2.09 0.15 41.26 74.14 94.32 0.60 18.54 52.47 82.21 4.63 95.37 60.39 67.12 82.51 0.00 4.93 90.43 21.08

DRA 29.74 10.91 45.89 0.00 2.24 73.54 20.33 3.74 22.12 28.70 86.55 0.00 0.75 65.92 12.71 98.65 3.44 1.94 35.13 2.99 0.75 77.73 9.42 1.49 99.25 28.40 14.80 82.96 2.24

SCRD4AD 43.98 20.63 47.23 8.97 1.49 60.69 52.02 5.23 71.45 52.77 58.89 32.29 1.64 66.52 86.55 97.31 26.16 18.09 25.56 50.67 22.42 97.31 26.01 70.55 81.61 20.78 28.55 94.77 5.23

EDC 9.08 0.75 1.64 0.15 0.30 9.12 3.74 6.73 1.05 1.94 28.25 3.29 0.75 4.19 75.64 0.30 0.30 10.16 26.76 4.04 11.06 4.19 2.24 18.98 10.16 2.24 22.42 1.35 2.39

SimpleNet 14.26 0.00 2.39 2.99 0.00 9.72 15.25 0.75 4.78 32.44 6.43 1.20 0.75 22.87 23.32 31.39 11.06 9.72 23.32 16.29 4.78 51.27 9.72 27.20 46.19 6.58 3.74 30.34 4.93

CIPL 43.99 24.22 62.93 2.09 19.28 28.25 39.46 10.16 66.82 80.87 31.69 16.29 16.14 81.61 62.48 100.00 6.13 14.35 33.48 29.00 36.62 76.83 39.91 48.13 99.85 35.28 60.99 98.80 10.01

D.2.3. C ATEGORY-L EVEL P ERFORMANCE ON C HE X PERT Table 5. Category-level results for disease screening on CheXpert regarding Spe@100%R (%). The best results are in bold, and the second-best results are underlined. * A variant of PatchCore that uses the same foundation models as ours. Category Mean Atelectasis Cardiomegaly Consolidation Edema Pleural Effusion

Ours 68.72 28.72 34.04 100.00 85.11 95.74

FM 60.00 35.11 25.53 100.00 77.66 61.70

PatchCore* PatchCore NFM-DRA DRA 15.53 11.06 13.62 18.72 8.51 13.83 17.02 28.72 11.70 2.13 5.32 9.57 34.04 10.64 17.02 20.21 11.70 2.13 5.32 26.60 11.70 26.60 23.40 8.51

19

SCRD4AD 5.53 3.19 4.26 17.02 0.00 3.19

EDC 14.89 0.00 4.26 26.60 15.96 27.66

SimpleNet 8.94 7.45 2.13 23.40 4.26 7.45

CIPL 49.15 30.85 12.77 55.32 82.98 63.83

Evidential Reasoning Advances Interpretable Real-World Disease Screening

D.2.4. C ATEGORY-L EVEL P ERFORMANCE ON D ERM 12345 Table 6. Category-level results for disease screening on Derm12345 regarding Spe@100%R (%). The best results are in bold, and the second-best results are underlined. * A variant of PatchCore that uses the same foundation models as ours. Category Mean ALM ANM LM LMM MEL AK BCC BD CH MPD SCC DFSP KS

Ours 78.49 31.30 91.94 59.07 96.81 29.71 75.85 82.05 95.49 84.04 99.29 81.20 98.78 94.83

FM 55.41 22.63 68.56 0.90 96.99 22.43 3.56 31.79 92.86 16.55 99.25 70.65 97.72 96.48

PatchCore* PatchCore NFM-DRA DRA 46.90 36.82 41.80 56.91 34.16 21.54 20.86 27.74 76.60 69.79 82.10 89.60 5.40 23.26 29.05 33.60 43.44 60.64 74.15 95.31 0.33 12.68 11.92 15.53 17.72 1.29 3.19 17.39 22.92 3.17 3.56 0.05 22.19 13.79 13.27 49.20 81.34 65.81 73.93 79.24 82.60 51.31 64.02 98.78 64.24 22.93 22.38 57.57 92.56 90.42 93.48 97.60 66.12 42.00 51.47 78.25

SCRD4AD 32.86 22.07 31.24 30.50 57.90 0.61 14.07 0.99 38.73 43.27 40.27 46.55 74.72 26.25

EDC 22.15 9.92 49.97 2.01 11.86 8.20 13.71 6.40 2.34 16.96 67.65 17.70 62.81 18.42

SimpleNet 20.94 4.58 31.80 0.92 53.83 8.59 1.01 1.27 2.66 10.38 67.11 23.62 53.45 12.94

CIPL 54.95 33.28 89.64 42.55 93.49 20.20 12.02 8.54 68.96 49.02 93.77 12.60 97.28 92.99

D.3. Deployment Efficiency We report the deployment cost of the dual knowledge banks. The ophthalmology, radiology, and dermatology banks contain 256k, 1,024k, and 196k vectors, respectively, with corresponding memory footprints of 1,002 MB, 4,004 MB, and 766 MB. In the largest radiology setting, the dual knowledge banks contain 1,024k vectors, require 4,004 MB of memory, and take approximately 0.6 s for retrieval, indicating that the evidence retrieval process remains lightweight for deployment. D.4. Stress Test on Pathological Bank Contamination The pathological knowledge bank inevitably contains normal regions because image-level pathological labels do not provide patch-level annotations. To evaluate the robustness of our framework to such contamination, we conduct a stress test on JSIEC by injecting 20% normal samples into the pathological bank during construction. We also evaluate a SoftPatchlike (Jiang et al., 2022) denoising strategy during knowledge bank construction. As shown in the results below, our model maintains stable performance under stress testing on the JSIEC dataset, and explicit denoising yields comparable scores. Table 7. Stress test and denoising results on JSIEC (%).

Metric AUROC Spe@99%R

Original 98.06 91.62

Stress test 97.71 90.10

Denoising strategy 96.90 92.11

D.5. Quantitative Interpretability Evaluation We evaluate localization interpretability using expert-annotated lesion masks from a reader study on 20 fundus images. As shown in Table 8, our approach achieves more accurate localization than PatchCore* and CIPL. Table 8. Reader study results for localization interpretability.

Method PatchCore* CIPL Ours

Dice 0.45 ± 0.13 0.58 ± 0.16 0.66 ± 0.06

20

IoU 0.30 ± 0.12 0.42 ± 0.16 0.50 ± 0.07

Record · ID 187294 · SHA-256 14080e1ca5984af4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.