arXiv:2605.26068v1 [cs.LG] 25 May 2026
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark Xu Yao∗
Siyuan Zhou∗
Zhenbo Wu∗
[email protected] Shanghai University of Finance and Economics Shanghai, China
[email protected] Shanghai University of Finance and Economics Shanghai, China
[email protected] Shanghai University of Finance and Economics Shanghai, China
Chaochuan Hou∗
Shuang Liang∗
Shiping Wang∗
[email protected] Shanghai University of Finance and Economics Shanghai, China
[email protected] Shanghai University of Finance and Economics Shanghai, China
[email protected] Ant Group Shanghai, China
Hailiang Huang†
Songqiao Han†
Minqi Jiang†
[email protected] Key Laboratory of Interdisciplinary Research of Computation and Economics Shanghai University of Finance and Economics Shanghai, China
[email protected] Key Laboratory of Interdisciplinary Research of Computation and Economics Shanghai University of Finance and Economics Shanghai, China
[email protected] Key Laboratory of Interdisciplinary Research of Computation and Economics Shanghai University of Finance and Economics Shanghai, China
Abstract Weakly supervised anomaly detection (WSAD) has developed in three primary directions: incomplete, inexact, and inaccurate supervision. However, these directions remain isolated, lacking a unified framework to assess whether they address unique challenges or share fundamental mechanics. This paper introduces WSADBench, the first benchmark that unifies evaluation across distinct weakly supervised scenarios, benchmarking diverse approaches from specialized WSAD methods to advanced tabular foundation models. WSADBench establishes standardized protocols to evaluate 36 algorithms across 4 modalities by systematically varying label quantity, granularity, and quality, revealing the performance boundaries of various methods. Based on over 700K experiments, WSADBench reveals four critical insights: (i) Strong intrinsic correlations exist between these weak supervision scenarios, challenging the isolation of current research directions. (ii) Specialized WSAD algorithms excel only in extreme label-scarcity regimes but are quickly dominated by tabular foundation models and general classification methods as supervision increases or in OOD scenarios. (iii) Unlabeled data shows inconsistent utility across settings, with marginal gains compared to label refinement. (iv) Models exhibit asymmetric sensitivity to ∗ Equal contribution. † Corresponding author.
This work is licensed under a Creative Commons Attribution 4.0 International License. KDD 2026, Jeju Island, Republic of Korea. © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2259-2/2026/08 https://doi.org/10.1145/3770855.3817536
different types of label noise. We release WSADBench as an opensource benchmark with code and datasets to facilitate future WSAD research: https://github.com/SUFE-AILAB/WSADBench.
CCS Concepts • Computing methodologies → Anomaly detection.
Keywords Anomaly Detection; Weakly Supervised Learning; Benchmark ACM Reference Format: Xu Yao, Siyuan Zhou, Zhenbo Wu, Chaochuan Hou, Shuang Liang, Shiping Wang, Hailiang Huang, Songqiao Han, and Minqi Jiang. 2026. Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD 2026), August 9–13, 2026, Jeju Island, Republic of Korea. ACM, New York, NY, USA, 49 pages. https://doi.org/10.1145/3770855.3817536
1
Introduction
Anomaly detection (AD) plays a critical role in risk-sensitive tasks [3, 38, 66], ranging from fraud prevention to medical diagnosis. Yet, practical systems must typically learn under imperfect supervision: labels that are scarce, coarse-grained, or corrupted by noise. This reality has given rise to Weakly Supervised Anomaly Detection (WSAD), where algorithms compensate for label deficiencies by leveraging unlabeled data, exploiting aggregate supervision signals, or explicitly modeling annotation noise. These include positiveunlabeled (PU) learning methods that treat anomaly detection as learning from labeled anomalies and unlabeled samples [39, 46], Multiple Instance Learning (MIL) approaches for coarse bag-level
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 1: Comparison of WSADBench with existing anomaly detection benchmarks. Benchmark
Year
Modality
Scope
# Dataset
# Model
Task Level Inst.
ADBench [23] AnoShift [14] UBnormal [2] BMAD [6] NLP-ADBench [29] MMAD [26] Text-ADBench [57]
2022 2022 2022 2024 2024 2025 2025
Tab., Img., Txt. Tab. Vid. Img. Txt. Img., Txt. Txt.
ML/DL ML/DL DL DL ML/DL DL ML/DL
57 1 10 6 8 4 12
30 11 3 15 19 12 12
✓ ✓ ✓ ✓ ✓ ✓ ✓
WSADBench (Ours)
2026
Tab., Img., Txt., Vid.
ML/DL
61
36
✓
labels [50, 52], and noise-robust techniques that explicitly model annotation errors [13, 67]. However, this parallel evolution of isolated research directions raises fundamental questions that remain unanswered. (i) Are these three supervision deficiencies truly independent research problems? For instance, bag-level labels in inexact supervision can be broadcast to all instances within each bag, converting the problem into inaccurate supervision with noisy instance-level labels. (ii) Do WSAD algorithms necessarily require modality-specific designs and specialized architectures? Recent tabular foundation models like TabPFN [21] and LimiX [64] achieve competitive performance on tabular tasks without anomaly-specific inductive biases, questioning whether specialized WSAD techniques remain necessary. (iii) Are existing evaluations even comparable across studies? Different works employ inconsistent feature extractors (e.g., I3D[9] vs. C3D[53] for video), incompatible ground-truth label definitions, and varied experimental protocols, preventing rigorous cross-method comparison. These three questions motivate a systematic rethinking of the WSAD landscape. We present WSADBench, a comprehensive benchmark designed to address these fundamental questions through three key design principles. First, to examine cross-scenario connections, we evaluate whether techniques designed for one supervision setting can transfer to others. Second, to assess algorithm necessity, we benchmark diverse approaches ranging from specialized WSAD techniques to general-purpose classifiers and tabular foundation models, determining when domain-specific designs provide value. Third, to enable fair comparison, we establish unified protocols with standardized feature extractors, consistent ground-truth alignments, and controlled experimental settings across diverse data modalities. Contributions. Driven by comprehensive evaluation spanning over 700k experiments, this work systematizes the WSAD domain and delivers the following contributions: • Unified framework for cross-scenario evaluation. We establish the first benchmark that systematically evaluates WSAD methods across different supervision deficiencies, revealing connections and transfer potential between traditionally isolated research directions. • Novel analytical perspectives and findings. Beyond standard leaderboard comparisons, we introduce analytical perspectives such as progressive OOD distribution shift, decoupled asymmetric noise injection, inclusion of tabular foundation models
OOD
Bag
✓
✓ ✓ ✓ ✓
✓
✓
Supervision Unsup.
Weak.
✓ ✓ ✓ ✓ ✓ ✓ ✓
✓
✓
TFM
✓ ✓ ✓ ✓ ✓
✓
in WSAD method comparison, and joint labeled–unlabeled data utility mapping. These perspectives reveal new insights into algorithmic robustness, noise sensitivity, and the effective boundaries of specialized WSAD designs. • Standardized evaluation protocols. Through systematic experiments varying label quantity, granularity, and quality, we standardize experimental settings, enabling rigorous method comparisons that were previously hindered by the inconsistent protocols prevalent in existing WSAD research. • Comprehensive coverage and open-source platform. We benchmark 36 diverse algorithms across 4 modalities, including specialized WSAD methods, deep learning, and tabular foundation models. To foster reproducibility, we publicly release all code, datasets and standardized preprocessing pipelines at https://github.com/SUFE-AILAB/WSADBench.
2 Related Work 2.1 Anomaly Detection Methods Traditional unsupervised methods operate on unlabeled data containing both normal and anomalous instances, distinguishing outliers based on distributional assumptions such as low-density regions (e.g., LOF [8]) or distance-based isolation (e.g., kNN [5]). Semi-supervised approaches assume access to clean normal samples for one-class learning [20, 45]. However, in real-world scenarios, only scarce label information is available, which is neither purely normal nor fully reliable. Exploiting such limited supervision has been shown to substantially improve detection performance [23], motivating the study of Weakly-Supervised Anomaly Detection (WSAD) [39]. Drawing upon the established taxonomy of weakly supervised learning [71], we introduce this framework to the AD domain and propose WSADBench, the first systematic WSAD benchmark that categorizes methods by three types of label imperfection: incomplete, inexact, and inaccurate supervision. Incomplete Supervision refers to scenarios where only a subset of anomalies are labeled. Methods in this category employ techniques such as anomaly score learning in DevNet [40] and PReNet [39], feature-guided strategies in FEAWAD [70], or data augmentation in RoSAS [58] to estimate class priors or re-weight unlabeled samples, mitigating bias from label scarcity. In contrast, Inexact Supervision deals with coarse-grained bag-level labels rather than instance-level annotations. Particularly prominent in video anomaly detection [1], temporal MIL methods [50] leverage ranking losses, feature magnitude
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 2: Classification of benchmarked AD methods. constraints introduced by RTFM [52], or center-based regularization employed in AR-Net [54] to localize anomalies within sequences. Complementing these settings, Inaccurate Supervision addresses the challenge of label noise, where annotations themselves may be erroneous. Although robust learning has been extensively studied in classification [49, 61], its application to AD remains less systematic, with emerging work on noise-tolerant losses and confidence-based filtering [47].
2.2
Anomaly Detection Benchmarks
Several benchmarks have been proposed for anomaly detection. Most benchmarks focus on unsupervised or semi-supervised settings, treating weak supervision as a minor extension rather than a systematic research direction. For instance, AnoShift [14] and our prior work ADBench [23] cover tabular, image, and text modalities but primarily evaluate unsupervised methods with limited weak supervision scenarios. Similarly, modality-specific benchmarks such as BMAD [6] for medical images, UBnormal [2] for video AD, NLP-ADBench [29] and Text-ADBench [57] for text, along with video AD surveys [1], provide valuable domain-focused evaluations but lack unified cross-modal assessment. Beyond standard AD, OpenOOD [60] addresses out-of-distribution detection, while MMAD [26] explores multimodal foundation models but remains confined to unsupervised settings. Furthermore, WRENCH [62], though designed for weak supervision, targets general classification rather than AD-specific challenges. We compare WSADBench with existing benchmarks, as summarized in Table 1. Additionally, current benchmarks lack systematic evaluation of tabular foundation models under weak supervision—while recent works [6, 26, 57] explore zero-shot capabilities, their robustness across different weak supervision scenarios (incomplete, inexact, inaccurate) and OOD settings remains unexplored. WSADBench addresses these gaps by providing the first comprehensive benchmark dedicated to weakly-supervised anomaly detection. It unifies evaluation across three supervision deficiency types, spans four modalities, and systematically assesses both specialized WSAD algorithms and tabular foundation models under standardized protocols, yielding critical insights into the effective boundaries and robustness of different approaches.
3
WSADBench
We design WSADBench (Figure 1) to address the fragmentation and inconsistency in current WSAD research. Our rationale comprises three core principles: (1) Cross-Scenario Transferability—we break scenario isolation by enabling methods specialized for one deficiency to be evaluated across all scenarios, testing whether these supervisions are independent problems or convertible formulations. (2) Broad Baseline Inclusion—we require specialized WSAD algorithms and general-purpose tabular foundation models to compete on the same stage, evaluating the necessity of model specialization. (3) Standardized Fair Evaluation—we enforce unified feature extractors, standardized metrics, and consistent evaluation protocols to resolve the inconsistent practices found in existing studies.
3.1
Unified Benchmarking Framework.
While these three regimes have traditionally been studied in isolation, WSADBench unifies them by establishing explicit connections
Method
Category
Weakly-supervised (Instance) DevNet [40] Score Learning DeepSAD [46] Score Learning PReNet [39] Score Learning REPEN [37] Repr. Learning XGBOD [65] Repr. Learning RoSAS [58] Data Aug. Dual-MGAN [30] Data Aug. FEAWAD [70] Reconstruction DDAE [47] Diffusion DAE SOEL-NTL [40] Pseudo-Labeling AA-BiGAN [51] GAN-based GANomaly [4] GAN-based Weakly-supervised (Bag) Sultani [50] Vanilla MIL RTFM [52] Magnitude MIL MGFN [11] Magnitude MIL AR-Net [54] Dynamic MIL VadCLIP [56] Language-Guided MIL UR-DMU [69] Uncertainty-Aware MIL GCN-Anomaly [68] Label Denoising PUMA [41] PU MIL
Method
Category
Unsupervised (Instance) IForest [32] Isolation-based AutoEncoder [72] Reconstruction VAE [27] Reconstruction PCA [48] Reconstruction DeepSVDD [45] Deep One-class ECOD [31] Probabilistic CBLOF [24] Cluster-based LOF [8] Density-based LUNAR [16] GNN-based Supervised (Instance) XGBoost [10] GBDT CatBoost [42] GBDT FTTransformer [19] Deep (Sup.) TabM [17] Deep (Sup.) TabR-S [18] Deep (Sup.) Foundation Models (Instance) TabPFN [21] Found. Model LimiX [64] Found. Model
between supervision scenarios: (i) Incomplete-to-Inexact Transferability. Methods designed for incomplete supervision (e.g., weaklysupervised AD algorithms) are evaluated on inexact supervision tasks (video MIL), revealing whether designs optimized for label scarcity can generalize to label granularity challenges. (ii) Inexactto-Incomplete Transferability. Specialized inexact methods (e.g., MIL models) are tested on incomplete supervision tasks with sparse instance-level labels, examining whether temporal or structural inductive biases transfer to traditional tabular settings. (iii) Robustness under Inaccurate Supervision. Methods from different supervision scenarios are evaluated under label noise and corruption, examining whether designs optimized for specific supervision types maintain robustness when label quality degrades. By systematically implementing these evaluations across diverse modalities and different models, we enable rigorous assessment of how tabular foundation models and specialized WSAD algorithms perform when supervision is imperfect in quantity, granularity, or quality—challenging the assumption that methods should only be tested within their original problem settings.
3.2
Datasets and Models
Datasets. To comprehensively verify algorithm detection performance across diverse data types, WSADBench incorporates 61 datasets covering four distinct modalities: Tabular, Computer Vision (CV), Natural Language Processing (NLP), and Video. The datasets for Tabular, CV, and NLP are adopted from ADBench [23] (e.g., Pima, CIFAR-10, Agnews). For Video modality, we re-collected and reprocessed real-world surveillance datasets, including UCF-Crime[50], XD-Violence[55], TAD[59], and ShanghaiTech[33]. We strictly follow the “Standard / One-vs-Rest” protocol for converting multiclass datasets into AD tasks. Detailed statistics for all datasets are provided in Appendix A.1. Anomaly Detection Methods. We re-evaluate the landscape of effective anomaly detection by expanding the comparison beyond specialized WSAD algorithms to include Generalist Baselines, such as Tabular Foundation Models and Generic Tree-based/Deep-learningbased Classifiers. This broadens the competitive landscape to determine whether specialized AD designs are truly necessary or
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Figure 1: Overview of the WSADBench. It integrates datasets spanning diverse modalities and varied supervision scenarios into a comprehensive evaluation pipeline. The analysis uncovers the limits and trade-offs of existing WSAD algorithms. if general-purpose learners suffice. We benchmark a comprehensive collection of 36 algorithms. Table 2 presents the complete list of algorithms evaluated in our experiments. Note that this table serves as an inventory of benchmarked methods rather than a strict taxonomy. Methods are organized by supervision granularity: instance-level supervision and bag-level supervision (most originally proposed for video anomaly detection).
3.3
Experimental Protocols
Inconsistent evaluation protocols often hinder fair comparison. We therefore enforce strict experimental standardization covering the following aspects: (i) Unified Feature Representation. Pretrained features alter method rankings (detailed in Appendix E.2.2). We ensure all methods in a modality use the same features to isolate algorithmic impact. (ii) Standardized Ground Truth Alignment. Subtle differences in ground truth definitions (e.g., frame-level vs. clip-level labels) can lead to contradictory conclusions (see Appendix E.2.3). We unify these definitions to prevent annotation-induced variance. (iii) Reproducible Configurations and Controlled Variation. We evaluate methods with default hyperparameters and fixed data splits, while systematically mapping performance boundaries through controlled variation of supervision parameters (e.g., label ratios, noise rates). (iv) Rigorous Metrics and Statistical Validation. We assess performance using both AUC-ROC (ranking quality) and AUC-PR (imbalanced robustness), reporting mean and standard deviation across multiple seeds with statistical hypothesis testing to ensure significance.
4
Experiments and Analysis
We evaluate WSAD methods across three dimensions: label Scarcity, Granularity, and Quality. We compare specialized WSAD methods
against baselines under varying supervision conditions, then investigate their generalization, robustness, and limitations. Finally, we provide a global performance summary over all evaluated scenarios.
4.1
Basic WSAD Experiments
4.1.1 Instance-Level Supervision. We evaluate instance-level weak supervision on Tabular, CV, and NLP benchmarks under two labeling configurations: (1) Ratio-labeled Anomaly (RLA, 𝛾𝑙𝑎 ∈ {1%, 5%, 10%}), where a fixed percentage of anomalies are labeled; and (2) Number-labeled Anomaly (NLA, 𝑁𝑙𝑎 ∈ {1, 5, 10}), where a fixed count of anomalies are labeled regardless of dataset size. Table 3 presents RLA results; NLA results are provided in Appendix E.1, with both settings exhibiting consistent patterns. The narrow effective scope of weak supervision. Our results demonstrate that specialized WSAD methods outperform supervised baselines only under extreme label scarcity. For example, at 𝛾𝑙𝑎 = 1%, DevNet achieves AUCPR 0.555 compared to 0.521 for the best supervised baseline CatBoost. However, this advantage diminishes rapidly as labeled data increases. This pattern also occurs in CV tasks, where Dual-MGAN leads at 𝛾𝑙𝑎 = 1% (AUCPR 0.520 vs. TabM 0.461) but falls behind as labels increase. These findings reveal that WSAD methods’ specialized inductive biases, while valuable for sparse supervision, become redundant once sufficient labeled data enables general supervised approaches to learn effective decision boundaries. 4.1.2 Bag-Level Supervision. Video anomaly detection exemplifies bag-level supervision, where video-level labels (bags) are used to train models that predict frame-level anomalies (instances). Table 4 shows our evaluation of representative MIL-based VAD methods across multiple pretrained feature extractors to assess their robustness to different backbones, addressing inconsistent backbone
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 3: AUCPR comparison of different models under varying label ratios. Tabular
Text
𝛾𝑙𝑎 = 10%
𝛾𝑙𝑎 = 50%
𝛾𝑙𝑎 = 100%
𝛾𝑙𝑎 = 1%
𝛾𝑙𝑎 = 10%
𝛾𝑙𝑎 = 50%
𝛾𝑙𝑎 = 100%
𝛾𝑙𝑎 = 1%
𝛾𝑙𝑎 = 10%
𝛾𝑙𝑎 = 50%
𝛾𝑙𝑎 = 100%
Weakly Sup.
Image
𝛾𝑙𝑎 = 1% XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
0.481(6) 0.379(10) 0.408(8) 0.308(15) 0.314(14) 0.555(1) 0.524(3) 0.547(2) 0.378(12) 0.518(5) 0.378(11) 0.320(13)
0.688(3) 0.568(12) 0.602(9) 0.310(16) 0.316(15) 0.666(6) 0.691(2) 0.676(4) 0.570(11) 0.603(8) 0.415(13) 0.331(14)
0.816(2) 0.752(8) 0.739(9) 0.321(16) 0.335(15) 0.713(11) 0.772(5) 0.721(10) 0.762(7) 0.647(12) 0.416(13) 0.386(14)
0.864(2) 0.813(7) 0.779(9) 0.351(16) 0.382(15) 0.722(11) 0.786(8) 0.729(10) 0.826(6) 0.671(12) 0.489(14) 0.517(13)
0.388(7) 0.268(16) 0.288(11) 0.332(8) 0.277(15) 0.459(4) 0.449(5) 0.468(2) 0.285(13) 0.520(1) 0.292(10) 0.321(9)
0.633(7) 0.474(12) 0.592(9) 0.320(14) 0.279(16) 0.694(1) 0.666(4) 0.688(3) 0.548(10) 0.689(2) 0.292(15) 0.332(13)
0.759(7) 0.722(11) 0.773(4) 0.330(15) 0.285(16) 0.782(2) 0.790(1) 0.781(3) 0.747(8) 0.764(6) 0.383(14) 0.404(13)
0.798(8) 0.826(1) 0.816(3) 0.366(15) 0.311(16) 0.798(7) 0.812(4) 0.795(9) 0.820(2) 0.776(11) 0.449(14) 0.592(13)
0.166(7) 0.105(14) 0.134(11) 0.067(16) 0.110(13) 0.272(2) 0.221(5) 0.261(3) 0.164(8) 0.253(4) 0.087(15) 0.172(6)
0.366(6) 0.225(12) 0.344(9) 0.078(16) 0.107(14) 0.512(2) 0.461(5) 0.513(1) 0.356(8) 0.498(3) 0.104(15) 0.176(13)
0.571(7) 0.503(11) 0.618(5) 0.071(16) 0.114(15) 0.661(2) 0.646(4) 0.668(1) 0.582(6) 0.647(3) 0.135(14) 0.194(13)
0.669(9) 0.726(2) 0.741(1) 0.073(16) 0.136(15) 0.705(4) 0.717(3) 0.702(6) 0.703(5) 0.685(8) 0.144(14) 0.224(13)
Sup.
Model
XGBoost CatBoost TabM TabR-S
0.241(16) 0.521(4) 0.468(7) 0.405(9)
0.622(7) 0.715(1) 0.675(5) 0.596(10)
0.795(3) 0.842(1) 0.787(4) 0.768(6)
0.858(3) 0.883(1) 0.830(5) 0.851(4)
0.283(14) 0.448(6) 0.461(3) 0.287(12)
0.596(8) 0.635(6) 0.660(5) 0.488(11)
0.744(9) 0.764(5) 0.741(10) 0.696(12)
0.795(10) 0.802(6) 0.762(12) 0.808(5)
0.151(9) 0.143(10) 0.278(1) 0.130(12)
0.361(7) 0.317(10) 0.478(4) 0.249(11)
0.557(9) 0.568(8) 0.520(10) 0.483(12)
0.667(10) 0.686(7) 0.552(12) 0.660(11)
0.365
Median Unsup.
0.360
usage in prior comparative studies. Detailed experimental settings and results are provided in Appendix E.2. Table 4: AUCPR comparison of VAD models across various pretraining methods and segmentations. i3d
x3d
sf50
mvit
sf
Mean
0.430 (1) 0.392 (2) 0.258 (7) 0.386 (3) 0.340 (5) 0.344 (4) 0.304 (6)
0.434 (1) 0.377 (3) 0.279 (7) 0.397 (2) 0.339 (6) 0.343 (5) 0.364 (4)
0.436 (1) 0.403 (3) 0.337 (6) 0.418 (2) 0.377 (4) 0.357 (5) 0.307 (7)
0.441 (1) 0.399 (2) 0.308 (7) 0.394 (3) 0.363 (4) 0.354 (5) 0.351 (6)
0.445 (1) 0.412 (3) 0.337 (7) 0.427 (2) 0.376 (4) 0.350 (6) 0.354 (5)
0.437 (1) 0.397 (2) 0.304 (7) 0.404 (3) 0.359 (5) 0.350 (6) 0.336 (4)
AR-Net Sultani MGFN RTFM UR-DMU VadClip GCN-Anomaly
0.429 (1) 0.377 (3) 0.297 (7) 0.380 (2) 0.353 (4) 0.348 (5) 0.303 (6)
0.433 (1) 0.398 (2) 0.291 (7) 0.379 (4) 0.350 (6) 0.351 (5) 0.381 (3)
0.436 (1) 0.400 (3) 0.308 (7) 0.408 (2) 0.387 (4) 0.359 (5) 0.349 (6)
0.442 (1) 0.397 (2) 0.333 (7) 0.389 (3) 0.371 (6) 0.372 (5) 0.375 (4)
0.445 (1) 0.412 (3) 0.346 (7) 0.423 (2) 0.374 (4) 0.350 (6) 0.373 (5)
0.437 (1) 0.397 (3) 0.315 (7) 0.396 (2) 0.367 (4) 0.356 (5) 0.356 (6)
Overall Mean
0.352
0.361
0.374
0.377
0.387
0.370
AR-Net Sultani MGFN RTFM UR-DMU VadClip GCN-Anomaly 200 Segments
Simpler methods prevail under standardized evaluation. A notable finding is that AR-Net [54], one of the earliest MIL-based methods, ranks first across all backbone–segmentation combinations, with mean AUCPR 0.437. This strong performance stems from a remarkably simple but effective design: adding the center loss to the foundational margin-ranking framework of Sultani [50], which encourages normal-snippet features to cluster tightly. In contrast, more recent and complex methods—such as MGFN (mean AUCPR 0.309, rank 7) and UR-DMU (0.363, rank 4)—perform notably worse and exhibit higher variance across backbones. We attribute this partly to evaluation protocol: most prior methods report the best epoch selected on the test set, whereas WSADBench adopts a strict last-epoch protocol without peeking at test labels. This change particularly penalizes methods with unstable convergence, as they can no longer identify and select the best-performing checkpoints.
4.2
Tabular
Label Ratio Setting AUCPR
32 Segments
across varying supervision levels, modalities, and dataset characteristics to assess the necessity of specialized WSAD designs. To adapt these models, training data formatting follows a PositiveUnlabeled (PU) learning proxy. This setup constructs the Context Set (𝑋𝑐𝑡𝑥 , 𝑌𝑐𝑡𝑥 ) by assigning 𝑦 = 1 to labeled anomalies and pseudolabeling unlabeled instances as 𝑦 = 0. During inference, in-context learning predicts anomaly probabilities for test instances 𝑋𝑞𝑢𝑒𝑟 𝑦 : 𝑌ˆ𝑞𝑢𝑒𝑟 𝑦 = 𝑓 (𝑋𝑞𝑢𝑒𝑟 𝑦 | 𝑋𝑐𝑡𝑥 , 𝑌𝑐𝑡𝑥 ).
Further Experiments and Analysis
4.2.1 Tabular Foundation Models: Challenging the Necessity of Specialized WSAD Methods. We evaluate tabular foundation models
0.6 0.4 1%
5%
10%
Label Count Setting AUCPR
Model
0.093
0.6 0.4 1
5
10
Image
0.7 0.6 0.5 0.4 0.3
0.6 0.5 0.4 0.3 0.2
Repr. Learning GAN-based Score Learning
Label Ratio Setting
1%
5%
10%
Text
0.5 0.4 0.3 0.2 0.1
Label Count Setting
Label Ratio Setting Models
1%
5%
10%
Label Count Setting 0.4 0.3 0.2 0.1
1
5
10
Data Aug. Pseudo-Labeling GBDT
1
5 Deep(Sup.) Found. Model Diffusion DAE
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL XGBoost CatBoost TabM TabR-S TabPFN LimiX DDAE
10
Figure 2: Tabular Foundation model performance (AUCPR) for different modalities under varying labeled anomaly ratios (upper) and counts (lower). Tabular Foundation models dominate specialized WSAD across supervision levels and modalities. As shown in Figure 2, tabular foundation models (e.g., LimiX, TabPFN) consistently match or exceed specialized WSAD methods across the majority of supervision settings. Specialized WSAD methods remain competitive only when labels are extremely scarce (𝑁𝑙𝑎 = 1). Once supervision exceeds this minimal threshold, tabular foundation models achieve substantially higher performance (e.g., LimiX 0.641 vs. PReNet 0.578 w.r.t. 𝑁𝑙𝑎 = 5 on tabular data). This finding challenges the traditional view that specialized anomaly detection designs are essential for weak supervision. In contrast, general-purpose pre-trained representations often prove sufficient or even superior.
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
4.2.2 The Value of Unlabeled Data. We investigate the value of unlabeled data across varying amounts of labeled anomalies, revealing its dependency on both supervision levels and model architecture.
1000
0.8
0.6
0.6
0.4
0.4
0.2
0.2 1
10
Nla
20
AUCPR
Repr. Learning GAN-based
50
50
Nu
200
1000
Fixed Nu = 1000
1.0
0.8
0.0
20
0.0
1
Score Learning Data Aug.
10
Pseudo-Labeling GBDT
Nla
20
Nu 50
20 1
10
20
1.0
1000 200
50
Nu 50
N la
REPEN
Nu 50
50
Deep(Sup.) Found. Model
20 1
(a) Marginal effect
10
20
50
0.8
N la
0.6
TabPFN
1.0 0.8 0.6 0.4 0.2 0.0
1000 200
20 1
10
Gradient Magnitude
200
Fixed Nu = 20
1.0
AUCPR
Nu
20
0.4 0.2
1.0 0.8 0.6 0.4 0.2 0.0
0.0
1000 200
50
Nu 50
N la
20 1
10
20
50
N la
(b) Joint effect
Figure 3: AUCPR results on classical datasets under varying labeled (𝑁𝑎 ) and unlabeled (𝑁𝑢 ) data amounts. learning approaches (e.g., DeepSAD, REPEN) require a minimum threshold of unlabeled samples to define normal patterns; insufficient unlabeled data can cause performance degradation when labeled anomalies are added, due to instability of learned features or overfitting to sparse anomalies. Surprisingly, despite pre-training on vast external corpora, tabular foundation models (e.g., LimiX, TabPFN) display scaling behaviors similar to specialized WSAD methods with respect to unlabeled data (Figure 3b): performance gains plateau as labeled data increases, while unlabeled data remains valuable when labeled examples are extremely scarce. This confirms that tabular foundation models inherently possess high label efficiency, enabling them to effectively maximize the utility of minimal supervision signals comparable to specialized methods. 4.2.3 Sensitivity to Label Noise. We investigate how label noise affects model performance on tabular datasets, revealing that noise impact varies systematically by corruption type and model architecture, and evaluate whether noise cleaning methods can effectively mitigate these effects. Label noise is simulated by injecting errors into the training set: Flip-Normal (Normal → Anomaly), FlipAbnormal (Anomaly → Normal), and Double-Noise (Both). We introduce label noise by setting specific error ratios, as real-world annotation errors affect both classes at similar rates. Fixed Flip abnormal ratio = 0.0
Fixed Flip abnormal ratio = 0.0
20
0.8
0
0.7
10 20 30
0.6 0.5 0.4
40 50
Models
0.9
10
0.3 0
0.01 0.05 0.1 0.25 flip normal ratio Repr. Learning GAN-based Score Learning
0.5
0.01 0.05 0.1 0.25 0.5 flip normal ratio Data Aug. Deep(Sup.) Found. Model Pseudo-Labeling Diffusion DAE GBDT
0
10
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX DDAE
Fixed Flip normal ratio = 0.0 0.9
0
0.8
5
0.7
10 15 25
0.4
30 35
0.3 0
0.01 0.05 0.1 0.25 flip abnormal ratio
Repr. Learning GAN-based Score Learning
FNR
0.01 0.05 0.1 0.25 flip abnormal ratio
0.5
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX DDAE
Deep(Sup.) Found. Model Diffusion DAE
0.01 0.05 0.144 0.1
FAR0.25
0.5 0.5 0.25 0.1 0.05 0.01
FNR
LimiX 1.0 0.8 0.6 0.4 0.2
1.0 0.8 0.6 0.4
AUCPR
1.0 0.8 0.6 0.4 0.2
0.01 0.277 0.05 0.1
0.5 0.5 0.25 0.1 0.05 0.01
0
Data Aug. Pseudo-Labeling GBDT
TabPFN
AUCPR
0.01 0.05 0.177 0.1
0.5
Models
(b) Flip abnormal
1.0 0.8 0.6 0.4 0.2
FAR0.25
0.6 0.5
20
REPEN
AUCPR
1.0 0.8 0.6 0.4 0.2
Fixed Flip normal ratio = 0.0
5
(a) Flip normal DeepSAD
AUCPR
Unlabeled data utility depends on label availability. Contrary to the expectation that unlabeled data compensates for supervision scarcity, Figure 3a reveals a more complex relationship. Under oneshot scenario (𝑁𝑙𝑎 = 1), incorporating massive unlabeled data yields limited performance gains, with AUCPR improvements of less than 0.05 for most models when scaling 𝑁𝑢 from 20 to 1000. This occurs because the supervision signal remains too weak to anchor effective learning, rendering distributional information from unlabeled samples difficult to exploit. Conversely, when labeled anomalies are abundant (𝑁𝑙𝑎 >20), unlabeled data becomes redundant, contributing AUCPR marginal gains of less than 0.1, as the decision boundary is sufficiently defined by labeled samples alone. Maximum utility occurs in moderate supervision regimes (𝑁𝑙𝑎 ∈ [10, 50]), where labeled samples provide sufficient anchors while unlabeled data effectively expands the normal manifold, yielding AUCPR improvements of up to 0.4 . These patterns highlight that the effectiveness of unlabeled data is highly dependent on the availability of labeled anomalies. Method-specific unlabeled data dependencies. Different methods exhibit distinct dependencies on unlabeled data. Representation
50
1000 200
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
1.0 0.8 0.6 0.4 0.2 0.0
0.2
0.01 0.05 0.141 0.1
FAR0.25
0.5 0.5 0.25 0.1 0.05 0.01
FNR
AUCPR Performance
Note: ∗ : 𝑝 < 0.05,∗∗ : 𝑝 < 0.01,∗∗∗ : 𝑝 < 0.001. Eff.:Effective, Feat:Feature, Sim:Similarity, Pts: Points.
20
Models
AUCPR
−0.48∗∗∗ 0.44∗∗∗ 0.33∗∗∗ −0.73∗∗∗ 0.24∗∗ 0.36∗∗∗ −0.12 −0.62∗∗∗ 0.62∗∗∗ −0.03
0.0
AUCPR (%)
LimiX - TabPFN
−0.37∗∗∗ 0.24∗∗ 0.02 −0.10 0.43∗∗∗ −0.61∗∗∗ 0.48∗∗∗ 0.18 −0.50∗∗∗ 0.41∗∗∗
0.2
0.0
AUCPR
TabPFN - DevNet
−0.44∗∗∗ 0.31∗∗∗ 0.12 −0.28∗∗ 0.44∗∗∗ −0.49∗∗∗ 0.40∗∗∗ −0.01 −0.31∗∗∗ 0.30∗∗∗
0.4
0.2
AUCPR (%)
LimiX - DevNet
0.6
0.4
AUCPR
Dimensionality Sparsity Ratio Avg.Feat.Sim. Eff. Rank Eff. Rank Ratio Fisher’s Ratio (F1) Borderline Pts. (N1) Intra/Inter Ratio (N2) RF’s F1 Local/Global Outlier Ratio
0.8
0.6
DeepSAD
1.0 0.8 0.6 0.4 0.2 0.0
AUCPR
Meta-Feature
0.8
DevNet
Fixed Nla = 50
1.0
AUCPR
Table 5: Spearman correlation (𝜌) between meta-features and pairwise AUCPR gaps across all datasets (𝛾𝑙𝑎 = 100%).
Fixed Nla = 1
1.0
AUCPR
Task characteristics define the boundaries of tabular foundation model superiority. Given the dominance of tabular foundation models on tabular data, we investigate model performance across modalities through a Spearman correlation analysis between meta-features and the performance gap, as detailed in Table 5. The explanation of all meta-features can be found in Table 20. First, Dimensionality shows a significant negative correlation with the performance advantage of tabular foundation models over specialized methods (𝜌 = −0.44). This suggests that while tabular foundation models excel on lower-dimensional tabular data, their advantage diminishes in high-dimensional feature spaces (e.g., embeddings from ViT or RoBERTa used in CV and NLP). Second, Linear Separability (Fisher’s Ratio) also negatively correlates with the tabular foundation model advantage. Specialized methods like DevNet are more efficient when boundaries are distinct, whereas tabular foundation models excel at modeling complex, non-linear boundaries where traditional methods struggle. These factors—dimensionality and separability—jointly characterize the conditions where tabular foundation models dominate versus where specialized methods remain competitive.
FAR0.25
0.5 0.5 0.25 0.1 0.05 0.01
FNR
(c) Stability landscape
Figure 4: Model sensitivity to label noise: AUCPR degradation under asymmetric noise types.
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Asymmetric sensitivity to label noise types. Normal label noise is more destructive than abnormal label noise. As illustrated in Figure. 4a and 4b, models generally exhibit a steeper performance decline under Flip-Normal noise (FNR) compared to Flip-Abnormal noise (FAR). This asymmetry stems from the anomaly-aware nature of WSAD methods. FNR (Normal → Anomaly) introduces false positive anomalies that directly corrupt the anomaly pattern WSAD models explicitly explore during training, severely distorting learned abnormal representations. In contrast, FAR (Anomaly → Normal) merely reduces labeled anomalies, which WSAD methods can tolerate due to their inherent robustness to unlabeled data noise, treating mislabeled anomalies as part of the unlabeled pool. Inconsistent robustness across models. As illustrated in Figure 4c, tabular foundation models demonstrate contrasting noise tolerance behaviors. TabPFN forms a high, flat plateau indicating broad noise tolerance with “cliff effect” only under extreme dualnoise conditions, while LimiX exhibits sharp surface decline at lower noise thresholds despite superior clean-data performance. Specialized WSAD methods (DeepSAD and REPEN) exhibit moderate noise robustness, degrading more gradually than LimiX but less robustly than TabPFN.
Using PCA-derived feature-space distances from the normal centroid (Figure 5a), we partition anomaly classes into Known (A𝑘𝑛𝑜𝑤𝑛 , used for training) and Unknown (A𝑢𝑛𝑘𝑛𝑜𝑤𝑛 , reserved for OOD testing), where Known anomalies are sampled at label ratio 𝑟𝑙𝑎 . This yields three distinct settings: Setting I (ID Far, OOD Near) assesses whether models trained on obvious anomalies can generalize to detect subtle threats; Setting II (ID Near, OOD Far) examines if fitting hard anomalies sacrifices the performance on distinct, easy outliers; and Setting III (ID Near, OOD Near) evaluates the boundary precision when all anomalies are near-normal.
Model SOEL-NTL TabM FTTransformer DDAE GANomaly DeepSAD RoSAS REPEN XGBOD DevNet CatBoost TabPFN
Flip-Normal
Flip-Abnormal
Average
9.61% 8.17% 2.13% 0.29% 0.49% -0.15% -0.24% -0.68% -1.00% -1.71% -1.56% -3.16%
0.87% 0.80% 0.41% 0.67% 0.25% 0.08% 0.15% -0.08% -0.32% 0.02% -0.46% -0.43%
5.24% 4.49% 1.27% 0.48% 0.37% -0.04% -0.04% -0.38% -0.66% -0.84% -1.01% -1.80%
Effectiveness of data cleaning method. We further evaluate whether automated data cleaning tools like Cleanlab [36] can mitigate label noise impact. As detailed in Table 6, external filtering may significantly benefit models lacking inherent robustness such as SOEL-NTL, TabM, and FTTransformer, which are the most sensitive to flip-normal noise (as shown in Figure 4a). Specifically, SOEL-NTL and TabM achieve substantial performance recovery under flip-normal noise, with AUCPR improvements of 9.61% and 8.17%, respectively. Conversely, this preprocessing proves detrimental to intrinsically robust architectures, as TabPFN and DevNet suffer performance degradations of 3.16% and 1.71% under the same conditions. This discrepancy suggests that cleaning helps sensitive models by correcting mislabeled samples but harms robust models by removing valuable difficult samples needed for generalization. 4.2.4 Out-of-Distribution Generalization Analysis. We evaluate robustness under distribution shift on tabular datasets by controlling the distance between training (ID) and testing (OOD) anomalies.
PCA Component 2
Setting I: ID Far, OOD Near
OOD
2
ID
dOOD
Setting II: ID Near, OOD Far
Setting III: ID Near, OOD Near
ID OOD
dID
dID dOOD
0
2 ID
dID
4 6
5.0
dOOD
OOD
2.5 0.0 2.5 5.0 PCA Component 1 Unselected Selected Normal
5.0
Classes bent color flip good scratch
2.5 0.0 2.5 5.0 PCA Component 1 Selected Known (ID) Selected Unknown (OOD)
5.0 2.5 0.0 2.5 5.0 PCA Component 1 Normal Center OOD Center ID Center
(a) PCA visualization of feature distributions. Setting I: Setting II: Setting III: ID Far, OOD Near ID Near, OOD Far ID Near, OOD Near
1.0 0.8
AUCPR
Table 6: Average absolute performance improvement (Δ AUCPR %) after Cleanlab denoising.
4
0.6 0.4 0.2 0.0
0
25
50 75 OOD Rate (%)
100
Repr. Learning GAN-based Score Learning
0
25
50 75 OOD Rate (%)
Data Aug. Pseudo-Labeling GBDT
100
0
25
50 75 OOD Rate (%)
Deep(Sup.) Found. Model Diffusion DAE
100
Model
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX DDAE Median Unsup.
(b) AUCPR performance (𝛾𝑙𝑎 = 10%).
Figure 5: Analysis of incomplete OOD across 3 Settings. WSAD methods struggle to generalize beyond known anomaly types. In Scenario I of Figure 5b, where ID anomalies are far from normal, WSAD methods (DevNet, PReNet, FEAWAD) achieve the highest AUCPR at OOD rate = 0%, surpassing even tabular foundation models. However, this advantage is narrow: performance drops sharply as the OOD rate increases, falling below tabular foundation models and GBDTs at moderate OOD rates. This performance reversal reveals an ID-specific optimization that collapses when novel anomaly types appear. In contrast, Settings II and III show no such reversal: When ID anomalies are near-normal, harder training conditions prevent WSAD methods from surpassing tabular foundation models, leaving them consistently below even at OOD rate = 0%. Non-WSAD methods exhibit varied OOD tolerance. Tabular foundation models (e.g., LimiX, TabPFN) and gradient boosting (e.g., CatBoost) degrade less severely than WSAD methods as OOD rates increase. They remain competitive even when unknown anomalies dominate (OOD rate → 100%). This tolerance likely stems from pre-trained or ensemble strategies that generalize beyond specific types. In contrast, reconstruction methods (e.g., GANomaly and AA-BiGAN) maintain stable performance across OOD rates by modeling normal patterns, but their absolute performance remains inferior, typically at or below unsupervised baselines. An exception occurs in Scenario II: when OOD anomalies are far from normal, GANomaly’s AUCPR rises with OOD rate, as the normal-centric objective naturally separates distant anomalies. Additional results for all datasets are shown in Figure 11 (Appendix E.3).
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
DevNet
AUCPR: 0.564
LimiX
AUCPR: 0.763
AUCPR: 0.968 0.6
TabPFN
LimiX
CatBoost
FEAWAD
XGBOD
0 0.5
10 20
10 0 10 PCA Component 1
20 Normal
10 0 10 PCA Component 1 ID Anom
20
10 0 10 PCA Component 1
TabM
TabR-S XGBoost PReNet RoSAS REPEN DeepSAD
AUCPR
PCA Component 2
GANomaly 10
Xu Yao et al.
DevNet
Dual-MGAN
0.4
OOD Anom DDAE
Figure 6: Anomaly score decision boundaries on Metal_nut (Scenario I, 𝛾𝑙𝑎 = 10%, OOD rate=50%). Redder regions indicate higher anomaly scores.
0.3
SOEL-NTL GANomaly AA-BiGAN Bubble Size 0.20
0.21
0.22
Log(Time)
0.23
0.24 Pseudo-Labeling
Standard Deviation
Figure 6 visualizes the output anomaly score for three representative models on Metal_nut (Scenario I). GANomaly detects both ID and OOD anomalies but assigns uniformly low scores with excessive false positives, as normal-only reconstruction cannot leverage supervision effectively. DevNet assigns high scores to ID anomalies but near-zero scores to OOD samples. Its decision regions narrowly surround known anomaly types, failing to generalize to nearby OOD samples. Conversely, LimiX constructs compact boundaries that cover both ID and OOD anomalies, achieving superior generalization without fragmenting the decision space. This reveals that WSAD methods memorize specific instances rather than learning generalizable patterns—a form of ID overfitting that prevents generalization even when OOD anomalies are nearby. 4.2.5 Can Methods Transfer Across Supervision Types? We evaluate whether methods designed for one supervision type transfer to another by testing incomplete supervision methods on the inexact scenario (video MIL tasks) and vice versa, examining how domainspecific inductive biases affect generalization between scenarios. Table 7: Supervision type transferability evaluation. Paradigm
Model
Category
Video MIL
Tabular MIL
MIL
AR-Net Sultani MGFN RTFM UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Vanilla MIL Magnitude MIL Magnitude MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.441 (3) 0.398 (7) 0.308 (12) 0.395 (8) 0.363 (9) 0.354 (10) 0.351 (11)
0.252 (2) 0.127 (8)
Broadcast
IForest CatBoost DevNet FEAWAD DeepSAD TabPFN
Isolation-based GBDT Score Learning Reconstruction Score Learning Found. Model
0.143 (13) 0.428 (5) 0.437 (4) 0.452 (2) 0.453 (1) 0.408 (6)
0.174 (7) 0.230 (6) 0.235 (5) 0.241 (3) 0.240 (4) 0.260 (1)
Transferability across supervision types reveals domainspecific constraints. Label broadcasting transforms inexact supervision into an inaccurate supervision problem. By treating video bags as i.i.d. tabular instances, bag-level labels propagate to instancelevel, introducing label noise as the primary challenge. Surprisingly, Table 7 illustrates that methods designed for incomplete supervision—despite lacking temporal modeling or MIL pooling—even outperform specialized inexact approaches on video MIL tasks: DeepSAD achieves mean AUCPR 0.453, surpassing the specialized MIL model AR-Net’s AUCPR 0.441. Conversely, specialized inexact methods (e.g., Sultani, RTFM) show limited transferability to static tabular MIL, lagging significantly behind tabular baselines despite
Repr. Learning
Data Aug.
Deep(Sup.)
GAN-based
Found. Model
Score Learning
GBDT
Diffusion DAE
Figure 7: Performance across all scenarios. Weakly Supervised Models
Supervised Models
Scarce Label
Scarce Label
OOD
Top 1
Top 1
Top 5
Top 5
Coarse Label
Top 10
Top 15
Top 15
Double Noise
Normal Noise
Double Noise
Abnormal Noise AA-BiGAN SOEL-NTL GANomaly
Coarse Label
OOD
Top 10
DDAE Dual-MGAN DevNet
Normal Noise
Abnormal Noise REPEN PReNet RoSAS
DeepSAD FEAWAD XGBOD
TabM XGBoost
TabR-S LimiX
CatBoost
TabPFN
Figure 8: Model ranking radar chart. competitive performance on video. Domain-specific inductive biases improve performance within their target modality but limit generalizability across scenarios. This highlights the importance of evaluating WSAD methods through a unified taxonomy.
4.3
Performance Summary
To conclude our analysis, we present a global overview of algorithmic performance across all settings. Figure 7 visualizes the trade-off among performance, stability, and efficiency, where models in the top-left region with smaller bubbles achieve the best balance. TabPFN dominates this space with the highest mean AUCPR (∼0.65), the lowest variability, and the smallest training cost. Most WSAD methods, regardless of category, cluster tightly in a narrow band, suggesting that different algorithmic designs yield only marginal differentiation under aggregated evaluation. GAN-based approaches fall to the bottom of the chart despite incurring the largest computational overhead. Figure 8 presents model rankings (based on mean AUCPR) across eight distinct scenarios spanning varying labeled anomaly ratios (100%, 50%, 1%), noise conditions (FNR, FAR, Double Noise), and three data modalities (Tabular, Inexact, VAD). Crucially, the ranking for each scenario-modality combination is calculated independently within its corresponding modality, without any cross-modality ranking aggregation. This cross-scenario comparison confirms that specialized WSAD algorithms are inferior to tabular foundation models in most scenarios, with no single method dominating all settings, underscoring the need for detailed per-scenario analysis and scenario-specific algorithm adaptation/selection.
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
5
Conclusion
This work presents WSADBench, the first unified evaluation framework for WSAD that enables fair comparison across diverse supervision settings and algorithms. We evaluate 36 algorithms on 61 datasets spanning tabular, image, text, and video modalities, conducting over 700K experiments. Results show that specialized WSAD algorithms excel only under extreme label scarcity; as labeled data increases or anomaly types shift, tabular foundation models and general classification methods consistently dominate. We also reveal strong correlations across supervision settings and asymmetric model sensitivity to different noise types. Our findings suggest two directions for future research. First, WSAD can benefit through either leveraging general-purpose foundation models to improve representation quality in downstream AD-specific models, or developing AD foundation models that capture universal abnormal patterns across domains. Second, algorithmic challenges such as poor OOD generalization and asymmetric noise sensitivity highlight the need for algorithms that learn robust, generalizable abnormal patterns. Going forward, We plan to expand WSADBench with additional modalities and continuously update the benchmark with emerging algorithms.
Acknowledgments This work was supported by the National Natural Science Foundation of China (Nos. 72271151, 72342009, 72442024, 72172085) and Ant Group. We acknowledge AI tools for assisting with LaTeX formatting, English grammar polishing, and literature search. Literature sorting and verification, as well as review of all AI-assisted content, were conducted manually.
References [1] Moshira Abdalla, Sajid Javed, Muaz Al Radi, Anwaar Ulhaq, and Naoufel Werghi. 2025. Video anomaly detection in 10 years: A survey and outlook. Neural Computing and Applications 37, 32 (2025), 26321–26364. [2] Andra Acsintoae, Andrei Florescu, Mariana-Iuliana Georgescu, Tudor Mare, Paul Sumedrea, Radu Tudor Ionescu, Fahad Shahbaz Khan, and Mubarak Shah. 2022. UBnormal: New Benchmark for Supervised Open-Set Video Anomaly Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 20125–20135. [3] Charu C. Aggarwal. 2013. Outlier Analysis. Springer. doi:10.1007/978-1-46146396-2 [4] Samet Akcay, Amir Atapour-Abarghouei, and Toby P Breckon. 2018. Ganomaly: Semi-supervised anomaly detection via adversarial training. In Asian conference on computer vision. Springer, 622–637. [5] Fabrizio Angiulli and Clara Pizzuti. 2002. Fast outlier detection in high dimensional spaces. In European conference on principles of data mining and knowledge discovery. Springer, 15–27. [6] Jinan Bao, Hanshi Sun, Hanqiu Deng, Yinsheng Brennan He, Zhaoxiang Zhang, and Xingyu Li. 2024. BMAD: Benchmarks for Medical Anomaly Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4042–4053. [7] Mohamed Bouadi, Pratinav Seth, Aditya Tanna, and Vinay Kumar Sankarapu. 2025. Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning. arXiv preprint arXiv:2511.02818 (2025). [8] Markus M Breunig, Hans-Peter Kriegel, Raymond T Ng, and Jörg Sander. 2000. LOF: identifying density-based local outliers. In Proceedings of the 2000 ACM SIGMOD international conference on Management of data. 93–104. [9] Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 6299–6308. [10] Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining. 785–794. [11] Yingxian Chen, Zhengzhe Liu, Baoheng Zhang, Wilton Fok, Xiaojuan Qi, and Yik-Chung Wu. 2023. Mgfn: Magnitude-contrastive glance-and-focus network for
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
weakly-supervised video anomaly detection. In Proceedings of the AAAI conference on artificial intelligence, Vol. 37. 387–395. [12] Choubo Ding, Guansong Pang, and Chunhua Shen. 2022. Catching both gray and black swans: Open-set supervised anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 7388–7398. [13] Yutao Dong, Qing Li, Richard O Sinnott, Yong Jiang, and Shutao Xia. 2021. ISP self-operated BGP anomaly detection based on weakly supervised learning. In 2021 IEEE 29th International Conference on Network Protocols (ICNP). IEEE, 1–11. [14] Marius Dragoi, Elena Burceanu, Emanuela Haller, Andrei Manolache, and Florin Brad. 2022. Anoshift: A distribution shift benchmark for unsupervised anomaly detection. Advances in Neural Information Processing Systems 35 (2022), 32854– 32867. [15] Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision. 6202–6211. [16] Adam Goodge, Bryan Hooi, See-Kiong Ng, and Wee Siong Ng. 2022. Lunar: Unifying local outlier detection methods via graph neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 6737–6745. [17] Yury Gorishniy, Akim Kotelnikov, and Artem Babenko. 2025. Tabm: Advancing tabular deep learning with parameter-efficient ensembling. In International Conference on Learning Representations, Vol. 2025. 77899–77935. [18] Yury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii, Akim Kotelnikov, and Artem Babenko. 2024. Tabr: Tabular deep learning meets nearest neighbors. In International Conference on Learning Representations, Vol. 2024. 18209–18249. [19] Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021. Revisiting deep learning models for tabular data. Advances in neural information processing systems 34 (2021), 18932–18943. [20] Sachin Goyal, Aditi Raghunathan, Moksh Jain, Harsha Vardhan Simhadri, and Prateek Jain. 2020. DROCC: Deep robust one-class classification. In International conference on machine learning. PMLR, 3711–3721. [21] Léo Grinsztajn, Klemens Flöge, Oscar Key, Felix Birkel, Philipp Jund, Brendan Roof, Benjamin Jäger, Dominik Safaric, Simone Alessi, Adrian Hayler, et al. 2025. Tabpfn-2.5: Advancing the state of the art in tabular foundation models. arXiv preprint arXiv:2511.08667 (2025). [22] Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Conference on Machine Learning. PMLR, 1321–1330. [23] Songqiao Han, Xiyang Hu, Hailiang Huang, Minqi Jiang, and Yue Zhao. 2022. ADBench: Anomaly Detection Benchmark. In NeurIPS. [24] Zengyou He, Xiaofei Xu, and Shengchun Deng. 2003. Discovering cluster-based local outliers. Pattern recognition letters 24, 9-10 (2003), 1641–1650. [25] Tin Kam Ho and Mitra Basu. 2002. Complexity measures of supervised classification problems. IEEE transactions on pattern analysis and machine intelligence 24, 3 (2002), 289–300. [26] Xi Jiang, Jian Li, Hanqiu Deng, Yong Liu, Bin-Bin Gao, Yifeng Zhou, Jialin Li, Chengjie Wang, and Feng Zheng. 2025. Mmad: A comprehensive benchmark for multimodal large language models in industrial anomaly detection. In International conference on learning representations. 87273–87295. [27] Diederik P. Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations. [28] Elizaveta Levina and Peter Bickel. 2004. Maximum likelihood estimation of intrinsic dimension. Advances in neural information processing systems 17 (2004). [29] Yuangang Li, Jiaqi Li, Zhuo Xiao, Tiankai Yang, Yi Nian, Xiyang Hu, and Yue Zhao. 2025. NLP-ADBench: NLP Anomaly Detection Benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2025. 2464–2474. doi:10. 18653/v1/2025.findings-emnlp.133 [30] Zhe Li, Chunhua Sun, et al. 2022. Dual-MGAN: An Efficient Approach for Semisupervised Outlier Detection with Few Identified Anomalies. TKDD (2022). [31] Zheng Li, Yue Zhao, Xiyang Hu, Nicola Botta, Cezar Ionescu, and George H Chen. 2022. Ecod: Unsupervised outlier detection using empirical cumulative distribution functions. IEEE Transactions on Knowledge and Data Engineering 35, 12 (2022), 12181–12193. [32] Fei Tony Liu, Kai Ming Ting, and Zhi-Hua Zhou. 2008. Isolation forest. In 2008 eighth ieee international conference on data mining. IEEE, 413–422. [33] Wen Liu, Weixin Luo, Dongze Lian, and Shenghua Gao. 2018. Future frame prediction for anomaly detection–a new baseline. In Proceedings of the IEEE conference on computer vision and pattern recognition. 6536–6545. [34] Junwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach, Jesse C. Cresswell, Keyvan Golestan, Guangwei Yu, Anthony L. Caterini, and Maksims Volkovs. 2025. TabDPT: Scaling Tabular Foundation Models on Real Data. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. [35] Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015. Obtaining Well Calibrated Probabilities Using Bayesian Binning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 29. [36] Curtis Northcutt, Lu Jiang, and Isaac Chuang. 2021. Confident learning: Estimating uncertainty in dataset labels. Journal of Artificial Intelligence Research 70 (2021), 1373–1411.
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
[37] Guansong Pang, Longbing Cao, Ling Chen, and Huan Liu. 2018. Learning representations of ultrahigh-dimensional data for random distance-based outlier detection. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 2041–2050. [38] Guansong Pang, Chunhua Shen, Longbing Cao, and Anton Van Den Hengel. 2021. Deep learning for anomaly detection: A review. ACM computing surveys (CSUR) 54, 2 (2021), 1–38. [39] Guansong Pang, Chunhua Shen, Huidong Jin, and Anton Van Den Hengel. 2023. Deep weakly-supervised anomaly detection. In Proceedings of the 29th ACM SIGKDD conference on knowledge discovery and data mining. 1795–1807. [40] Guansong Pang, Chunhua Shen, and Anton Van Den Hengel. 2019. Deep anomaly detection with deviation networks. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. 353–362. [41] Lorenzo Perini, Vincent Vercruyssen, and Jesse Davis. 2023. Learning from positive and unlabeled multi-instance bags in anomaly detection. In Proceedings of the 29th ACM SIGKDD conference on knowledge discovery and data mining. 1897–1906. [42] Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and Andrey Gulin. 2018. CatBoost: unbiased boosting with categorical features. Advances in neural information processing systems 31 (2018). [43] Jingang Qu, David Holzmüller, Gaël Varoquaux, and Marine Le Morvan. 2026. TabICLv2: A better, faster, scalable, and open tabular foundation model. (2026). [44] Olivier Roy and Martin Vetterli. 2007. The effective rank: A measure of effective dimensionality. In 2007 15th European signal processing conference. IEEE, 606–610. [45] Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Müller, and Marius Kloft. 2018. Deep one-class classification. In International conference on machine learning. PMLR, 4393–4402. [46] Lukas Ruff, Robert A. Vandermeulen, Nico Görnitz, Alexander Binder, Emmanuel Müller, Klaus-Robert Müller, and Marius Kloft. 2020. Deep Semi-Supervised Anomaly Detection. In International Conference on Learning Representations. [47] Timur Sattarov, Marco Schreyer, and Damian Borth. 2025. Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular Data. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 2478–2489. [48] Mei-Ling Shyu, Shu-Ching Chen, Kanoksri Sarinnapakorn, and LiWu Chang. 2003. A novel anomaly detection scheme based on principal component classifier. Technical report, Miami Univ Coral Gables Fl Dept of Electrical and Computer Engineering. [49] Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee. 2022. Learning from noisy labels with deep neural networks: A survey. IEEE transactions on neural networks and learning systems 34, 11 (2022), 8135–8153. [50] Waqas Sultani, Chen Chen, and Mubarak Shah. 2018. Real-world anomaly detection in surveillance videos. In Proceedings of the IEEE conference on computer vision and pattern recognition. 6479–6488. [51] Bowen Tian, Qinliang Su, and Jian Yin. 2022. Anomaly Detection by Leveraging Incomplete Anomalous Knowledge with Anomaly-Aware Bidirectional GANs. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22. 2255–2261. [52] Yu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh, Johan W Verjans, and Gustavo Carneiro. 2021. Weakly-supervised video anomaly detection with robust temporal feature magnitude learning. In Proceedings of the IEEE/CVF international conference on computer vision. 4975–4986. [53] Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3D convolutional networks. In Proceedings of the IEEE international conference on computer vision. 4489–4497. [54] Boyang Wan, Yuming Fang, Xue Xia, and Jiajie Mei. 2020. Weakly Supervised Video Anomaly Detection via Center-Guided Discriminative Learning. In 2020 IEEE International Conference on Multimedia and Expo (ICME). 1–6. doi:10.1109/ ICME46284.2020.9102722 [55] Peng Wu, Jing Liu, Yujia Shi, Yujia Sun, Fangtao Shao, Zhaoyang Wu, and Zhiwei Yang. 2020. Not only look, but also listen: Learning multimodal violence detection under weak supervision. In European conference on computer vision. Springer, 322–339. [56] Peng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou, Qingsen Yan, Peng Wang, and Yanning Zhang. 2024. Vadclip: Adapting vision-language models for weakly supervised video anomaly detection. In Proceedings of the AAAI conference on artificial intelligence, Vol. 38. 6074–6082. [57] Feng Xiao and Jicong Fan. 2025. Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding. arXiv preprint arXiv:2507.12295 (2025). [58] Hongzuo Xu, Yijie Wang, Guansong Pang, Songlei Jian, Ning Liu, and Yongjun Wang. 2023. RoSAS: Deep semi-supervised anomaly detection with contamination-resilient continuous supervision. Information Processing & Management 60, 5 (2023), 103459. [59] Yajun Xu, Huan Hu, Chuwen Huang, Yibing Nan, Yuyao Liu, Kai Wang, Zhaoxiang Liu, and Shiguo Lian. 2025. TAD: A Large-Scale Benchmark for Traffic Accidents Detection From Video Surveillance. IEEE Access 13 (2025), 2018–2033.
Xu Yao et al.
[60] Jingkang Yang, Pengyun Wang, Dejian Zou, Zitang Zhou, Kunyuan Ding, Wenxuan Peng, Haoqi Wang, Dilip Chen, Bo Li, Yiyou Sun, et al. 2022. OpenOOD: Benchmarking generalized out-of-distribution detection. In Advances in Neural Information Processing Systems, Vol. 35. 32598–32611. [61] Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2021. Understanding deep learning (still) requires rethinking generalization. Commun. ACM 64, 3 (2021), 107–115. [62] Jieyu Zhang, Yue Yu, Yinghao Li, Yujing Wang, Yaming Yang, Mao Yang, and Alexander Ratner. 2021. WRENCH: A Comprehensive Benchmark for Weak Supervision. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks. [63] Xiyuan Zhang et al. 2025. Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. [64] Xingxuan Zhang, Gang Ren, Han Yu, Hao Yuan, Hui Wang, Jiansheng Li, Jiayun Wu, Lang Mo, Li Mao, Mingchao Hao, et al. 2025. Limix: Unleashing structured-data modeling capability for generalist intelligence. arXiv preprint arXiv:2509.03505 (2025). [65] Yue Zhao and Maciej K Hryniewicki. 2018. Xgbod: improving supervised outlier detection with unsupervised representation learning. In 2018 International Joint Conference on Neural Networks (IJCNN). IEEE, 1–8. [66] Yue Zhao, Zain Nasrullah, and Zheng Li. 2019. Pyod: A python toolbox for scalable outlier detection. Journal of machine learning research 20, 96 (2019), 1–7. [67] Yue Zhao, Guoqing Zheng, Subhabrata Mukherjee, Robert McCann, and Ahmed Awadallah. 2023. Admoe: Anomaly detection with mixture-of-experts from noisy labels. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 4937–4945. [68] Jia-Xing Zhong, Nannan Li, Weijie Kong, Shan Liu, Thomas H Li, and Ge Li. 2019. Graph convolutional label noise cleaner: Train a plug-and-play action classifier for anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 1237–1246. [69] Hang Zhou, Junqing Yu, and Wei Yang. 2023. Dual memory units with uncertainty regulation for weakly supervised video anomaly detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 3769–3777. [70] Yingjie Zhou, Xucheng Song, Yanru Zhang, Fanxing Liu, Ce Zhu, and Lingqiao Liu. 2021. Feature encoding with autoencoders for weakly supervised anomaly detection. IEEE Transactions on Neural Networks and Learning Systems 33, 6 (2021), 2454–2465. [71] Zhi-Hua Zhou. 2018. A brief introduction to weakly supervised learning. National science review 5, 1 (2018), 44–53. [72] Bo Zong, Qi Song, Martin Renqiang Min, Wei Cheng, Cristian Lumezanu, Daeki Cho, and Haifeng Chen. 2018. Deep autoencoding gaussian mixture model for unsupervised anomaly detection. In International conference on learning representations.
A Benchmark Details A.1 Dataset Summaries and Processing Specs We detail the diverse collection of datasets evaluated in WSADBench, spanning tabular, image, text and video modalities. These benchmarks are selected to ensure broad representativeness across varying data structures and AD scenarios. A.1.1 Statistical Summaries of Benchmarks. We evaluate our benchmark on a diverse suite of datasets covering tabular, image, and text modalities. The tabular benchmarks span domains including network intrusion, medical diagnosis, and industrial fault detection. The visual and natural language benchmarks cover critical defect detection and text classification tasks. Table 9 summarizes the comprehensive statistics and properties for these three modalities. We also incorporate 4 standard video anomaly detection (VAD) benchmarks, UCF-Crime, ShanghaiTech, XD-Violence, and TAD, covering diverse anomaly types in surveillance and action settings. Table 10 summarizes their detailed characteristics, including video counts and clip distributions. A.1.2 Out-of-Distribution Datasets. We incorporate diverse anomaly datasets featuring various anomaly types for out-of-distribution generalization research. Table 8 summarizes the detailed statistics and properties of these OOD datasets.
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
Table 8: Description of OOD datasets Test (Normal)
Test (Ano)
A.3
Type
Dataset
Train (Normal)
Ano Classes
Source
Texture
AITEX Carpet ELPV
1,692 280 1,131
564 28 377
183 89 715
12 5 2
Link Link Link
Object
Mastcam Metal Nut
9,302 220
426 22
881 93
11 4
Link Link
Medical
HyperKvasir
2,021
674
757
4
Link
A.1.3 Video Feature Specifications and Standardization. To unify VAD with other modalities under MIL, we decouple feature extraction from weakly supervised reasoning. This design isolates algorithm evaluation from upstream video perception. Extracting features using pre-trained models before applying WSAD algorithms is standard in VAD [50, 52, 69]. Under this setting, pre-trained feature extractors capture temporal dynamics and spatial patterns, preserving the complexity of VAD. Additionally, historical VAD evaluations often suffered from inconsistencies due to different pre-trained backbones or checkpoints. This mismatch makes it difficult to determine whether performance gains stem from the algorithm or superior feature representations. To ensure fair comparisons, we constructed a standardized feature extraction pipeline from raw videos. Our pipeline reads video frames, applying RGB normalization and TenCrop spatial augmentations. We extract dense, non-overlapping clips using sliding windows and group them into segments. A producer-consumer architecture feeds these segments into identical frozen pre-trained 3D models, such as I3D [9] and SlowFast [15]. This process generated over 2.4TB of standardized, memory-mapped embedding files. We open-source these features to establish a fair benchmark for evaluating pure weakly supervised reasoning.
A.2
Details of Tabular MIL Bags Generation
We formulate the bag generation process as a sampling procedure from decoupled distributions of normal (X𝑛 ) and abnormal (X𝑎 ) data. The goal is to generate a dataset of bags B = {𝐵 1, 𝐵 2, . . . , 𝐵𝐾 }, where each bag 𝐵𝑖 contains 𝐿 instances (𝐿 = 𝑁𝑠𝑎𝑚𝑝𝑙𝑒𝑠 ). For each bag 𝐵𝑖 , a label 𝑌𝑖 ∈ {0, 1} is determined by a random variable 𝑝 ∼ Uniform(0, 1) and a threshold 𝑃: Normal Bags (𝑌𝑖 = 0): Condition: 𝑝 < 𝑃. The bag is composed entirely of normal instances: 𝐵𝑖 = {𝑥 𝑗 }𝐿𝑗=1,
where 𝑥 𝑗 ∼ X𝑛
Abnormal Bags (𝑌𝑖 = 1): Condition: 𝑝 ≥ 𝑃. The bag is constructed by injecting a random number of anomalies. Let 𝑚 be the number of abnormal instances, sampled uniformly such that 𝑚 ∼ 𝑈 {1, . . . , ⌊𝐿/2⌋ − 1}. The bag is composed of: 𝐿−𝑚 𝐵𝑖 = {𝑥 𝑗 }𝑚 𝑗=1 ∪ {𝑥𝑘 }𝑘=1 ,
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
where 𝑥 𝑗 ∼ X𝑎 , 𝑥𝑘 ∼ X𝑛
This protocol ensures that normal bags are pure, while abnormal bags contain a minority of anomalies mixed with background normal samples, mimicking real-world MIL scenarios. Following the experimental setup in [41], we set the number of instances per bag to 10. Our experiments were conducted on the Tabular datasets, which were preprocessed into bag-structured formats suitable for Multi-Instance Learning.
Experiment Setting
This section provides additional details on experiment setting. Hyperparameter Configurations. For fair comparison and reproducibility, all baseline methods are evaluated using their official default hyperparameters without additional tuning. Classical unsupervised baselines (e.g., LOF, IForest) follow scikit-learn defaults, while deep models adopt consistent optimization strategies (Adam optimizer, learning rate 1e-3 unless otherwise noted). A complete list of parameters—including embedding dimensions, hidden layer widths, loss coefficients, and training epochs—for each model–dataset pair is provided in the supplementary material (see repository link in Appendix D). Computational Resources. . Classical anomaly detection models are run on an Intel(R) Xeon(R) Platinum 8358 CPU @ 2.60GHz, 1TB RAM, 128-core server. For deep learning models, experiments are conducted on a server equipped with eight NVIDIA A800-SXM480GB GPUs. Evaluation Metrics. We employ two standard metrics: Area Under the Receiver Operating Characteristic curve (AUCROC) and Area Under the Precision-Recall curve (AUCPR). AUCROC evaluates the global ranking quality by plotting the True Positive Rate (TPR) against the False Positive Rate (FPR) across all thresholds. It is robust to decision thresholds but can be overly optimistic in highly imbalanced scenarios. AUCPR plots Precision against Recall, focusing specifically on the minority class (anomalies). Given the severe class imbalance in anomaly detection, AUCPR is often considered a more informative metric for practical deployment performance.
B
Calibration Analysis
Anomaly detection benchmarks typically evaluate performance via ranking metrics such as AUC-PR and AUC-ROC. These metrics measure the model’s ability to separate anomalies from normal instances but do not assess probability reliability. Expected Calibration Error (ECE) [22, 35] is a complementary metric that quantifies whether predicted probabilities match empirical anomaly frequencies. We report ECE to provide a reliability perspective beyond ranking quality, evaluating how well each model’s output scores can serve as probability estimates. Model Selection. We evaluate five models spanning three algorithm families: specialized WSAD methods (DevNet, XGBOD), general supervised learning (CatBoost), and tabular foundation models (TabPFN, LimiX). These models are competitive on the 47 tabular datasets and cover the main supervision paradigms in WSADBench. Calibration. To obtain comparable probability estimates, we reserve a validation split for calibration. Specifically, we partition the training data into training and validation subsets, resulting in a 60%/10%/30% train/validation/test split. Under extreme label scarcity, we ensure the validation subset contains at least one anomalous sample to guarantee calibrator convergence. Let 𝑓 (𝑥) ∈ R be the mapping function outputting the raw anomaly score for sample 𝑥. We fit a Platt scaling calibrator on the validation
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 9: Summary of tabular, image, and text datasets evaluated in WSADBench Dataset
#Samples
#Features
Anomaly %
Domain
Dataset
#Samples
Tabular Datasets (Part 1) ALOI annthyroid backdoor breastw campaign cardio Cardiotocography
49,534 7,200 95,329 683 41,188 1,831 2,114
celeba census cover donors fault fraud glass
214
7
80 567,498 1,966 351 6,435 1,600
19 3 1,555 33 36 32
Hepatitis http InternetAds Ionosphere landsat letter
27 6 196 9 62 21 21
3.04 7.42 2.44 34.99 11.27 9.61 22.04
Image Healthcare Network Healthcare Finance Healthcare Healthcare
202,599
39
2.24
299,285 286,048 619,326 1,941 284,807
500 10 10 27 29
6.20 0.96 5.93 34.67 0.17 4.21
Forensic
16.25 0.39 18.72 35.90 20.71 6.25
Healthcare Web Image Physical Astronautics Image
TAD UCF-Crime XD-Violence ShanghaiTech
500 1900 4751 437
Domain
Dataset
18 10 6 100 166 64 10
Image
pendigits
6,870
16
2.27
Image
Sociology Botany Sociology Physical Finance
Pima satellite satimage-2 shuttle skin
768 6,435 5,803 49,097 245,057
8 36 36 9 3
34.90 31.64 1.22 7.15 20.75
Healthcare Astronautics Astronautics Astronautics Image
smtp
95,156
3
0.03
Web
SpamBase speech Stamps thyroid vertebral vowels
4,207 3,686 340 3,772 240 1,456
57 400 9 6 6 12
39.91 1.65 9.12 2.47 12.50 3.43
Document Linguistics Document Healthcare Biology Linguistics
Clip Range Distribution (%)
Max
Min
Mean
Median
Std
≤50
51–100
101–200
201–500
>500
938 61032 16225 261
2 7 3 13
67.62 453.32 246.56 45.66
21 133 158 46
118.27 2055.3 478.77 24.85
74.60% 13.30% 11.20% 67.00%
7.80% 23.60% 19.80% 31.10%
7.00% 27.70% 29.80% 1.40%
9.00% 21.80% 30.00% 0.50%
1.60% 13.50% 9.20% 0.00%
Table 11: Expected Calibration Error (ECE) across different supervision settings. Model
𝛾𝑙𝑎 = 1%
𝛾𝑙𝑎 = 10%
𝛾𝑙𝑎 = 50%
𝛾𝑙𝑎 = 100%
𝑁𝑙𝑎 = 1
𝑁𝑙𝑎 = 5
𝑁𝑙𝑎 = 10
𝑁𝑙𝑎 = 50
DevNet CatBoost XGBOD TabPFN LimiX
0.0265(5) 0.0014(1) 0.0014(1) 0.0023(3) 0.0026(4)
0.0248(1) 0.0315(2) 0.0329(3) 0.0407(4) 0.0435(5)
0.0241(1) 0.0599(3) 0.0587(2) 0.0645(4) 0.0670(5)
0.0235(1) 0.0378(2) 0.0381(3) 0.0388(4) 0.0388(5)
0.0285(5) 0.0011(2) 0.0011(3) 0.0021(4) 0.0011(1)
0.0246(5) 0.0017(2) 0.0015(1) 0.0019(3) 0.0032(4)
0.0297(5) 0.0044(1) 0.0044(2) 0.0057(3) 0.0069(4)
0.0251(1) 0.0497(2) 0.0530(3) 0.0663(4) 0.0664(5)
Table 12: AUCPR performance of additional tabular foundation models across different supervision settings. Model
𝛾𝑙𝑎 = 1%
𝛾𝑙𝑎 = 10%
𝛾𝑙𝑎 = 50%
𝛾𝑙𝑎 = 100%
𝑁𝑙𝑎 = 1
𝑁𝑙𝑎 = 5
𝑁𝑙𝑎 = 10
𝑁𝑙𝑎 = 50
XGBOD DevNet CatBoost TabPFN TabICLv2 Orion-MSP TabDPT Mitra
0.481(8) 0.555(3) 0.521(5) 0.528(4) 0.570(1) 0.562(2) 0.505(6) 0.485(7)
0.688(7) 0.666(8) 0.715(6) 0.770(2) 0.773(1) 0.745(3) 0.743(4) 0.719(5)
0.816(7) 0.713(8) 0.842(5) 0.864(2) 0.872(1) 0.844(4) 0.850(3) 0.831(6)
0.864(6) 0.722(8) 0.883(4) 0.900(2) 0.902(1) 0.872(5) 0.885(3) 0.863(7)
0.385(8) 0.517(1) 0.457(4) 0.415(6) 0.482(3) 0.492(2) 0.397(7) 0.429(5)
0.580(8) 0.610(7) 0.635(5) 0.667(3) 0.710(1) 0.685(2) 0.649(4) 0.623(6)
0.657(7) 0.648(8) 0.689(5) 0.735(2) 0.761(1) 0.733(3) 0.709(4) 0.688(6)
0.791(7) 0.697(8) 0.809(5) 0.842(2) 0.859(1) 0.835(3) 0.823(4) 0.800(6)
subset to map raw scores to the binary label space 𝑌 ∈ {0, 1}. The calibrated probability is formulated as: 1 𝑝ˆ (𝑦 = 1 | 𝑥) = 𝜎 (𝐴 · 𝑓 (𝑥) + 𝐵) = , 1 + exp(𝐴 · 𝑓 (𝑥) + 𝐵) where parameters 𝐴, 𝐵 are optimized by minimizing the negative log-likelihood loss L on the validation subset. Finally, ECE is computed on the calibrated test probabilities. We partition the predicted probability range [0, 1] into 𝑀 = 10 equally spaced bins. The ECE aggregates the absolute difference between average confidence and empirical accuracy across all bins, weighted by sample frequency. Evaluation and Analysis. We evaluate ECE across both RLA and NLA settings, covering 𝛾𝑙𝑎 ∈ {1%, 10%, 50%, 100%} and 𝑁𝑙𝑎 ∈ {1, 5, 10, 50}, on all 47 tabular datasets. The computed ECE values are averaged across all 47 tabular datasets, with results reported in
4.05 35.16 2.32 9.21 3.17 2.88 9.46
#Samples
#Features
Anomaly %
Domain
Tabular Datasets (Part 3)
148 19,020 11,183 7,603 3,062 5,216 5,393
Clip Number Statistics
Video Num
Anomaly %
Lymphography magic.gamma mammography mnist musk optdigits PageBlocks
Table 10: Statistics of video anomaly detection Benchmarks. Dataset
#Features
Tabular Datasets (Part 2) Healthcare Physical Healthcare Image Chemistry Image Document
Waveform WBC WDBC Wilt wine WPBC yeast
3,443 223 367 4,819 129 198 1,484
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
5,263 6,315 10,000 5,354 5,208
Agnews Amazon Imdb Yelp 20newsgroups
10,000 10,000 10,000 10,000 11,905
21 9 30 5 13 33 8
2.90 4.48 2.72 5.33 7.75 23.74 34.16
Physics Healthcare Healthcare Botany Chemistry Healthcare Biology
5.00 5.00 5.00 23.50 5.00
Image Image Image Image Image
5.00 5.00 5.00 5.00 4.96
NLP NLP NLP NLP NLP
Image Datasets 512 512 512 512 512 Text Datasets 768 768 768 768 768
Table 11. Overall, the ECE results complement the ranking-based evaluation. The best-calibrated method varies across supervision regimes. CatBoost and XGBOD perform well under extreme label scarcity, whereas DevNet improves when more labeled anomalies are available. TabPFN and LimiX do not show a consistent ECE advantage in this experiment.
C
Further Experimental Results on Tabular Foundation Models
To solidify our findings on the dominance of tabular foundation models, we expanded our evaluation to include four additional stateof-the-art tabular foundation models representing diverse design strategies: TabICLv2 [43] and Orion-MSP [7] extend the in-context learning (ICL) paradigm with improved attention mechanisms for better scalability; TabDPT [34] incorporates self-supervised learning on real data to scale with larger corpora; and Mitra [63] focuses on synthetic prior design for improved generalization. Together, these models cover the major development directions in current tabular foundation model research. As shown in Table 12, all four newly evaluated tabular foundation models consistently outperform the WSAD baseline (DevNet) and the supervised baseline (XGBOD) across supervision levels, reinforcing our conclusion that modern tabular foundation models hold a strong advantage in WSAD tasks. Among these models, TabICLv2 achieves the best overall performance under ratio-based supervision (𝛾𝑙𝑎 ), while Orion-MSP demonstrates a relative advantage under the extremely scarce one-shot setting (𝑁𝑙𝑎 = 1). The performance gap between tabular foundation models and non-foundation baselines widens as supervision increases, suggesting that tabular foundation models benefit more from additional labeled anomalies due to stronger in-context learning capacity.
D
Open-Source Repository
To facilitate reproducibility and future research, we release the full implementation of the WSADBench benchmark at: https://github.com/SUFE-AILAB/WSADBench We encourage the research community to use and extend WSADBench to explore new directions in WSAD.
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 13: Specifications and average performance of pretrained video backbones. Inputs denote (Frames × Sample Rate), and metrics are averaged over downstream algorithms.
and Table 57 report the AUCPR performance under varying 𝛾𝑙𝑎 and 𝑛𝑙𝑎 settings compared to these unsupervised benchmarks. These comparisons highlight the performance gains achieved by introducing varying degrees of incomplete supervision.
Backbone
Abbr.
Input (𝑇 × 𝜏)
Params (M)
Feat. Dim.
AUCROC
AUCPR
I3D_R50 X3D_M MViT_B SlowFast_R50 SlowFast_R101
i3d x3d mvit sf50 sf
8×8 16×5 32×3 8×8 8×8
28.04 3.79 36.61 34.57 62.83
2048 2048 768 2304 2304
0.8199 (5) 0.8278 (3) 0.8274 (4) 0.8393 (2) 0.8489 (1)
0.3528 (5) 0.3654 (4) 0.3778 (2) 0.3772 (3) 0.3875 (1)
Table 14: Performance comparison between 32-seg and 200seg settings (AUCPR). Model AR-Net Sultani MGFN RTFM UR-DMU VadCLIP GCN-Anomaly
32-seg
200-seg
Improvement (%)
0.430 ± 0.189 (1) 0.392 ± 0.153 (2) 0.258 ± 0.124 (7) 0.386 ± 0.162 (3) 0.340 ± 0.169 (5) 0.344 ± 0.140 (4) 0.304 ± 0.189 (6)
0.429 ± 0.188 (1) 0.377 ± 0.149 (3) 0.297 ± 0.129 (7) 0.380 ± 0.162 (2) 0.353 ± 0.169 (4) 0.348 ± 0.142 (5) 0.303 ± 0.203 (6)
-0.13 -3.89 15.03 -1.56 3.56 1.20 -0.31
E Extended Results for WSAD Performance E.1 Incomplete Supervision This section presents the detailed experimental results corresponding to the incomplete supervision analysis in the main paper. Tables 23, 24, 29, and 30 present the Tabular Benchmark results. Tables 25, 26, 31, and 32 summarize the CV Datasets. Similarly, Tables 28, 27, 33, and 34 provide the NLP Datasets performance. For multimodality datasets, Tables 35–46 exhibit the Ratio-labeled Anomaly (𝛾𝑙𝑎 ) results. Tables 47–54 provide the corresponding Number-labeled Anomaly (𝑁𝑙𝑎 ) performance. These tables offer detailed rankings, average scores, and standard deviations for each algorithm. They serve as an extended reference for the benchmark evaluation. To rigorously compare the performance differences among algorithms under extremely limited supervision settings, we employ the Nemenyi post-hoc test to generate Critical Difference (CD) diagrams, as shown in Figure 9. Figure 9a visualizes the statistical ranking on the Tabular Benchmark with only one labeled anomaly (𝑁𝑙𝑎 = 1). Similarly, Figure 9b and Figure 9c illustrate the model comparisons on image and text datasets under the 1% labeled anomaly ratio (𝛾𝑙𝑎 = 1%), respectively. These diagrams provide a clear perspective on the relative effectiveness of different methods in data-scarce scenarios. Furthermore, Figure 9 bottom row aggregates performance by algorithm categories to align with Section 4.1.1. Across all modalities under extreme scarcity (𝑁𝑙𝑎 = 1 or 𝛾𝑙𝑎 = 1%), tabular foundation models exhibit no statistically significant superiority. Specifically, they share critical difference bars with other top-performing categories, including score learning, GBDTs, and data augmentation. These statistical ties indicate that no single algorithm category, including tabular foundation models, achieves absolute dominance under such scarce supervision. To further demonstrate the effectiveness of using limited labeled anomalies, we compare the weakly supervised performance against the best-performing unsupervised baselines (PCA for Tabular, CBLOF for CV, and AutoEncoder for NLP). Table 55, Table 56,
E.2
Inexact Supervision
This section provides three supplementary analyses to extend the main text (Section 4.1.2), detailing experimental configurations, comprehensive performance across different feature backbones, and additional findings on model sensitivity. Specifically, Table 58 includes the comprehensive performance of general-purpose tabular classifiers applied to video anomaly detection tasks, providing the detailed data supporting the interoperability discussion in Section 4.2.5. E.2.1 Benchmark Results on Video AD. Following standard MILbased VAD protocols (e.g., Sultani), we extract clip-level features (16-frame non-overlapping clips) using pre-trained backbones. To handle variable video lengths, we employ temporal pooling to standardize each video into a fixed sequence of 𝑛 segments (default 𝑛 = 32, with 𝑛 = 200 for granularity analysis). Evaluation is performed on the test set using the final epoch model, with frame-level scores aligned to calculate AUCROC and AUCPR. The performance is summarized in Table 59 and Table 60. Selection of Pre-trained Models Significantly Impacts Average VAD Performance. Our experiments reveal distinct performance disparities driven solely by the choice of feature extractor. As summarized in Table 13, the top-performing backbone, SlowFast-R101, achieves an average AUC-ROC of 0.8489 across all downstream algorithms, whereas the older I3D-R50 backbone reaches only 0.8199. This substantial gap highlights that upgrading the upstream feature representation is as critical as designing sophisticated anomaly detection heads. Feature Quality Weighs More Than Parameter Count. Counterintuitively, larger models do not guarantee better features for VAD. As detailed in Table 13, X3D-M, despite having only 3.79M parameters, outperforms the much larger I3D-R50 (28.04M). Similarly, MViT-B achieves competitive results with a compact 768dimensional feature space, outperforming 2048-dimensional features from ResNet-based models. This suggests that the semantic discriminability of features is more critical than raw model capacity or embedding dimensionality. Temporal Granularity Differentially Impacts Model Architectures. Table 14 reveals that finer temporal segmentation (𝑛 = 200 vs 𝑛 = 32) benefits complex models while hindering simpler ones. Attention-based models like MGFN gain significantly (15% AUCPR increase) from higher resolution, leveraging their "focus" modules to pinpoint anomalies. Conversely, simpler MIL models like Sultani suffer a performance drop (3.89% AUCPR decrease), likely because finer granularity introduces noise that overwhelms their coarse ranking constraints. Models with robust feature aggregation, such as AR-Net and RTFM, remain stable across granularities. E.2.2 Relative Contribution of Backbone and Model. We examine the relative contributions of feature representation (pre-trained
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
CD = 5.07 (p=0.05)
CD = 6.21 (p=0.05)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
1
28
2
3
2
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
Average Rank (AUCPR) - Lower is Better
Average Rank (AUCPR) - Lower is Better
(a) Tabular Datasets (𝑁𝑙𝑎 = 1)
(b) Image Datasets (𝛾𝑙𝑎 = 1%)
CD = 1.98
1
4
CD = 11.80 (p=0.05)
23
24
25
26
1
2
3
4
5
4
5
6
7
8
9
10
1
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
(c) Text Datasets (𝛾𝑙𝑎 = 1%)
CD = 3.76
2
3
Average Rank by Category (AUCPR) [ nla = 1 ] Lower is Better
4
5
6
7
8
Average Rank by Category (AUCPR) [ rla = 1% ] Lower is Better
(d) Tabular Datasets (𝑁𝑙𝑎 = 1)
7
Average Rank (AUCPR) - Lower is Better
CD = 1.73
3
6
(e) Image Datasets (𝛾𝑙𝑎 = 1%)
9
10
1
2
3
4
5
6
7
8
9
10
Average Rank by Category (AUCPR) [ rla = 1% ] Lower is Better
(f) Text Datasets (𝛾𝑙𝑎 = 1%)
Figure 9: Critical Difference (CD) diagrams under extremely limited supervision for each modality. The top row (a-c) compares individual models, whereas the bottom row (d-f) compares algorithm categories. Average ranks are based on AUCPR (lower rank is better). Horizontal bars indicate groups without statistically significant differences (Nemenyi test, 𝛼 = 0.05). Table 15: Multi-factor ANOVA results quantifying the contribution of factors to AUCROC and AUCPR variance. AUCROC
Factor (Interaction) Dataset Dataset × Classifier Classifier Dataset × Pretraining Pretraining Classifier × Pretraining Seed Residual
AUCPR
Sum Sq.
𝑃-Value
Con.(%)
Sum Sq.
𝑃-Value
Con.(%)
0.7251 0.3141 0.2572 0.2037 0.0774 0.0699 0.0036 0.6279
3.00 × 10−104 2.08 × 10−44 6.47 × 10−44 1.19 × 10−31 5.02 × 10−15 5.37 × 10−6 0.467 -
31.82 13.78 11.29 8.94 3.40 3.07 0.16 27.55
16.5964 1.0222 1.2465 0.1916 0.1042 0.1266 0.0130 0.8800
< 10−300 1.44 × 10−92 8.26 × 10−117 7.23 × 10−21 1.88 × 10−14 6.88 × 10−9 0.056 -
82.24 5.07 6.18 0.95 0.52 0.63 0.06 4.36
Table 16: Comparision of Weakly Supervised Video Anomaly Detection at clip and frame levels. Model AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
AUC (Clip)
AUC (Frame)
PR (Clip)
PR (Frame)
0.783±0.008 (2) 0.798±0.006 (1) 0.770±0.015 (4) 0.778±0.002 (3) 0.769±0.005 (5) 0.753±0.011 (7) 0.767±0.006 (6)
0.806±0.001 (1) 0.795±0.006 (3) 0.779±0.014 (5) 0.806±0.001 (1) 0.760±0.014 (6) 0.755±0.010 (7) 0.793±0.002 (4)
0.196±0.006 (3) 0.172±0.004 (6) 0.203±0.021 (1) 0.199±0.003 (2) 0.178±0.011 (4) 0.159±0.013 (7) 0.178±0.006 (5)
0.207±0.000 (3) 0.171±0.003 (6) 0.209±0.024 (2) 0.216±0.002 (1) 0.185±0.017 (5) 0.162±0.013 (7) 0.188±0.003 (4)
Table 17: Ground truth alignment methods in different WSAD approaches Method MGFN Sultani VadCLIP GCN-Anomaly AR-Net RTFM UR-DMU WSADBench(Ours)
Aligned with frames
Aligned with test clips
Unclear
✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
backbone) and inductive bias (classifier design) to VAD performance, as well as their potential synergistic effects. By analyzing these factors, we aim to uncover whether performance gains are driven primarily by better features or superior detection algorithms.
To decouple the impact of feature representation and inductive bias, we conduct an N-way Analysis of Variance (ANOVA) across 4 datasets, 5 backbones, and 7 algorithms. We model performance (𝑦) as a linear combination of main effects (Dataset 𝐷, Pre-training 𝑃, Classifier 𝐶, Seed 𝑆) and their interactions: 𝑦 = 𝜇 + 𝛼 𝐷 + 𝛽𝑃 + 𝛾𝐶 + 𝛿𝑆 + (𝛼𝛽)𝐷𝑃 + (𝛼𝛾)𝐷𝐶 + (𝛽𝛾)𝑃𝐶 + 𝜖 (1) Contribution rates, derived from the Sum of Squares, quantify the variance explained by each factor. Statistical significance is assessed at 𝑃 < 0.05. Detailed results for AUCROC and AUCPR are provided in Table 15. The ANOVA results reveal two pivotal findings. First, across both metrics the classifier consistently explains more variance than the backbone (AUCROC: 11.29% vs. 3.40%; AUCPR: 6.18% vs. 0.52%), confirming that inductive bias outweighs feature representation. Second, the Dataset x Classifier interaction under AUCROC (13.78%) exceeds the Classifier main effect (11.29%), indicating that the optimal algorithm varies across datasets—no single classifier universally dominates. E.2.3 Ground Truth Alignment. During baseline reproduction, we observed inconsistencies in how different methods align predicted anomaly scores with ground truth (GT) labels. Table 17 summarizes these divergent protocols, which we hypothesize introduce systematic bias and potentially alter model rankings. We compare two prevalent alignment strategies: Protocol 1: Frame-canonical Alignment. This approach prioritizes the physical duration of the raw video. Spatially, scores from multiple crops are averaged to a single value per clip. Temporally, clip-level scores are interpolated to restore frame-level resolution and then truncated to match the exact length of the original video (𝐿𝑟𝑎𝑤 ). This method strictly evaluates valid frames, ignoring padding artifacts introduced by feature extraction. Protocol 2: Feature-grid Alignment. This approach operates on the feature grid, preserving all crops and padding. Spatially, crops
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
are treated as independent instances (flattened). Temporally, the ground truth is padded with "normal" labels to match the feature grid size (𝐿𝑔𝑟𝑖𝑑 > 𝐿𝑟𝑎𝑤 ) and tiled to align with each crop. This artificially inflates the evaluation set with non-existent padding frames and treats spatially disjoint crops as independent samples. We empirically quantify the impact of these strategies using the UCF-Crime dataset with I3D features, standardized to 32 temporal segments. Analysis. Quantitative results (Table 16) demonstrate that alignment strategies significantly impact model rankings. Protocol 1 (Frame-canonical) yields higher performance for spatially sensitive models like AR-Net and Sultani (0.806 AUC), whereas Protocol 2 (Feature-grid) results in a performance drop (e.g., Sultani falls to 0.778), shifting the lead to MGFN (0.798 AUC). This shift can be explained by a structural conflict we term the "Spatial Sparsity Penalty." Protocol 2 mechanically broadcasts the frame-level anomaly label to all spatial crops. Consequently, in scenarios where the anomaly is localized (e.g., a specific object) against a normal background, Protocol 2 treats the entire grid as anomalous. Models like Sultani, designed to suppress background scores via sparsity constraints, are thus penalized for correctly predicting low scores on these background regions (False Negatives under Protocol 2). In contrast, Protocol 1 averages spatial scores to evaluate the frame as a single unit. This allows localized anomalies to sufficiently elevate the frame score, avoiding penalties for validly suppressed background regions. Therefore, we adopt Protocol 1 as the standard for VAD benchmarking to ensure fairness for spatially discriminative methods.
employing a 7:3 stratified train-test split. Training is restricted to A𝑘𝑛𝑜𝑤𝑛 under limited supervision, where only a fraction (𝛾𝑙𝑎 ) of samples is labeled, and the rest are treated as unlabeled background to simulate realistic label noise. The test set evaluates robustness by mixing ID (A𝑘𝑛𝑜𝑤𝑛 ) and OOD (A𝑢𝑛𝑘𝑛𝑜𝑤𝑛 ) anomalies, with the proportion of unknown classes explicitly controlled by the OOD rate. Figure 10 presents the complete performance landscape across OOD datasets. These detailed visualizations serve as supplementary empirical evidence reinforcing the robustness analysis discussed in Section 4.2.4, demonstrating that the observed trends are consistent across varying evaluation metrics (AUCROC and AUCPR).
1.00 0.95 0.90 0.85 0.80 0.75 0.70 0.65 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0
RLA = 0.1
RLA = 0.5
RLA = 1.0 Model
25
50
75
OOD Rate (%)
100
0
Repr. Learning GAN-based Score Learning
25
50
75
Tabular MIL Bag Construction Sensitivity
Table 18: AUCPR sensitivity of models to Tabular MIL bag construction hyperparameters. Parentheses denote rank within each column. prob = 0.3, 𝛾𝑙𝑎 = 100%
OOD Generalization
AUCPR
AUCROC
E.3
E.4
To verify that benchmark conclusions are not sensitive to bag construction choices, we conduct a sensitivity analysis on tabular Multiple Instance Learning (MIL). We vary two key hyperparameters independently: bag size and abnormal-bag ratio. For the bag size analysis, we set the abnormal-bag probability to 0.3 and label ratio to 100%, while varying the bag size in {10, 20, 30}. For the abnormalbag ratio analysis, we fix the bag size to 10 and label ratio to 100%, varying the abnormal-bag probability in {0.1, 0.2, 0.3}. We evaluate five representative models: DeepSAD [46], REPEN [37], DevNet [40], TabPFN [21], and Sultani [50].
100
OOD Rate (%) Data Aug. Pseudo-Labeling GBDT
0
25
50
75
OOD Rate (%) Deep(Sup.) Found. Model Diffusion DAE
100
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX DDAE Median Unsup.
Figure 10: Visualization of incomplete OOD results. Row 1 shows the AUCROC performance, and Row 2 shows the AUCPR performance under OOD conditions. Unlike the controlled OOD sampling based on feature distance presented in the main text, this section evaluates generalization under a more generalized setting without explicit distance constraints. We follow the protocol proposed by Ding et al. [12], where anomalies are split solely based on distinct semantic classes: one class is designated as Known (A𝑘𝑛𝑜𝑤𝑛 ) and all others as Unknown (A𝑢𝑛𝑘𝑛𝑜𝑤𝑛 ). To ensure rigorous evaluation, we select datasets containing at least two anomaly types and a minimum of 15 samples,
bag = 10, 𝛾𝑙𝑎 = 100%
Model
bag=10
bag=20
bag=30
prob=0.1
prob=0.2
prob=0.3
DeepSAD REPEN DevNet TabPFN Sultani
0.302(1) 0.276(3) 0.265(5) 0.284(2) 0.272(4)
0.328(1) 0.303(2) 0.302(3) 0.286(5) 0.296(4)
0.322(1) 0.309(2) 0.309(2) 0.274(5) 0.303(4)
0.286(1) 0.255(3) 0.239(5) 0.277(2) 0.251(4)
0.285(1) 0.258(3) 0.245(5) 0.266(2) 0.252(4)
0.302(1) 0.276(3) 0.265(5) 0.284(2) 0.272(4)
Table 18 reports the average AUCPR values and corresponding model ranks under these configurations. The experimental results demonstrate that relative model rankings remain highly consistent across the evaluated settings. For instance, DeepSAD consistently dominates the comparison by achieving the top rank in all test cases. Similarly, Sultani consistently occupies the fourth rank under every hyperparameter setting. Although minor rank shifts exist between REPEN and DevNet under larger bag sizes, their overall performance differences remain small. These findings verify that the relative performance of anomaly detection methods is robust to the specific choices of MIL dataset construction.
F
Algorithms List
We provide a comprehensive overview of the 36 algorithms evaluated in WSADBench, spanning diverse supervision paradigms including weakly-supervised (instance/bag), unsupervised, supervised, and tabular foundation models. Detailed attributes for each method, such as publication venue, backbone architecture, and code availability, are summarized in Table 19.
Setting III: Setting II: ID Near, OOD Near ID Near, OOD Far
Setting I: ID Far, OOD Near
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
RLA=0.1 (AUCROC)
1.0
RLA=0.5 (AUCROC)
Xu Yao et al.
RLA=1.0 (AUCROC)
0.8
0.8
0.6
0.7
0.4
0.6
0.2
0.5 1.0
0.0 1.0
0.9
0.8
0.8
0.6
0.7
0.4
0.6
0.2
0.5 1.0
0.0 1.0
0.9
0.8
0.8
0.6
0.7
0.4
0.6
0.2
0.5
0
25
50
75
OOD Rate (%)
100
0
25
50
75
OOD Rate (%)
Repr. Learning GAN-based
100
0
25
50
75
OOD Rate (%)
Score Learning Data Aug.
RLA=0.1 (AUCPR)
1.0
0.9
100
0.0
RLA=0.5 (AUCPR)
RLA=1.0 (AUCPR)
Model
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX DDAE Median Unsup.
0
25
50
75
OOD Rate (%)
Pseudo-Labeling GBDT
100
0
25
50
75
OOD Rate (%)
Deep(Sup.) Found. Model
100
0
25
50
75
OOD Rate (%)
100
Diffusion DAE
Figure 11: Visualization of AUCROC and AUCPR distributions under three different incomplete OOD settings.
F.1
Weakly-supervised (Instance) • REPEN [37]: A representation learning method that leverages limited labeled anomalies to learn low-dimensional embeddings where anomalies are significantly distant from normal data, optimized via a triplet ranking loss. • DevNet [40]: An end-to-end framework that learns anomaly scores by enforcing a Gaussian prior on normal data and maximizing the Z-score deviation of anomalies from the normal distribution. • DeepSAD [46]: A deep semi-supervised extension of DeepSVDD that minimizes the distance of normal samples to the hypersphere center while penalizing the inverse distance of labeled anomalies to push them away. • FEAWAD [70]: A reconstructive framework that utilizes an autoencoder to map instances into a latent space where anomalies are poorly reconstructed, while simultaneously training a classifier to separate encoded features of labeled anomalies from normal data. • PReNet [39]: A pairwise relation prediction network that learns anomaly scores by contrasting ordinal pairs of labeled anomalies and unlabeled samples, enforcing a large margin between known anomalies and the unlabeled distribution. • RoSAS [58]: A robust semi-supervised framework that employs a purity-aware label smoothing strategy to mitigate noise in unlabeled data, combined with a deviation loss to learn discriminative anomaly scores. • XGBOD [65]: An extreme gradient boosting-based outlier detection framework that enhances unsupervised outlier scores with new features generated from various extraction methods, optimized under a semi-supervised setting. • AA-BiGAN [51]: An augmented adversarial bi-directional GAN that addresses class imbalance by augmenting labeled anomalies and enforces cycle-consistency to map normal data to a latent distribution distinct from anomalies.
• Dual-MGAN [30]: A dual-memory GAN framework that employs separate memory modules for normal and abnormal patterns to mitigate mode collapse, learning discriminative features through adversarial training. • PUMA [41]: A framework that formulates anomaly detection as a Positive-Unlabeled (PU) learning problem under multi-instance constraints, estimating the class prior to robustly learn from bags with unknown instance labels. • GANomaly [4]: A semi-supervised adversarial framework that employs an encoder-decoder-encoder architecture to minimize the reconstruction error of normal samples in both image and latent spaces. • DDAE [47]: A denoising diffusion autoencoder that utilizes diffusion noise to corrupt input data and trains a denoiser to reconstruct the original signal, using the reconstruction error as the anomaly score.
F.2
Unsupervised (Instance) • DeepSVDD [45]: A deep one-class classification method that maps normal data into a minimum-volume hypersphere in feature space. • CBLOF [24]: A clustering-based method that assigns anomaly scores based on the size of the cluster a point belongs to and its distance to the cluster center. • LOF [8]: A density-based algorithm that measures the local density deviation of a given data point with respect to its neighbors. • AutoEncoder [72]: A deep neural network trained to reconstruct input data, identifying anomalies through high reconstruction errors. • VAE [27]: A probabilistic generative model that learns the underlying data distribution, detecting anomalies by their low likelihood or high reconstruction loss.
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 19: Overview of evaluated algorithms Supervision Type
Category
Venue
Backbone
Modality
Code
XGBOD DeepSAD REPEN AA-BiGAN Dual-MGAN DevNet FEAWAD PReNet RoSAS SOEL-NTL DDAE GANomaly
Repr. Learning Score Learning Repr. Learning GAN-based Data Aug. Score Learning Reconstruction Score Learning Data Aug. Pseudo-Labeling Diffusion DAE GAN-based
IJCNN’18 ICLR’20 KDD’18 IJCAI’22 TKDD’22 KDD’19 TNNLS’21 KDD’23 IP&M’23 ICML’23 KDD’25 ACCV’18
Tree-based MLP MLP GAN GAN MLP AE MLP MLP MLP Diffusion AE GAN
Tabular Tabular Tabular Tabular Tabular Tabular Tabular Tabular Tabular Tabular Tabular Tabular
Link Link Link Link Link Link Link Link Link Link Link Link
Sultani VadCLIP MGFN RTFM AR-Net UR-DMU GCN-Anomaly PUMA
Vanilla MIL Lang-Guided MIL Magnitude MIL Magnitude MIL Dynamic MIL Uncertainty MIL Label Denoising PU MIL
CVPR’18 AAAI’24 AAAI’23 ICCV’21 ICME’20 AAAI’23 CVPR’19 KDD’23
MLP Attention CNN+Attn CNN+Attn MLP Attention GCN AE
Video Video Video Video Video Video Video Tabular
Link Link Link Link Link Link Link Link
Unsupervised (Instance)
IForest AutoEncoder DeepSVDD VAE PCA ECOD CBLOF LOF LUNAR
Isolation-based Reconstruction Deep One-class Reconstruction Reconstruction Probabilistic Cluster-based Density-based GNN-based
ICDM’08 ICLR’18 ICML’18 ICLR’14 ICDM’03 workshop TKDE’22 PRL’03 SIGMOD’00 AAAI’22
Tree-based MLP MLP MLP Linear Probabilistic Cluster-based Density-based GNN
Tabular Tabular Tabular Tabular Tabular Tabular Tabular Tabular Tabular
Link Link Link Link Link Link Link Link Link
Supervised (Instance)
XGBoost CatBoost FTTransformer TabM TabR-S
GBDT GBDT Deep (Sup.) Deep (Sup.) Deep (Sup.)
KDD’16 NeurIPS’18 NeurIPS’21 ICLR’25 ICLR’24
GBDT GBDT Transformer MLP MLP
Tabular Tabular Tabular Tabular Tabular
Link Link Link Link Link
Foundation Models (Instance)
TabPFN LimiX
Found. Model Found. Model
ICLR’23 ArXiv’25
Transformer Transformer
Tabular Tabular
Link Link
Weakly-supervised (Instance)
Weakly-supervised (Bag)
Method
• ECOD [31]: An unsupervised method that estimates the underlying distribution of data using empirical cumulative distribution functions. • LUNAR [16]: A graph neural network-based approach that learns to classify nodes as normal or anomalous based on their local neighborhood structure. • IForest [32]: An ensemble method that isolates anomalies using random trees, exploiting the property that anomalies are susceptible to isolation (shorter path lengths). • PCA [48]: A linear dimensionality reduction technique that detects anomalies by projecting data onto principal components and measuring reconstruction error.
F.3
Weakly-supervised (Bag) • AR-Net [54]: An MIL framework incorporating Center Loss to enforce intra-class compactness of normal features and a dynamic instance selection strategy to filter noisy labels.
• Sultani [50]: The seminal MIL-based method that introduces a ranking loss to maximize the separability between the highest-scoring instance in abnormal bags and that in normal bags. • MGFN [11]: An attention-based network that integrates feature magnitude learning with a glance-and-focus mechanism to capture global context and refine local anomaly localization. • VadCLIP [56]: A vision-language framework that adapts CLIP for VAD using dual-branch prompt learning to align video segments with anomaly-related textual descriptions. • UR-DMU [69]: An uncertainty-regulated dual-memory unit that uses separate memory banks for normal/abnormal patterns and an uncertainty module to weigh feature reliability and filter noise. • GCN-Anomaly [68]: A graph-based framework that treats anomaly detection as a label noise cleaning problem, using
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 20: Comprehensive description and interpretation of dataset meta-features. Category
Meta-Feature
Sym.
Definition
Interpretation
Basic Stats
Sample Size Dimensionality Sparsity Ratio
𝑁 𝐷 𝜌𝑠𝑝
Total number of instances. Number of features. Fraction of zero elements.
Larger 𝑁 improves model stability but increases training cost. High 𝐷 with low 𝑁 increases risk of overfitting (curse of dimensionality). High values indicate sparse data (e.g., text), suitable for specialized sparse solvers.
Intrinsic Dim. & Corr.
Avg. Feat. Sim. Effective Rank Eff. Rank Ratio Intrinsic Dim. ID Ratio
𝜌¯𝑓 𝐸𝑅 𝑅𝐸𝑅 𝐼𝐷 𝑅𝐼 𝐷
Mean pairwise correlation. Number of significant principal components. 𝐸𝑅 normalized by 𝐷 . MLE of manifold dimension. 𝐼 𝐷 normalized by 𝐷 .
High values imply high redundancy and potential for feature selection. Indicates the true linear dimensionality of the data. Low values imply features reside in a low-dimensional linear subspace. High values indicate complex non-linear structures. Measures compactness of the data manifold relative to feature space.
Complexity
Fisher’s Ratio Borderline Pts. Intra/Inter Ratio RF’s F1
𝐹1 𝑁1 𝑁2 −
Max Fisher discriminant ratio. Fraction of points in MST crossing class boundary. Ratio of within-class to between-class distance. F1-score of Random Forest.
High values indicate high linear separability between classes. High values imply complex, overlapping decision boundaries. Low values indicate distinct, compact clusters (easier classification). Serves as a proxy for overall task difficulty (High = Easy).
Anom. Cluster. degree
𝐶𝑙𝑢𝑠
Clustering coefficient of anomalies.
Local/Global Outlier Ratio
𝐿𝐺𝑅
Ratio of local to global outliers.
High values mean anomalies form dense micro-clusters rather than being scattered. High values indicate anomalies are context-dependent rather than extreme values.
Anomaly Topology
Note: Complexity measures (𝐹 1, 𝑁 1, 𝑁 2) are adopted from Ho & Basu [25]; 𝐼 𝐷 uses the MLE estimator [28]; 𝐸𝑅 is based on [44].
GCNs to propagate pseudo-labels based on feature similarity and temporal consistency. • RTFM [52]: A robust temporal feature magnitude learning method that enforces separability in the feature magnitude domain with additional temporal smoothness and sparsity constraints.
F.4
Supervised (Instance) • XGBoost [10]: A scalable tree boosting system that implements gradient-boosted decision trees with regularization to control overfitting and optimize computational efficiency. • CatBoost [42]: A gradient boosting framework specialized for categorical features, using ordered boosting to mitigate prediction shift and handle categorical variables without extensive preprocessing. • FTTransformer [19]: A Transformer-based architecture adapted for tabular data, treating features as tokens and applying self-attention to capture complex feature interactions. • TabM [17]: A modern deep tabular learning framework that leverages model ensembling and advanced regularization techniques to improve robustness on heterogeneous tabular datasets. • TabR-S [18]: A retrieval-augmented tabular deep learning model that enhances prediction by retrieving and aggregating information from similar training examples (neighbors) during inference.
F.5
Tabular Foundation Models (Instance) • TabPFN [21]: A Prior-Data Fitted Network that learns a Bayesian inference algorithm on synthetic datasets, enabling efficient in-context learning for tabular classification without gradient-based training. • LimiX [64]: A general-purpose large model for structured data modeling that leverages massive cross-table pretraining to transfer knowledge to downstream tasks, exhibiting strong zero-shot generalization capabilities.
G
Dataset Meta-features and Correlation Analysis G.1 Meta-feature Definitions To systematically characterize the intrinsic properties of datasets and investigate their impact on anomaly detection performance, we extracted a comprehensive set of meta-features. The definitions and mathematical formulations of some key meta-features are detailed 𝑁 denote a dataset with 𝑁 samples and below. Let D = {(𝑥𝑖 , 𝑦𝑖 )}𝑖=1 𝐷 features, where 𝑦𝑖 ∈ {0, 1} represents the label (0 for normal, 1 for abnormal). Average Feature Similarity (Avg. Feat. Sim., 𝜌¯𝑓 ): This metric quantifies feature redundancy by averaging the absolute Pearson correlation coefficients between all pairs of features. 𝜌¯𝑓 =
𝐷 −1 ∑︁ 𝐷 ∑︁ 2 |corr(𝑓𝑖 , 𝑓 𝑗 )| 𝐷 (𝐷 − 1) 𝑖=1 𝑗=𝑖+1
where 𝑓𝑖 denotes the 𝑖-th feature vector. A higher 𝜌¯𝑓 indicates strong linear dependencies among features. Intrinsic Dimensionality (𝐼𝐷): We estimate the non-linear manifold dimension of the data using the Maximum Likelihood Estimation (MLE) method based on nearest neighbor distances, which can explain 95% variance. For each sample 𝑥𝑖 , let 𝑟𝑘 (𝑥𝑖 ) be the distance to its 𝑘-th nearest neighbor. The 𝐼𝐷 is estimated as: # −1 " 𝑁 𝑘 −1 1 ∑︁ 1 ∑︁ 𝑟𝑘 (𝑥𝑖 ) 𝐼𝐷 = ln 𝑁 𝑖=1 𝑘 − 1 𝑗=1 𝑟 𝑗 (𝑥𝑖 ) This value represents the minimum number of variables required to describe the data without significant information loss. ID Ratio (𝑅𝐼 𝐷 ): Defined as the ratio of the intrinsic dimensionality to the original feature space dimension (𝐷), this metric measures the compactness of the data manifold relative to the embedding space. 𝐼𝐷 𝑅𝐼 𝐷 = 𝐷 A lower 𝑅𝐼 𝐷 suggests that the data resides on a low-dimensional manifold within the high-dimensional feature space. Fisher’s Discriminant Ratio (𝐹 1): To assess linear separability, we compute the maximum Fisher’s discriminant ratio across all
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark N 1.00 0.00 0.22 0.15 -0.40 -0.20 -0.01 0.14 0.14 0.04 -0.32 0.03 -0.42 -0.15 -0.06 -0.39 0.38 -0.46 D 0.00 1.00 0.73 -0.87 -0.27 0.37 -0.23 0.71 -0.74 -0.49 -0.21 -0.01 0.10 0.57 -0.17 -0.12 0.03 -0.10
RER 0.15 -0.87 -0.39 1.00 0.01 -0.35 0.17 -0.43 0.89 0.46 -0.03 -0.05 -0.18 -0.48 0.20 -0.01 0.16 -0.01
ID 0.14 0.71 0.94 -0.43 -0.51 0.15 -0.21 1.00 -0.18 -0.46 -0.52 -0.12 -0.04 0.60 -0.05 -0.33 0.36 -0.40
0.14
0.30
0.03
-0.34
0.07
-0.34
0.34
0.04
-0.02
0.34
-0.54 -0.19
0.12
-0.54
1.00
-0.42
-0.50
0.13
-0.09 -0.20
0.09
0.20
0.06
-0.03 -0.10 -0.01
0.10
-0.18
0.48
0.63
0.86
-0.42
1.00
0.50
0.21
0.08
0.16
N2 -0.15 0.57 0.64 -0.48 -0.25 0.01 0.24 0.60 -0.34 -0.44 -0.61 -0.59 0.60 1.00 -0.23 -0.03 -0.15 -0.07
0.75
Lim Dev
-0.06 -0.17 -0.14
0.20
0.00
0.11
-0.05 -0.05
TabP Dev
-0.39 -0.12 -0.34 -0.01
0.30
-0.00
0.27
Lim TabP
0.38
0.16
-0.41
0.07
Cat Dev
-0.46 -0.10 -0.36 -0.01
0.28
0.05
0.03
0.27
0.28
0.24
0.02
0.02
0.04
-0.23
1.00
0.50
0.37
0.39
-0.33 -0.03
0.21
0.13
-0.23
0.48
-0.03
0.50
1.00
-0.51
0.80
-0.31
0.36
0.30
0.08
-0.17
0.22
-0.42 -0.15
0.37
-0.51
1.00
0.15
-0.40 -0.06
0.16
0.35
-0.09
0.42
0.39
0.80
-0.44
Ca t
Lim
Lim
Dev
N2
0.18
Ta bP
RID
0.11
0.24
(a) 𝛾𝑙𝑎 = 1%
-0.07
0.75
0.50
0.00
sp
0.04
-0.49 -0.45
0.46
-0.11 -0.13
0.15
-0.46
0.34
1.00
0.16
0.09
-0.01 -0.44
0.31
0.24
0.44
0.25
0.46
0.00
RF_f1 -0.32 -0.21 -0.56 -0.03 0.44 0.13 -0.37 -0.52 -0.13 0.16 1.00 0.74 -0.35 -0.61 -0.31 -0.50 0.62 0.26 F1 0.03 -0.01 -0.21 -0.05 0.15 0.21 -0.55 -0.12 -0.06 0.09 0.74 1.00 -0.70 -0.59 -0.49 -0.61 0.36 -0.02
0.25
0.25
N1 -0.42 0.10 0.00 -0.18 0.07 -0.06 0.52 -0.04 -0.17 -0.01 -0.35 -0.70 1.00 0.60 0.40 0.48 -0.12 0.17 0.50
0.50
N2 -0.15 0.57 0.64 -0.48 -0.25 0.01 0.24 0.60 -0.34 -0.44 -0.61 -0.59 0.60 1.00 -0.01 0.18 -0.62 -0.43 Lim Dev
-0.04 -0.44 -0.28
0.44
0.12
-0.12
0.30
-0.29
0.42
0.31
-0.31 -0.49
0.40
-0.01
1.00
0.91
0.11
0.70
TabP Dev
-0.03 -0.37 -0.10
0.43
0.02
-0.17
0.41
-0.12
0.42
0.24
-0.50 -0.61
0.48
0.18
0.91
1.00
-0.16
0.51
-0.44
Lim TabP
-0.11 -0.48 -0.73
0.24
0.33
-0.07 -0.03 -0.72
0.11
0.44
0.62
0.36
-0.12 -0.62
0.11
-0.16
1.00
0.54
1.00
Cat Dev
-0.15 -0.58 -0.65
0.44
0.33
-0.13
0.38
0.46
0.26
-0.02
0.17
0.70
0.51
0.54
1.00
0.75
0.10
-0.61
(b) 𝛾𝑙𝑎 = 5%
-0.43
0.75
Dev
0.39
Cat Dev
-0.01 -0.44
1.00
0.33
Ca t
Lim TabP
0.09
N2
0.86
0.16
Dev
0.63
-0.54
1.00
Dev
0.12
1.00
0.34
Ta bP
0.70
0.70
-0.46
F1 0.03 -0.01 -0.21 -0.05 0.15 0.21 -0.55 -0.12 -0.06 0.09 0.74 1.00 -0.70 -0.59 0.02 -0.23 0.22 -0.09
N
1.00
0.13
0.15
0.33
ID 0.14 0.71 0.94 -0.43 -0.51 0.15 -0.21 1.00 -0.18 -0.46 -0.52 -0.12 -0.04 0.60 -0.29 -0.12 -0.72 -0.61
N1 -0.42 0.10 0.00 -0.18 0.07 -0.06 0.52 -0.04 -0.17 -0.01 -0.35 -0.70 1.00 0.60 0.04 0.48 -0.42 0.42
Dev
-0.02
0.44
Ta bP
0.08
-0.23
Dev
0.02
0.04
sp
0.00
F1
0.06
-0.08 -0.03 -0.02
N1
0.00
0.08
RF _f1
0.20
0.15
f
-0.18
0.15
ID
-0.15 -0.10
Clu s LG R
0.25
-0.14 -0.14
D
0.10
0.04
N
0.19
-0.40
ER
-0.18
RER
Lim Dev TabP Dev
-0.11 -0.13
0.02
RID 0.14 -0.74 -0.24 0.89 -0.06 -0.38 0.07 -0.18 1.00 0.34 -0.13 -0.06 -0.17 -0.34 0.42 0.42 0.11 0.38
RF_f1 -0.32 -0.21 -0.56 -0.03 0.44 0.13 -0.37 -0.52 -0.13 0.16 1.00 0.74 -0.35 -0.61 0.02 0.13 -0.17 0.35 0.25
N1 -0.42 0.10 0.00 -0.18 0.07 -0.06 0.52 -0.04 -0.17 -0.01 -0.35 -0.70 1.00 0.60 0.08 0.44 -0.54 0.48 N2 -0.15 0.57 0.64 -0.48 -0.25 0.01 0.24 0.60 -0.34 -0.44 -0.61 -0.59 0.60 1.00 -0.02 0.13 -0.19 0.18
0.46
0.12
Lim
F1 0.03 -0.01 -0.21 -0.05 0.15 0.21 -0.55 -0.12 -0.06 0.09 0.74 1.00 -0.70 -0.59 0.02 -0.23 0.34 -0.18
-0.49 -0.45
-0.25
Lim
RF_f1 -0.32 -0.21 -0.56 -0.03 0.44 0.13 -0.37 -0.52 -0.13 0.16 1.00 0.74 -0.35 -0.61 0.06 0.04 -0.02 0.10
0.04
0.07
Ta bP
sp
0.15
LGR -0.01 -0.23 -0.17 0.17 -0.01 -0.44 1.00 -0.21 0.07 0.15 -0.37 -0.55 0.52 0.24 0.30 0.41 -0.03 0.10 0.25
RID 0.14 -0.74 -0.24 0.89 -0.06 -0.38 0.07 -0.18 1.00 0.34 -0.13 -0.06 -0.17 -0.34 0.28 -0.03 0.30 -0.06 0.00
0.44
Clus -0.20 0.37 0.15 -0.35 -0.08 1.00 -0.44 0.15 -0.38 -0.13 0.13 0.21 -0.06 0.01 -0.12 -0.17 -0.07 -0.13
0.50
sp
-0.01
-0.08 -0.01 -0.51 -0.06 -0.11
F1
0.04
1.00
N1
-0.02
0.01
RID
0.00
-0.40 -0.27 -0.57
RF _f1
-0.01 -0.44
f
f
0.09
D 0.00 1.00 0.73 -0.87 -0.27 0.37 -0.23 0.71 -0.74 -0.49 -0.21 -0.01 0.10 0.57 -0.44 -0.37 -0.48 -0.58
RER 0.15 -0.87 -0.39 1.00 0.01 -0.35 0.17 -0.43 0.89 0.46 -0.03 -0.05 -0.18 -0.48 0.44 0.43 0.24 0.44
0.75
0.28
ID
0.16
-0.41
Clu s LG R
1.00
0.30
N
0.34
0.00
Dev
-0.46
-0.25
Dev
0.15
0.07
Ta bP
-0.11 -0.13
0.15
Ca t
0.46
0.44
Lim
-0.49 -0.45
-0.08 -0.01 -0.51 -0.06 -0.11
N1
0.04
1.00
LGR -0.01 -0.23 -0.17 0.17 -0.01 -0.44 1.00 -0.21 0.07 0.15 -0.37 -0.55 0.52 0.24 -0.05 0.27 -0.31 0.15 0.25
RID 0.14 -0.74 -0.24 0.89 -0.06 -0.38 0.07 -0.18 1.00 0.34 -0.13 -0.06 -0.17 -0.34 0.00 -0.03 0.11 -0.10 sp
0.01
Clus -0.20 0.37 0.15 -0.35 -0.08 1.00 -0.44 0.15 -0.38 -0.13 0.13 0.21 -0.06 0.01 0.11 -0.00 0.07 0.05
0.50
LGR -0.01 -0.23 -0.17 0.17 -0.01 -0.44 1.00 -0.21 0.07 0.15 -0.37 -0.55 0.52 0.24 -0.18 0.08 -0.34 0.06 ID 0.14 0.71 0.94 -0.43 -0.51 0.15 -0.21 1.00 -0.18 -0.46 -0.52 -0.12 -0.04 0.60 0.20 -0.08 0.34 -0.03
-0.40 -0.27 -0.57
N2
Clus -0.20 0.37 0.15 -0.35 -0.08 1.00 -0.44 0.15 -0.38 -0.13 0.13 0.21 -0.06 0.01 0.25 0.15 0.07 0.20
f
ER 0.22 0.73 1.00 -0.39 -0.57 0.15 -0.17 0.94 -0.24 -0.45 -0.56 -0.21 0.00 0.64 -0.28 -0.10 -0.73 -0.65
D
0.75
0.09
Dev
-0.34
Lim
0.15
Ta bP
-0.25 -0.10
sp
0.07
F1
0.15
ID
0.44
RID
-0.08 -0.01 -0.51 -0.06 -0.11
RF _f1
1.00
f
0.01
Clu s LG R
-0.40 -0.27 -0.57
D
f
RER
RER 0.15 -0.87 -0.39 1.00 0.01 -0.35 0.17 -0.43 0.89 0.46 -0.03 -0.05 -0.18 -0.48 -0.15 -0.14 0.03 -0.20
N 1.00 0.00 0.22 0.15 -0.40 -0.20 -0.01 0.14 0.14 0.04 -0.32 0.03 -0.42 -0.15 -0.04 -0.03 -0.11 -0.15 1.00
ER 0.22 0.73 1.00 -0.39 -0.57 0.15 -0.17 0.94 -0.24 -0.45 -0.56 -0.21 0.00 0.64 -0.14 -0.34 0.27 -0.36
ER
1.00
ER
D 0.00 1.00 0.73 -0.87 -0.27 0.37 -0.23 0.71 -0.74 -0.49 -0.21 -0.01 0.10 0.57 0.19 0.04 0.14 0.13 ER 0.22 0.73 1.00 -0.39 -0.57 0.15 -0.17 0.94 -0.24 -0.45 -0.56 -0.21 0.00 0.64 0.10 -0.14 0.30 -0.09
RER
N 1.00 0.00 0.22 0.15 -0.40 -0.20 -0.01 0.14 0.14 0.04 -0.32 0.03 -0.42 -0.15 -0.18 -0.40 0.39 -0.50
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
(c) 𝛾𝑙𝑎 = 100%
Figure 12: Heatmap of spearman correlations between dataset meta-features (denoted by symbols) and model performance under varying supervision ratios. Red (high) indicates positive correlation, while blue (low) indicates negative. Δ𝐴−𝐵 denotes the performance gap between model A and model B. Table 21: Spearman correlation (𝜌) of meta-features across different model comparisons on 𝛾𝑙𝑎 = 1% of all tabular datasets. Meta-Feature Sample Size Dimensionality Avg. Feat. Sim. Effective Rank Eff. Rank Ratio Intrinsic Dim. Fisher’s Ratio Borderline Pts. Intra/Inter Ratio Anom. Cluster. degree Local/Global Outlier Ratio
LimiX - DevNet
TabPFN - DevNet
LimiX - TabPFN
CatB - DevNet
−0.18∗ 0.19∗ −0.10 0.10 −0.15 0.20∗ 0.02 0.08 −0.02 0.25∗∗ −0.18∗
−0.40∗∗∗ 0.04 0.15 −0.14 −0.14 −0.08 −0.23∗∗ 0.44∗∗∗ 0.13 0.15 0.08
0.39∗∗∗ 0.14 −0.34∗∗∗ 0.30∗∗∗ 0.03 0.34∗∗∗ 0.34∗∗∗ −0.54∗∗∗ −0.19∗ 0.07 −0.34∗∗∗
−0.50∗∗∗ 0.13 0.09 −0.09 −0.20∗ −0.03 −0.18∗ 0.48∗∗∗ 0.18∗ 0.20∗ 0.06
Note: ∗ 𝑝 < 0.05,∗∗ 𝑝 < 0.01,∗∗∗ 𝑝 < 0.001. Non-significant entries are unstarred.
individual feature dimensions. (𝜇 0,𝑗 − 𝜇 1,𝑗 ) 2 𝐹 1 = max 𝑗=1...𝐷 𝜎 2 + 𝜎 2 0,𝑗 1,𝑗 2 are the mean and variance of the 𝑗-th feature where 𝜇𝑐,𝑗 and 𝜎𝑐,𝑗 for class 𝑐. A higher 𝐹 1 indicates that at least one feature provides strong discriminative power between normal and abnormal classes. Fraction of Borderline Points (𝑁 1): This complexity measure quantifies the intricacy of the decision boundary. We construct a Minimum Spanning Tree (MST) over the entire dataset and calculate the fraction of edges connecting samples from different classes. ∑︁ 1 𝑁1 = I(𝑦𝑖 ≠ 𝑦 𝑗 ) 𝑁 (𝑥𝑖 ,𝑥 𝑗 ) ∈MST
High values of 𝑁 1 imply that normal and abnormal samples are interleaved, indicating a complex classification boundary. Ratio of Intra/Inter Class Nearest Neighbor Distances (𝑁 2): This metric evaluates the clustering structure by comparing within-class compactness to between-class separation. Í𝑁 𝑑 (𝑥𝑖 , 𝑁 𝑁 same (𝑥𝑖 )) 𝑁 2 = Í𝑖=1 𝑁 𝑖=1 𝑑 (𝑥𝑖 , 𝑁 𝑁 diff (𝑥𝑖 )) where 𝑑 (·) is the Euclidean distance, and 𝑁 𝑁 same (𝑁 𝑁 diff ) denotes the nearest neighbor from the same (different) class. Lower values indicate distinct, compact clusters.
Table 22: Spearman correlation (𝜌) of meta-features across different model comparisons on 𝛾𝑙𝑎 = 5% of all tabular datasets. Meta-Feature Sample Size Sparsity Ratio Avg. Feat. Sim. Effective Rank Eff. Rank Ratio Intrinsic Dim. ID Ratio Fisher’s Ratio Borderline Pts. Intra/Inter Ratio RF’s F1 Local/Global Outlier Ratio
LimiX - DevNet
TabPFN - DevNet
LimiX - TabPFN
CatB - DevNet
−0.06 0.24∗∗ 0.00 −0.14 0.20∗ −0.05 0.28∗∗ 0.02 0.04 −0.23∗ 0.02 −0.05
−0.39∗∗∗ 0.21∗ 0.30∗∗ −0.34∗∗∗ −0.01 −0.33∗∗∗ −0.03 −0.23∗ 0.48∗∗∗ −0.03 0.13 0.27∗∗
0.38∗∗∗ 0.08 −0.41∗∗∗ 0.27∗∗ 0.16 0.36∗∗∗ 0.30∗∗∗ 0.22∗ −0.42∗∗∗ −0.15 −0.17 −0.31∗∗∗
−0.46∗∗∗ 0.16 0.28∗∗ −0.36∗∗∗ −0.01 −0.40∗∗∗ −0.06 −0.09 0.42∗∗∗ −0.07 0.35∗∗∗ 0.15
Note: ∗ 𝑝 < 0.05,∗∗ 𝑝 < 0.01,∗∗∗ 𝑝 < 0.001. Non-significant entries are unstarred.
Anomaly Clustering Degree (𝐶𝑙𝑢𝑠): This feature measures the tendency of anomalies to form dense micro-clusters rather than being randomly scattered. It is calculated as the average local clustering coefficient of the subgraph induced by abnormal samples in the 𝑘-nearest neighbor graph. ∑︁ 1 𝐶𝑙𝑢𝑠 = 𝐶 (𝑥) 𝑁𝑎𝑛𝑜𝑚 𝑥 ∈ S𝑎𝑛𝑜𝑚
where 𝐶 (𝑥) is the local clustering coefficient of node 𝑥. High values suggest that anomalies exhibit local structural patterns.
G.2
Correlation with Model Performance
To visualize the relationship between these meta-features and model performance, we present the Spearman correlation heatmaps across varying supervision levels in Figure 12 (𝛾𝑙𝑎 = 1% in a, 𝛾𝑙𝑎 = 5% in b, and 𝛾𝑙𝑎 = 100% in c). These visualizations highlight how the importance of specific dataset characteristics shifts as supervision signals become more abundant. While the main text focuses on the meta-feature analysis under full supervision (𝛾𝑙𝑎 = 100%), we also investigate how these correlations evolve under limited supervision. Table 21 and Table 22 present the Spearman correlations between meta-features and model performance gaps at 𝛾𝑙𝑎 = 1% and 𝛾𝑙𝑎 = 5%, respectively. These results highlight the consistency of certain metafeatures (e.g., dimensionality) in influencing model superiority even when labeled data is scarce.
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 23: The average ± standard deviation and ranking of AUCPR under different 𝛾𝑙𝑎 (=1%, 5%, 10%, 25%, 50%, 100%) settings on tabular datasets. Label Ratio (𝛾𝑙𝑎 )
Type
Model
Key Mechanism 1%
5%
10%
25%
50%
100%
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Aug. Data Aug. Pseudo Labeling Diffusion DAE
0.481 ± 0.257 (9) 0.379 ± 0.251 (13) 0.408 ± 0.259 (11) 0.308 ± 0.267 (17) 0.314 ± 0.261 (16) 0.555 ± 0.314 (2) 0.524 ± 0.296 (5) 0.547 ± 0.309 (3) 0.378 ± 0.249 (14) 0.518 ± 0.293 (7) 0.275 ± 0.227 (18) 0.320 ± 0.240 (15)
0.623 ± 0.261 (7) 0.497 ± 0.276 (14) 0.538 ± 0.275 (11) 0.306 ± 0.263 (18) 0.315 ± 0.255 (17) 0.635 ± 0.304 (6) 0.646 ± 0.286 (4) 0.644 ± 0.295 (5) 0.498 ± 0.258 (13) 0.580 ± 0.297 (10) 0.295 ± 0.233 (19) 0.323 ± 0.243 (16)
0.688 ± 0.257 (5) 0.568 ± 0.283 (15) 0.602 ± 0.276 (12) 0.310 ± 0.269 (18) 0.316 ± 0.260 (17) 0.666 ± 0.302 (8) 0.691 ± 0.281 (4) 0.676 ± 0.292 (6) 0.570 ± 0.272 (14) 0.603 ± 0.296 (11) 0.310 ± 0.245 (19) 0.331 ± 0.244 (16)
0.767 ± 0.248 (4) 0.670 ± 0.278 (14) 0.686 ± 0.278 (13) 0.306 ± 0.267 (19) 0.319 ± 0.254 (18) 0.695 ± 0.293 (10) 0.741 ± 0.275 (6) 0.704 ± 0.286 (9) 0.689 ± 0.267 (11) 0.626 ± 0.295 (15) 0.349 ± 0.260 (16) 0.347 ± 0.251 (17)
0.816 ± 0.239 (4) 0.752 ± 0.263 (11) 0.739 ± 0.272 (12) 0.321 ± 0.280 (19) 0.335 ± 0.264 (18) 0.713 ± 0.287 (14) 0.772 ± 0.261 (8) 0.721 ± 0.281 (13) 0.762 ± 0.262 (10) 0.647 ± 0.297 (15) 0.416 ± 0.276 (16) 0.386 ± 0.269 (17)
0.864 ± 0.219 (4) 0.813 ± 0.243 (10) 0.779 ± 0.265 (12) 0.351 ± 0.286 (19) 0.382 ± 0.291 (18) 0.722 ± 0.284 (14) 0.786 ± 0.259 (11) 0.729 ± 0.281 (13) 0.826 ± 0.240 (9) 0.671 ± 0.298 (15) 0.489 ± 0.299 (17) 0.517 ± 0.325 (16)
Supervised
XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
GBDT GBDT Deep (Sup.) Deep (Sup.) Deep (Sup.) Found. Model Found. Model
0.241 ± 0.266 (19) 0.521 ± 0.281 (6) 0.515 ± 0.322 (8) 0.468 ± 0.298 (10) 0.405 ± 0.276 (12) 0.528 ± 0.260 (4) 0.585 ± 0.269 (1)
0.495 ± 0.311 (15) 0.650 ± 0.276 (3) 0.621 ± 0.294 (8) 0.615 ± 0.288 (9) 0.517 ± 0.274 (12) 0.710 ± 0.251 (2) 0.716 ± 0.253 (1)
0.622 ± 0.291 (10) 0.715 ± 0.271 (3) 0.664 ± 0.291 (9) 0.675 ± 0.274 (7) 0.596 ± 0.280 (13) 0.770 ± 0.243 (2) 0.771 ± 0.246 (1)
0.738 ± 0.268 (7) 0.783 ± 0.257 (3) 0.729 ± 0.283 (8) 0.744 ± 0.268 (5) 0.686 ± 0.282 (12) 0.831 ± 0.223 (2) 0.833 ± 0.230 (1)
0.795 ± 0.264 (5) 0.842 ± 0.227 (3) 0.782 ± 0.269 (7) 0.787 ± 0.251 (6) 0.768 ± 0.265 (9) 0.864 ± 0.212 (2) 0.869 ± 0.204 (1)
0.858 ± 0.231 (5) 0.883 ± 0.201 (3) 0.853 ± 0.237 (6) 0.830 ± 0.243 (8) 0.851 ± 0.239 (7) 0.900 ± 0.189 (2) 0.903 ± 0.181 (1)
Median Unsup.
ECOD
Probabilistic
0.365 ± 0.264
Table 24: The average ± standard deviation and ranking of AUCPR under different 𝑁𝑙𝑎 (=1, 3, 5, 10, 15, 20, 50) settings on tabular datasets. Type
Label Count (𝑁𝑙𝑎 )
Model 1
3
5
10
15
20
50
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
0.385 ± 0.244 (11) 0.332 ± 0.223 (15) 0.353 ± 0.235 (13) 0.322 ± 0.274 (16) 0.310 ± 0.257 (18) 0.517 ± 0.315 (1) 0.469 ± 0.295 (5) 0.508 ± 0.309 (3) 0.346 ± 0.233 (14) 0.479 ± 0.290 (4) 0.277 ± 0.227 (19) 0.321 ± 0.242 (17)
0.518 ± 0.255 (10) 0.423 ± 0.255 (14) 0.482 ± 0.261 (12) 0.284 ± 0.244 (19) 0.313 ± 0.257 (17) 0.576 ± 0.317 (5) 0.568 ± 0.303 (6) 0.578 ± 0.310 (3) 0.420 ± 0.232 (15) 0.547 ± 0.298 (7) 0.294 ± 0.238 (18) 0.323 ± 0.243 (16)
0.580 ± 0.258 (8) 0.474 ± 0.266 (15) 0.537 ± 0.274 (12) 0.282 ± 0.245 (19) 0.314 ± 0.258 (17) 0.610 ± 0.311 (4) 0.608 ± 0.300 (6) 0.608 ± 0.304 (5) 0.484 ± 0.250 (14) 0.571 ± 0.301 (9) 0.289 ± 0.230 (18) 0.328 ± 0.245 (16)
0.657 ± 0.260 (5) 0.556 ± 0.283 (14) 0.607 ± 0.286 (11) 0.282 ± 0.249 (19) 0.319 ± 0.258 (18) 0.648 ± 0.304 (7) 0.662 ± 0.292 (4) 0.650 ± 0.300 (6) 0.554 ± 0.273 (15) 0.603 ± 0.299 (12) 0.333 ± 0.246 (17) 0.337 ± 0.247 (16)
0.697 ± 0.260 (4) 0.604 ± 0.290 (14) 0.643 ± 0.291 (10) 0.299 ± 0.259 (19) 0.318 ± 0.258 (18) 0.666 ± 0.302 (9) 0.687 ± 0.289 (5) 0.667 ± 0.295 (8) 0.602 ± 0.281 (15) 0.621 ± 0.305 (12) 0.354 ± 0.259 (16) 0.349 ± 0.257 (17)
0.723 ± 0.253 (4) 0.637 ± 0.287 (13) 0.667 ± 0.290 (10) 0.288 ± 0.249 (19) 0.323 ± 0.265 (18) 0.676 ± 0.296 (8) 0.703 ± 0.285 (5) 0.675 ± 0.292 (9) 0.634 ± 0.284 (14) 0.623 ± 0.299 (15) 0.384 ± 0.281 (16) 0.364 ± 0.270 (17)
0.791 ± 0.241 (4) 0.715 ± 0.270 (10) 0.714 ± 0.280 (11) 0.331 ± 0.274 (19) 0.341 ± 0.275 (18) 0.697 ± 0.291 (14) 0.744 ± 0.275 (7) 0.701 ± 0.286 (13) 0.719 ± 0.280 (9) 0.634 ± 0.301 (15) 0.418 ± 0.289 (16) 0.411 ± 0.306 (17)
Supervised
XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
0.430 ± 0.266 (8) 0.457 ± 0.274 (6) 0.448 ± 0.310 (7) 0.404 ± 0.279 (10) 0.370 ± 0.275 (12) 0.415 ± 0.255 (9) 0.515 ± 0.261 (2)
0.505 ± 0.259 (11) 0.577 ± 0.272 (4) 0.541 ± 0.317 (8) 0.527 ± 0.307 (9) 0.458 ± 0.270 (13) 0.600 ± 0.269 (2) 0.641 ± 0.263 (1)
0.558 ± 0.259 (11) 0.635 ± 0.272 (3) 0.567 ± 0.315 (10) 0.582 ± 0.306 (7) 0.513 ± 0.273 (13) 0.667 ± 0.266 (2) 0.688 ± 0.261 (1)
0.630 ± 0.266 (9) 0.689 ± 0.269 (3) 0.612 ± 0.321 (10) 0.643 ± 0.297 (8) 0.569 ± 0.281 (13) 0.735 ± 0.253 (2) 0.745 ± 0.251 (1)
0.670 ± 0.264 (7) 0.729 ± 0.263 (3) 0.629 ± 0.327 (11) 0.681 ± 0.291 (6) 0.613 ± 0.288 (13) 0.771 ± 0.244 (2) 0.778 ± 0.241 (1)
0.695 ± 0.261 (7) 0.750 ± 0.254 (3) 0.653 ± 0.323 (11) 0.701 ± 0.285 (6) 0.641 ± 0.292 (12) 0.793 ± 0.238 (2) 0.799 ± 0.235 (1)
0.766 ± 0.253 (5) 0.809 ± 0.236 (3) 0.705 ± 0.315 (12) 0.758 ± 0.265 (6) 0.725 ± 0.283 (8) 0.842 ± 0.221 (2) 0.844 ± 0.218 (1)
Median Unsup.
ECOD
0.365 ± 0.264
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 25: The average ± standard deviation and ranking of AUCPR under different 𝛾𝑙𝑎 (=1%, 5%, 10%, 25%, 50%, 100%) settings on image datasets (preprocessed by Vision Transformer). Label Ratio (𝛾𝑙𝑎 )
Type
Model
Key Mechanism 1%
5%
10%
25%
50%
100%
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Aug. Data Aug. Pseudo Labeling Diffusion DAE
0.388 ± 0.242 (9) 0.268 ± 0.168 (18) 0.288 ± 0.198 (13) 0.332 ± 0.246 (10) 0.277 ± 0.282 (17) 0.459 ± 0.296 (5) 0.449 ± 0.290 (7) 0.468 ± 0.300 (3) 0.285 ± 0.203 (15) 0.520 ± 0.307 (1) 0.292 ± 0.231 (12) 0.321 ± 0.225 (11)
0.571 ± 0.301 (9) 0.395 ± 0.226 (14) 0.481 ± 0.264 (11) 0.318 ± 0.220 (16) 0.279 ± 0.282 (18) 0.629 ± 0.307 (4) 0.601 ± 0.304 (7) 0.624 ± 0.307 (5) 0.457 ± 0.271 (12) 0.636 ± 0.311 (2) 0.287 ± 0.235 (17) 0.325 ± 0.226 (15)
0.633 ± 0.307 (9) 0.474 ± 0.249 (14) 0.592 ± 0.292 (11) 0.320 ± 0.227 (16) 0.279 ± 0.280 (18) 0.694 ± 0.312 (2) 0.666 ± 0.304 (6) 0.688 ± 0.309 (4) 0.548 ± 0.287 (12) 0.689 ± 0.314 (3) 0.292 ± 0.238 (17) 0.332 ± 0.231 (15)
0.711 ± 0.311 (9) 0.606 ± 0.278 (13) 0.711 ± 0.307 (10) 0.321 ± 0.230 (17) 0.283 ± 0.283 (18) 0.756 ± 0.308 (2) 0.750 ± 0.305 (4) 0.752 ± 0.303 (3) 0.659 ± 0.294 (12) 0.740 ± 0.314 (6) 0.338 ± 0.269 (16) 0.358 ± 0.247 (15)
0.759 ± 0.307 (9) 0.722 ± 0.291 (13) 0.773 ± 0.305 (6) 0.330 ± 0.236 (17) 0.285 ± 0.285 (18) 0.782 ± 0.299 (4) 0.790 ± 0.296 (3) 0.781 ± 0.296 (5) 0.747 ± 0.292 (10) 0.764 ± 0.313 (8) 0.383 ± 0.291 (16) 0.404 ± 0.274 (15)
0.798 ± 0.301 (10) 0.826 ± 0.280 (3) 0.816 ± 0.298 (5) 0.366 ± 0.272 (17) 0.311 ± 0.295 (18) 0.798 ± 0.291 (9) 0.812 ± 0.284 (6) 0.795 ± 0.290 (11) 0.820 ± 0.277 (4) 0.776 ± 0.311 (13) 0.449 ± 0.317 (16) 0.592 ± 0.349 (15)
Supervised
XGBoost CatBoost TabM TabR-S TabPFN LimiX
GBDT GBDT Deep (Sup.) Deep (Sup.) Found. Model Found. Model
0.283 ± 0.241 (16) 0.448 ± 0.275 (8) 0.461 ± 0.293 (4) 0.287 ± 0.210 (14) 0.458 ± 0.266 (6) 0.518 ± 0.301 (2)
0.534 ± 0.284 (10) 0.579 ± 0.316 (8) 0.605 ± 0.306 (6) 0.411 ± 0.242 (13) 0.639 ± 0.305 (1) 0.633 ± 0.309 (3)
0.596 ± 0.294 (10) 0.635 ± 0.322 (8) 0.660 ± 0.313 (7) 0.488 ± 0.269 (13) 0.702 ± 0.307 (1) 0.679 ± 0.306 (5)
0.680 ± 0.304 (11) 0.713 ± 0.321 (8) 0.720 ± 0.325 (7) 0.600 ± 0.291 (14) 0.774 ± 0.294 (1) 0.749 ± 0.295 (5)
0.744 ± 0.307 (11) 0.764 ± 0.313 (7) 0.741 ± 0.334 (12) 0.696 ± 0.296 (14) 0.817 ± 0.274 (1) 0.797 ± 0.285 (2)
0.795 ± 0.303 (12) 0.802 ± 0.301 (8) 0.762 ± 0.340 (14) 0.808 ± 0.280 (7) 0.849 ± 0.258 (1) 0.831 ± 0.278 (2)
Median Unsup.
IForest
Isolation-based
0.360 ± 0.270
Table 26: The average ± standard deviation and ranking of AUCPR under different 𝑁𝑙𝑎 (=1, 3, 5, 10, 15, 20, 50) settings on image datasets (preprocessed by Vision Transformer). Label Count (𝑁𝑙𝑎 )
Type
Model 1
3
5
10
15
20
50
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
0.300 ± 0.198 (12) 0.217 ± 0.147 (16) 0.213 ± 0.175 (17) nan ± nan (18) 0.282 ± 0.284 (14) 0.333 ± 0.249 (7) 0.328 ± 0.242 (9) 0.349 ± 0.260 (5) 0.362 ± 0.166 (4) 0.421 ± 0.290 (2) 0.318 ± 0.262 (11) 0.319 ± 0.223 (10)
0.404 ± 0.248 (9) 0.269 ± 0.168 (18) 0.298 ± 0.203 (15) 0.315 ± 0.228 (13) 0.275 ± 0.281 (17) 0.467 ± 0.288 (4) 0.448 ± 0.274 (7) 0.479 ± 0.292 (3) 0.384 ± 0.193 (10) 0.522 ± 0.300 (2) 0.304 ± 0.259 (14) 0.322 ± 0.225 (12)
0.471 ± 0.271 (9) 0.302 ± 0.180 (17) 0.350 ± 0.215 (12) 0.327 ± 0.231 (13) 0.278 ± 0.281 (18) 0.541 ± 0.299 (4) 0.515 ± 0.295 (7) 0.541 ± 0.304 (3) 0.457 ± 0.208 (11) 0.557 ± 0.308 (2) 0.305 ± 0.250 (16) 0.323 ± 0.226 (14)
0.557 ± 0.297 (9) 0.376 ± 0.213 (14) 0.447 ± 0.243 (12) 0.324 ± 0.227 (16) 0.278 ± 0.282 (18) 0.625 ± 0.309 (3) 0.599 ± 0.303 (7) 0.619 ± 0.310 (5) 0.497 ± 0.238 (11) 0.627 ± 0.313 (2) 0.366 ± 0.258 (15) 0.324 ± 0.225 (17)
0.598 ± 0.305 (9) 0.423 ± 0.232 (15) 0.517 ± 0.265 (12) 0.324 ± 0.217 (17) 0.284 ± 0.283 (18) 0.670 ± 0.315 (1) 0.642 ± 0.307 (6) 0.660 ± 0.312 (4) 0.537 ± 0.265 (11) 0.667 ± 0.317 (3) 0.459 ± 0.284 (13) 0.329 ± 0.230 (16)
0.631 ± 0.308 (8) 0.455 ± 0.239 (14) 0.564 ± 0.277 (11) 0.322 ± 0.228 (16) 0.286 ± 0.286 (17) 0.694 ± 0.314 (2) 0.670 ± 0.307 (6) 0.686 ± 0.311 (4) 0.541 ± 0.282 (12) 0.688 ± 0.316 (3) 0.184 ± 0.180 (18) 0.334 ± 0.231 (15)
0.710 ± 0.316 (8) 0.599 ± 0.279 (13) 0.707 ± 0.308 (10) 0.345 ± 0.237 (16) 0.280 ± 0.281 (17) 0.756 ± 0.314 (2) 0.755 ± 0.312 (3) 0.753 ± 0.310 (5) 0.658 ± 0.298 (12) 0.743 ± 0.318 (6) 0.187 ± 0.197 (18) 0.364 ± 0.254 (15)
Supervised
XGBoost CatBoost TabM TabR-S TabPFN LimiX
0.231 ± 0.151 (15) 0.403 ± 0.269 (3) 0.344 ± 0.256 (6) 0.284 ± 0.241 (13) 0.331 ± 0.234 (8) 0.422 ± 0.282 (1)
0.379 ± 0.233 (11) 0.443 ± 0.274 (8) 0.462 ± 0.281 (5) 0.285 ± 0.204 (16) 0.450 ± 0.262 (6) 0.524 ± 0.298 (1)
0.460 ± 0.269 (10) 0.491 ± 0.287 (8) 0.522 ± 0.291 (6) 0.308 ± 0.206 (15) 0.536 ± 0.287 (5) 0.565 ± 0.304 (1)
0.550 ± 0.293 (10) 0.566 ± 0.310 (8) 0.602 ± 0.304 (6) 0.385 ± 0.235 (13) 0.627 ± 0.306 (1) 0.622 ± 0.310 (4)
0.590 ± 0.301 (10) 0.598 ± 0.320 (8) 0.641 ± 0.311 (7) 0.435 ± 0.253 (14) 0.670 ± 0.310 (2) 0.653 ± 0.310 (5)
0.622 ± 0.304 (10) 0.624 ± 0.324 (9) 0.665 ± 0.314 (7) 0.469 ± 0.261 (13) 0.697 ± 0.309 (1) 0.678 ± 0.307 (5)
0.704 ± 0.315 (11) 0.708 ± 0.327 (9) 0.728 ± 0.328 (7) 0.583 ± 0.293 (14) 0.772 ± 0.305 (1) 0.754 ± 0.302 (4)
Median Unsup.
IForest
0.360 ± 0.270
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 27: The average ± standard deviation and ranking of AUCPR under different 𝑁𝑙𝑎 (=1, 3, 5, 10, 15, 20, 50) settings on text datasets (preprocessed by RoBERTa). Label Count (𝑁𝑙𝑎 )
Type
Model 1
3
5
10
15
20
50
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
0.132 ± 0.073 (9) 0.088 ± 0.038 (18) 0.108 ± 0.057 (15) 0.065 ± 0.025 (19) 0.109 ± 0.054 (14) 0.171 ± 0.114 (5) 0.146 ± 0.085 (8) 0.184 ± 0.112 (2) 0.127 ± 0.069 (12) 0.176 ± 0.098 (3) 0.156 ± 0.129 (7) 0.172 ± 0.102 (4)
0.180 ± 0.105 (9) 0.112 ± 0.064 (17) 0.146 ± 0.080 (13) 0.077 ± 0.040 (19) 0.111 ± 0.054 (18) 0.306 ± 0.184 (1) 0.242 ± 0.145 (6) 0.294 ± 0.185 (3) 0.193 ± 0.109 (8) 0.284 ± 0.162 (5) 0.145 ± 0.092 (14) 0.173 ± 0.102 (11)
0.221 ± 0.125 (8) 0.130 ± 0.074 (16) 0.172 ± 0.088 (13) 0.076 ± 0.048 (19) 0.110 ± 0.052 (18) 0.352 ± 0.193 (3) 0.289 ± 0.173 (7) 0.340 ± 0.188 (4) 0.217 ± 0.126 (10) 0.338 ± 0.184 (5) 0.159 ± 0.097 (14) 0.173 ± 0.101 (12)
0.282 ± 0.147 (9) 0.171 ± 0.092 (16) 0.231 ± 0.124 (12) 0.076 ± 0.051 (19) 0.111 ± 0.055 (17) 0.436 ± 0.201 (2) 0.365 ± 0.191 (7) 0.433 ± 0.187 (3) 0.284 ± 0.134 (8) 0.419 ± 0.181 (4) 0.181 ± 0.129 (14) 0.175 ± 0.102 (15)
0.334 ± 0.166 (9) 0.214 ± 0.128 (14) 0.301 ± 0.158 (11) 0.085 ± 0.079 (19) 0.111 ± 0.056 (17) 0.501 ± 0.199 (2) 0.441 ± 0.187 (7) 0.501 ± 0.195 (1) 0.331 ± 0.164 (10) 0.487 ± 0.191 (3) 0.187 ± 0.156 (15) 0.178 ± 0.104 (16)
0.376 ± 0.178 (9) 0.247 ± 0.146 (14) 0.353 ± 0.173 (11) 0.091 ± 0.085 (19) 0.106 ± 0.051 (17) 0.544 ± 0.193 (1) 0.495 ± 0.192 (7) 0.543 ± 0.190 (2) 0.379 ± 0.176 (8) 0.536 ± 0.195 (3) 0.100 ± 0.118 (18) 0.180 ± 0.105 (15)
0.535 ± 0.183 (9) 0.436 ± 0.217 (13) 0.565 ± 0.193 (8) 0.117 ± 0.093 (18) 0.107 ± 0.056 (19) 0.645 ± 0.173 (2) 0.625 ± 0.176 (6) 0.644 ± 0.168 (3) 0.523 ± 0.186 (10) 0.630 ± 0.183 (5) 0.117 ± 0.171 (17) 0.194 ± 0.109 (16)
Supervised
XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
0.106 ± 0.052 (16) 0.120 ± 0.054 (13) 0.105 ± 0.076 (17) 0.163 ± 0.092 (6) 0.131 ± 0.081 (10) 0.129 ± 0.069 (11) 0.187 ± 0.105 (1)
0.176 ± 0.104 (10) 0.156 ± 0.095 (12) 0.125 ± 0.117 (16) 0.292 ± 0.161 (4) 0.133 ± 0.086 (15) 0.226 ± 0.150 (7) 0.303 ± 0.180 (2)
0.221 ± 0.126 (9) 0.195 ± 0.116 (11) 0.114 ± 0.141 (17) 0.352 ± 0.168 (2) 0.152 ± 0.090 (15) 0.304 ± 0.159 (6) 0.356 ± 0.195 (1)
0.281 ± 0.149 (10) 0.251 ± 0.155 (11) 0.094 ± 0.140 (18) 0.437 ± 0.185 (1) 0.209 ± 0.131 (13) 0.401 ± 0.173 (6) 0.411 ± 0.197 (5)
0.335 ± 0.165 (8) 0.276 ± 0.174 (12) 0.100 ± 0.150 (18) 0.487 ± 0.197 (4) 0.261 ± 0.152 (13) 0.461 ± 0.185 (6) 0.472 ± 0.195 (5)
0.375 ± 0.183 (10) 0.323 ± 0.191 (12) 0.116 ± 0.179 (16) 0.528 ± 0.198 (4) 0.287 ± 0.171 (13) 0.513 ± 0.186 (5) 0.511 ± 0.201 (6)
0.512 ± 0.190 (11) 0.512 ± 0.202 (12) 0.207 ± 0.215 (15) 0.614 ± 0.210 (7) 0.411 ± 0.193 (14) 0.667 ± 0.164 (1) 0.631 ± 0.184 (4)
Median Unsup.
VAE
0.093 ± 0.036
Table 28: The average ± standard deviation and ranking of AUCPR under different 𝛾𝑙𝑎 (=1%, 5%, 10%, 25%, 50%, 100%) settings on text datasets (preprocessed by RoBERTa). Type
Model
Label Ratio (𝛾𝑙𝑎 )
Key Mechanism 1%
5%
10%
25%
50%
100%
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Aug. Data Aug. Pseudo Labeling Diffusion DAE
0.166 ± 0.087 (9) 0.105 ± 0.046 (16) 0.134 ± 0.058 (13) 0.072 ± 0.023 (19) 0.110 ± 0.055 (15) 0.272 ± 0.172 (2) 0.221 ± 0.132 (6) 0.261 ± 0.154 (4) 0.164 ± 0.076 (10) 0.253 ± 0.149 (5) 0.087 ± 0.041 (18) 0.172 ± 0.102 (8)
0.284 ± 0.147 (8) 0.171 ± 0.100 (15) 0.248 ± 0.132 (11) 0.072 ± 0.021 (19) 0.113 ± 0.056 (16) 0.430 ± 0.219 (1) 0.379 ± 0.203 (7) 0.424 ± 0.215 (2) 0.277 ± 0.131 (10) 0.404 ± 0.207 (5) 0.101 ± 0.075 (17) 0.173 ± 0.103 (14)
0.366 ± 0.172 (8) 0.225 ± 0.125 (14) 0.344 ± 0.186 (11) 0.070 ± 0.021 (19) 0.107 ± 0.050 (16) 0.512 ± 0.219 (2) 0.461 ± 0.222 (7) 0.513 ± 0.214 (1) 0.356 ± 0.170 (10) 0.498 ± 0.209 (5) 0.104 ± 0.084 (17) 0.176 ± 0.106 (15)
0.485 ± 0.184 (10) 0.344 ± 0.151 (14) 0.499 ± 0.203 (8) 0.070 ± 0.020 (19) 0.112 ± 0.055 (17) 0.607 ± 0.190 (3) 0.576 ± 0.195 (5) 0.610 ± 0.182 (2) 0.485 ± 0.155 (9) 0.574 ± 0.199 (6) 0.104 ± 0.072 (18) 0.183 ± 0.110 (15)
0.571 ± 0.174 (9) 0.503 ± 0.159 (13) 0.618 ± 0.187 (7) 0.071 ± 0.021 (19) 0.114 ± 0.056 (18) 0.661 ± 0.159 (3) 0.646 ± 0.167 (6) 0.668 ± 0.155 (2) 0.582 ± 0.172 (8) 0.647 ± 0.165 (5) 0.135 ± 0.103 (17) 0.194 ± 0.121 (16)
0.669 ± 0.165 (11) 0.726 ± 0.146 (4) 0.741 ± 0.141 (3) 0.073 ± 0.026 (19) 0.136 ± 0.081 (18) 0.705 ± 0.152 (6) 0.717 ± 0.148 (5) 0.702 ± 0.147 (8) 0.703 ± 0.135 (7) 0.685 ± 0.163 (10) 0.144 ± 0.105 (17) 0.224 ± 0.146 (16)
Supervised
XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
GBDT GBDT Deep (Sup.) Deep (Sup.) Deep (Sup.) Found. Model Found. Model
0.151 ± 0.079 (11) 0.143 ± 0.062 (12) 0.094 ± 0.077 (17) 0.278 ± 0.163 (1) 0.130 ± 0.074 (14) 0.184 ± 0.088 (7) 0.266 ± 0.155 (3)
0.280 ± 0.145 (9) 0.224 ± 0.123 (12) 0.084 ± 0.123 (18) 0.407 ± 0.212 (4) 0.189 ± 0.093 (13) 0.394 ± 0.188 (6) 0.420 ± 0.198 (3)
0.361 ± 0.179 (9) 0.317 ± 0.165 (12) 0.086 ± 0.140 (18) 0.478 ± 0.209 (6) 0.249 ± 0.111 (13) 0.503 ± 0.191 (4) 0.506 ± 0.213 (3)
0.478 ± 0.184 (11) 0.466 ± 0.204 (12) 0.177 ± 0.157 (16) 0.542 ± 0.181 (7) 0.365 ± 0.125 (13) 0.622 ± 0.163 (1) 0.580 ± 0.187 (4)
0.557 ± 0.174 (11) 0.568 ± 0.195 (10) 0.257 ± 0.198 (15) 0.520 ± 0.208 (12) 0.483 ± 0.145 (14) 0.703 ± 0.137 (1) 0.650 ± 0.176 (4)
0.667 ± 0.167 (12) 0.686 ± 0.164 (9) 0.286 ± 0.215 (15) 0.552 ± 0.229 (14) 0.660 ± 0.159 (13) 0.782 ± 0.118 (1) 0.746 ± 0.135 (2)
Median Unsup.
VAE
Reconstruction
0.093 ± 0.036
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 29: The average ± standard deviation and ranking of AUCROC under different 𝛾𝑙𝑎 (=1%, 5%, 10%, 25%, 50%, 100%) settings on tabular datasets. Type
Model
Label Ratio (𝛾𝐿 )
Key Mechanism 1%
5%
10%
25%
50%
100%
Semi-supervised
XGBOD DeepSAD REPEN AABiGAN GANomaly DevNet FEAWAD PReNet RoSAS DualMGAN SOEL-NTL AnoDDAE
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Augmentation Data Augmentation Pseudo Labeling Diffusion DAE
0.807 ± 0.154 (5) 0.755 ± 0.159 (10) 0.755 ± 0.157 (9) 0.663 ± 0.219 (16) 0.677 ± 0.205 (15) 0.802 ± 0.174 (6) 0.747 ± 0.182 (12) 0.796 ± 0.177 (7) 0.633 ± 0.182 (18) 0.809 ± 0.178 (4) 0.663 ± 0.199 (17) 0.748 ± 0.181 (11)
0.869 ± 0.135 (4) 0.818 ± 0.144 (12) 0.828 ± 0.138 (11) 0.654 ± 0.229 (19) 0.683 ± 0.199 (17) 0.859 ± 0.151 (5) 0.833 ± 0.154 (10) 0.856 ± 0.154 (6) 0.712 ± 0.173 (16) 0.848 ± 0.156 (8) 0.678 ± 0.197 (18) 0.750 ± 0.181 (15)
0.903 ± 0.116 (4) 0.854 ± 0.133 (13) 0.865 ± 0.131 (10) 0.662 ± 0.222 (19) 0.682 ± 0.202 (17) 0.881 ± 0.138 (6) 0.876 ± 0.134 (7) 0.882 ± 0.141 (5) 0.764 ± 0.161 (15) 0.862 ± 0.146 (12) 0.678 ± 0.205 (18) 0.755 ± 0.179 (16)
0.935 ± 0.096 (4) 0.900 ± 0.117 (9) 0.898 ± 0.125 (10) 0.675 ± 0.226 (19) 0.686 ± 0.197 (18) 0.898 ± 0.128 (11) 0.908 ± 0.120 (7) 0.898 ± 0.132 (12) 0.853 ± 0.137 (15) 0.876 ± 0.139 (14) 0.719 ± 0.188 (17) 0.764 ± 0.178 (16)
0.954 ± 0.072 (4) 0.928 ± 0.099 (7) 0.916 ± 0.118 (11) 0.681 ± 0.215 (19) 0.703 ± 0.194 (18) 0.911 ± 0.106 (13) 0.927 ± 0.101 (8) 0.913 ± 0.107 (12) 0.901 ± 0.125 (14) 0.887 ± 0.125 (15) 0.752 ± 0.191 (17) 0.785 ± 0.179 (16)
0.968 ± 0.057 (4) 0.949 ± 0.083 (8) 0.930 ± 0.109 (12) 0.698 ± 0.225 (19) 0.729 ± 0.194 (18) 0.917 ± 0.103 (14) 0.940 ± 0.088 (11) 0.918 ± 0.106 (13) 0.945 ± 0.084 (9) 0.899 ± 0.116 (15) 0.790 ± 0.187 (17) 0.825 ± 0.182 (16)
Supervised
XGB CatB FTTransformer TabMCls TabR-S TabPFN LimiX
GBDT GBDT Deep Learning Deep Learning Deep Learning Found. Model Found. Model
0.603 ± 0.157 (19) 0.839 ± 0.157 (3) 0.775 ± 0.200 (8) 0.746 ± 0.194 (13) 0.742 ± 0.187 (14) 0.855 ± 0.147 (2) 0.872 ± 0.141 (1)
0.788 ± 0.177 (14) 0.892 ± 0.129 (3) 0.853 ± 0.155 (7) 0.835 ± 0.160 (9) 0.801 ± 0.169 (13) 0.916 ± 0.112 (2) 0.919 ± 0.112 (1)
0.862 ± 0.144 (11) 0.915 ± 0.112 (3) 0.870 ± 0.154 (9) 0.875 ± 0.140 (8) 0.843 ± 0.154 (14) 0.935 ± 0.097 (2) 0.938 ± 0.090 (1)
0.914 ± 0.113 (5) 0.938 ± 0.098 (3) 0.907 ± 0.130 (8) 0.910 ± 0.126 (6) 0.887 ± 0.138 (13) 0.953 ± 0.082 (2) 0.956 ± 0.076 (1)
0.937 ± 0.103 (5) 0.956 ± 0.080 (3) 0.926 ± 0.133 (9) 0.931 ± 0.101 (6) 0.925 ± 0.102 (10) 0.965 ± 0.063 (2) 0.968 ± 0.059 (1)
0.967 ± 0.060 (5) 0.971 ± 0.054 (3) 0.954 ± 0.088 (6) 0.942 ± 0.110 (10) 0.952 ± 0.084 (7) 0.976 ± 0.051 (2) 0.979 ± 0.042 (1)
Median Unsup.
VAE
Reconstruction
0.740 ± 0.194
Table 30: The average ± standard deviation and ranking of AUCROC under different 𝑁𝑙𝑎 (=1, 3, 5, 10, 15, 20, 50) settings on tabular datasets. Type
Label Count (𝑁𝑙𝑎 )
Model 1
3
5
10
15
20
50
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
0.737 ± 0.179 (9) 0.726 ± 0.161 (12) 0.714 ± 0.164 (13) 0.690 ± 0.209 (15) 0.672 ± 0.204 (17) 0.759 ± 0.197 (5) 0.699 ± 0.198 (14) 0.757 ± 0.196 (6) 0.616 ± 0.175 (19) 0.784 ± 0.186 (4) 0.667 ± 0.203 (18) 0.748 ± 0.181 (8)
0.807 ± 0.153 (6) 0.770 ± 0.159 (11) 0.778 ± 0.163 (10) 0.640 ± 0.221 (19) 0.682 ± 0.200 (16) 0.809 ± 0.176 (5) 0.767 ± 0.190 (12) 0.803 ± 0.180 (7) 0.666 ± 0.176 (17) 0.820 ± 0.169 (4) 0.662 ± 0.210 (18) 0.750 ± 0.180 (15)
0.840 ± 0.143 (4) 0.794 ± 0.154 (11) 0.807 ± 0.160 (9) 0.648 ± 0.216 (19) 0.680 ± 0.205 (17) 0.836 ± 0.157 (6) 0.794 ± 0.182 (12) 0.831 ± 0.160 (7) 0.708 ± 0.182 (16) 0.840 ± 0.156 (5) 0.673 ± 0.202 (18) 0.751 ± 0.181 (15)
0.880 ± 0.127 (4) 0.831 ± 0.151 (12) 0.847 ± 0.153 (9) 0.634 ± 0.220 (19) 0.687 ± 0.199 (18) 0.864 ± 0.140 (5) 0.834 ± 0.164 (11) 0.858 ± 0.146 (8) 0.753 ± 0.186 (15) 0.860 ± 0.141 (7) 0.699 ± 0.204 (17) 0.753 ± 0.182 (16)
0.900 ± 0.114 (4) 0.849 ± 0.146 (12) 0.868 ± 0.144 (9) 0.661 ± 0.215 (19) 0.686 ± 0.201 (18) 0.877 ± 0.133 (6) 0.859 ± 0.152 (11) 0.872 ± 0.136 (7) 0.786 ± 0.179 (15) 0.868 ± 0.137 (8) 0.710 ± 0.197 (17) 0.759 ± 0.181 (16)
0.911 ± 0.107 (4) 0.863 ± 0.141 (12) 0.879 ± 0.138 (8) 0.660 ± 0.211 (19) 0.682 ± 0.206 (18) 0.885 ± 0.126 (6) 0.873 ± 0.146 (11) 0.881 ± 0.130 (7) 0.805 ± 0.176 (15) 0.873 ± 0.134 (10) 0.728 ± 0.205 (17) 0.764 ± 0.183 (16)
0.939 ± 0.087 (4) 0.901 ± 0.120 (10) 0.900 ± 0.129 (11) 0.680 ± 0.221 (19) 0.696 ± 0.205 (18) 0.903 ± 0.112 (8) 0.909 ± 0.119 (6) 0.902 ± 0.115 (9) 0.865 ± 0.153 (14) 0.879 ± 0.132 (13) 0.748 ± 0.203 (17) 0.778 ± 0.186 (16)
Supervised
XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
0.749 ± 0.183 (7) 0.800 ± 0.174 (2) 0.735 ± 0.204 (10) 0.683 ± 0.219 (16) 0.729 ± 0.194 (11) 0.788 ± 0.179 (3) 0.831 ± 0.164 (1)
0.786 ± 0.165 (8) 0.848 ± 0.160 (3) 0.780 ± 0.210 (9) 0.765 ± 0.200 (13) 0.757 ± 0.180 (14) 0.861 ± 0.148 (2) 0.871 ± 0.149 (1)
0.820 ± 0.148 (8) 0.872 ± 0.144 (3) 0.788 ± 0.208 (13) 0.802 ± 0.187 (10) 0.777 ± 0.180 (14) 0.885 ± 0.132 (2) 0.893 ± 0.128 (1)
0.861 ± 0.139 (6) 0.892 ± 0.131 (3) 0.802 ± 0.219 (14) 0.843 ± 0.159 (10) 0.805 ± 0.172 (13) 0.911 ± 0.115 (2) 0.915 ± 0.110 (1)
0.882 ± 0.127 (5) 0.908 ± 0.115 (3) 0.810 ± 0.226 (14) 0.865 ± 0.149 (10) 0.821 ± 0.170 (13) 0.927 ± 0.102 (2) 0.932 ± 0.094 (1)
0.894 ± 0.120 (5) 0.916 ± 0.113 (3) 0.823 ± 0.223 (14) 0.875 ± 0.140 (9) 0.837 ± 0.168 (13) 0.935 ± 0.095 (2) 0.938 ± 0.089 (1)
0.926 ± 0.102 (5) 0.940 ± 0.091 (3) 0.864 ± 0.194 (15) 0.909 ± 0.116 (7) 0.879 ± 0.146 (12) 0.952 ± 0.079 (2) 0.956 ± 0.070 (1)
Median Unsup.
VAE
0.740 ± 0.192
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 31: The average ± standard deviation and ranking of AUCROC under different 𝛾𝑙𝑎 (=1%, 5%, 10%, 25%, 50%, 100%) settings on image datasets (preprocessed by Vision Transformer). Label Ratio (𝛾𝑙𝑎 )
Type
Model
Key Mechanism 1%
5%
10%
25%
50%
100%
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Aug. Data Aug. Pseudo-Labeling Diffusion DAE
0.752 ± 0.168 (6) 0.688 ± 0.139 (16) 0.695 ± 0.119 (15) 0.745 ± 0.172 (7) 0.718 ± 0.169 (12) 0.739 ± 0.165 (9) 0.704 ± 0.166 (13) 0.744 ± 0.169 (8) 0.633 ± 0.168 (18) 0.805 ± 0.169 (3) 0.733 ± 0.153 (11) 0.734 ± 0.173 (10)
0.844 ± 0.152 (6) 0.750 ± 0.148 (14) 0.799 ± 0.136 (11) 0.728 ± 0.171 (16) 0.721 ± 0.169 (17) 0.846 ± 0.147 (4) 0.805 ± 0.153 (10) 0.839 ± 0.148 (8) 0.765 ± 0.154 (13) 0.861 ± 0.149 (2) 0.720 ± 0.159 (18) 0.738 ± 0.170 (15)
0.873 ± 0.144 (6) 0.786 ± 0.143 (14) 0.847 ± 0.138 (11) 0.736 ± 0.169 (16) 0.723 ± 0.166 (17) 0.887 ± 0.134 (3) 0.848 ± 0.140 (10) 0.878 ± 0.136 (5) 0.825 ± 0.147 (12) 0.889 ± 0.136 (2) 0.717 ± 0.159 (18) 0.743 ± 0.169 (15)
0.906 ± 0.130 (7) 0.847 ± 0.135 (14) 0.895 ± 0.133 (11) 0.743 ± 0.169 (16) 0.724 ± 0.167 (18) 0.920 ± 0.120 (2) 0.908 ± 0.120 (6) 0.912 ± 0.122 (5) 0.879 ± 0.130 (12) 0.916 ± 0.125 (3) 0.732 ± 0.171 (17) 0.762 ± 0.162 (15)
0.925 ± 0.118 (7) 0.900 ± 0.125 (13) 0.921 ± 0.125 (10) 0.742 ± 0.169 (17) 0.734 ± 0.161 (18) 0.934 ± 0.109 (4) 0.935 ± 0.107 (3) 0.929 ± 0.111 (5) 0.914 ± 0.121 (11) 0.927 ± 0.119 (6) 0.754 ± 0.168 (16) 0.794 ± 0.159 (15)
0.940 ± 0.110 (6) 0.938 ± 0.115 (9) 0.939 ± 0.116 (7) 0.766 ± 0.171 (17) 0.752 ± 0.161 (18) 0.941 ± 0.104 (5) 0.946 ± 0.100 (3) 0.936 ± 0.107 (11) 0.942 ± 0.110 (4) 0.934 ± 0.116 (13) 0.768 ± 0.171 (16) 0.866 ± 0.164 (15)
Supervised
XGBoost CatBoost TabM TabR-S TabPFN LimiX
GBDT GBDT Deep (Sup.) Deep (Sup.) Found. Model Found. Model
0.665 ± 0.185 (17) 0.813 ± 0.156 (1) 0.762 ± 0.177 (5) 0.699 ± 0.146 (14) 0.811 ± 0.159 (2) 0.805 ± 0.176 (4)
0.833 ± 0.145 (9) 0.846 ± 0.162 (5) 0.841 ± 0.159 (7) 0.770 ± 0.150 (12) 0.873 ± 0.142 (1) 0.856 ± 0.155 (3)
0.862 ± 0.140 (9) 0.863 ± 0.160 (8) 0.871 ± 0.148 (7) 0.801 ± 0.157 (13) 0.900 ± 0.129 (1) 0.881 ± 0.139 (4)
0.898 ± 0.133 (10) 0.899 ± 0.145 (9) 0.900 ± 0.141 (8) 0.859 ± 0.143 (13) 0.929 ± 0.112 (1) 0.915 ± 0.116 (4)
0.922 ± 0.121 (9) 0.923 ± 0.128 (8) 0.912 ± 0.140 (12) 0.900 ± 0.128 (14) 0.946 ± 0.099 (1) 0.938 ± 0.104 (2)
0.938 ± 0.115 (10) 0.939 ± 0.115 (8) 0.921 ± 0.136 (14) 0.936 ± 0.107 (12) 0.957 ± 0.091 (1) 0.952 ± 0.099 (2)
Meadian Unsup.
CBLOF
Cluster-based
0.801 ± 0.172
Table 32: The average ± standard deviation and ranking of AUCROC under different 𝑁𝑙𝑎 (=1, 3, 5, 10, 15, 20, 50) settings on image datasets (preprocessed by Vision Transformer). Label Count (𝑁𝑙𝑎 )
Type
Model 1
3
5
10
15
20
50
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
0.708 ± 0.130 (8) 0.662 ± 0.130 (12) 0.639 ± 0.102 (15) nan ± nan (18) 0.719 ± 0.169 (7) 0.662 ± 0.147 (13) 0.624 ± 0.138 (16) 0.682 ± 0.157 (10) 0.595 ± 0.122 (17) 0.758 ± 0.165 (4) 0.732 ± 0.164 (6) 0.732 ± 0.172 (5)
0.767 ± 0.151 (6) 0.690 ± 0.135 (17) 0.705 ± 0.117 (14) 0.733 ± 0.168 (11) 0.715 ± 0.171 (12) 0.746 ± 0.160 (9) 0.702 ± 0.157 (15) 0.753 ± 0.164 (7) 0.627 ± 0.162 (18) 0.806 ± 0.169 (4) 0.706 ± 0.165 (13) 0.735 ± 0.173 (10)
0.802 ± 0.152 (5) 0.709 ± 0.136 (17) 0.738 ± 0.124 (11) 0.735 ± 0.167 (13) 0.718 ± 0.169 (14) 0.795 ± 0.154 (8) 0.748 ± 0.160 (10) 0.790 ± 0.159 (9) 0.714 ± 0.152 (15) 0.814 ± 0.165 (4) 0.675 ± 0.161 (18) 0.736 ± 0.171 (12)
0.841 ± 0.152 (6) 0.743 ± 0.142 (14) 0.788 ± 0.129 (11) 0.733 ± 0.170 (16) 0.718 ± 0.170 (17) 0.845 ± 0.149 (4) 0.804 ± 0.151 (10) 0.836 ± 0.151 (9) 0.763 ± 0.154 (12) 0.855 ± 0.153 (2) 0.662 ± 0.170 (18) 0.737 ± 0.170 (15)
0.861 ± 0.148 (6) 0.765 ± 0.141 (14) 0.818 ± 0.136 (11) 0.732 ± 0.168 (16) 0.721 ± 0.173 (17) 0.872 ± 0.144 (3) 0.833 ± 0.145 (10) 0.861 ± 0.146 (5) 0.805 ± 0.151 (12) 0.878 ± 0.147 (2) 0.720 ± 0.168 (18) 0.741 ± 0.170 (15)
0.872 ± 0.145 (6) 0.782 ± 0.140 (14) 0.838 ± 0.136 (11) 0.735 ± 0.168 (16) 0.717 ± 0.179 (17) 0.888 ± 0.137 (2) 0.855 ± 0.139 (10) 0.877 ± 0.141 (5) 0.819 ± 0.145 (12) 0.888 ± 0.141 (3) 0.526 ± 0.170 (18) 0.745 ± 0.168 (15)
0.905 ± 0.136 (7) 0.849 ± 0.136 (14) 0.894 ± 0.137 (11) 0.752 ± 0.169 (16) 0.719 ± 0.170 (17) 0.919 ± 0.123 (2) 0.911 ± 0.126 (6) 0.913 ± 0.124 (5) 0.876 ± 0.136 (12) 0.916 ± 0.128 (4) 0.521 ± 0.175 (18) 0.766 ± 0.164 (15)
Supervised
XGBoost CatBoost TabM TabR-S TabPFN LimiX
0.653 ± 0.107 (14) 0.794 ± 0.157 (1) 0.691 ± 0.171 (9) 0.680 ± 0.160 (11) 0.764 ± 0.160 (3) 0.774 ± 0.174 (2)
0.751 ± 0.143 (8) 0.813 ± 0.156 (1) 0.769 ± 0.169 (5) 0.698 ± 0.148 (16) 0.812 ± 0.156 (2) 0.810 ± 0.173 (3)
0.796 ± 0.150 (7) 0.830 ± 0.157 (2) 0.802 ± 0.165 (6) 0.714 ± 0.144 (16) 0.838 ± 0.155 (1) 0.827 ± 0.168 (3)
0.839 ± 0.149 (8) 0.844 ± 0.164 (5) 0.839 ± 0.159 (7) 0.757 ± 0.151 (13) 0.868 ± 0.146 (1) 0.852 ± 0.158 (3)
0.857 ± 0.146 (8) 0.852 ± 0.165 (9) 0.860 ± 0.156 (7) 0.779 ± 0.154 (13) 0.886 ± 0.141 (1) 0.867 ± 0.152 (4)
0.870 ± 0.144 (8) 0.859 ± 0.163 (9) 0.872 ± 0.152 (7) 0.797 ± 0.151 (13) 0.897 ± 0.134 (1) 0.879 ± 0.144 (4)
0.903 ± 0.136 (8) 0.896 ± 0.150 (10) 0.903 ± 0.144 (9) 0.856 ± 0.146 (13) 0.927 ± 0.119 (1) 0.917 ± 0.121 (3)
Median Unsup.
IForest
0.763 ± 0.174
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 33: The average ± standard deviation and ranking of AUCROC under different 𝛾𝑙𝑎 (=1%, 5%, 10%, 25%, 50%, 100%) settings on text datasets (preprocessed by RoBERTa). Label Ratio (𝛾𝑙𝑎 )
Type
Model
Key Mechanism 1%
5%
10%
25%
50%
100%
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Aug. Data Aug. Pseudo-Labeling
0.698 ± 0.088 (8) 0.611 ± 0.086 (16) 0.633 ± 0.072 (13) 0.606 ± 0.079 (17) 0.665 ± 0.121 (11) 0.745 ± 0.122 (5) 0.689 ± 0.101 (9) 0.749 ± 0.110 (4) 0.664 ± 0.109 (12) 0.770 ± 0.094 (3) 0.612 ± 0.101 (15)
0.795 ± 0.098 (7) 0.671 ± 0.098 (13) 0.730 ± 0.091 (12) 0.592 ± 0.078 (17) 0.667 ± 0.126 (14) 0.851 ± 0.110 (4) 0.795 ± 0.108 (8) 0.854 ± 0.105 (2) 0.771 ± 0.121 (10) 0.847 ± 0.100 (6) 0.619 ± 0.112 (16)
0.844 ± 0.086 (7) 0.710 ± 0.101 (14) 0.785 ± 0.104 (12) 0.584 ± 0.079 (17) 0.662 ± 0.116 (15) 0.891 ± 0.089 (3) 0.839 ± 0.103 (9) 0.893 ± 0.085 (2) 0.817 ± 0.107 (11) 0.893 ± 0.073 (1) 0.619 ± 0.116 (16)
0.892 ± 0.069 (8) 0.785 ± 0.085 (13) 0.862 ± 0.091 (12) 0.583 ± 0.079 (18) 0.670 ± 0.122 (15) 0.932 ± 0.050 (1) 0.898 ± 0.076 (7) 0.925 ± 0.056 (2) 0.880 ± 0.073 (11) 0.912 ± 0.082 (5) 0.627 ± 0.119 (17)
0.917 ± 0.053 (9) 0.856 ± 0.063 (13) 0.911 ± 0.066 (11) 0.587 ± 0.082 (18) 0.674 ± 0.112 (16) 0.945 ± 0.038 (2) 0.925 ± 0.060 (6) 0.944 ± 0.038 (3) 0.912 ± 0.054 (10) 0.938 ± 0.042 (5) 0.653 ± 0.128 (17)
0.942 ± 0.040 (9) 0.925 ± 0.051 (12) 0.944 ± 0.045 (8) 0.588 ± 0.089 (18) 0.700 ± 0.119 (16) 0.956 ± 0.030 (3) 0.946 ± 0.044 (7) 0.952 ± 0.032 (4) 0.941 ± 0.048 (10) 0.948 ± 0.035 (5) 0.671 ± 0.125 (17)
Supervised
XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
GBDT GBDT Deep (Sup.) Deep (Sup.) Deep (Sup.) Found. Model Found. Model
0.676 ± 0.086 (10) 0.719 ± 0.082 (7) 0.557 ± 0.153 (18) 0.774 ± 0.105 (2) 0.621 ± 0.106 (14) 0.734 ± 0.095 (6) 0.793 ± 0.098 (1)
0.790 ± 0.096 (9) 0.761 ± 0.087 (11) 0.395 ± 0.231 (18) 0.851 ± 0.097 (3) 0.665 ± 0.112 (15) 0.847 ± 0.093 (5) 0.859 ± 0.093 (1)
0.840 ± 0.083 (8) 0.820 ± 0.089 (10) 0.359 ± 0.255 (18) 0.882 ± 0.078 (6) 0.726 ± 0.122 (13) 0.888 ± 0.076 (5) 0.890 ± 0.085 (4)
0.891 ± 0.068 (9) 0.888 ± 0.074 (10) 0.651 ± 0.224 (16) 0.909 ± 0.055 (6) 0.782 ± 0.119 (14) 0.923 ± 0.056 (3) 0.918 ± 0.066 (4)
0.920 ± 0.047 (8) 0.921 ± 0.052 (7) 0.732 ± 0.189 (15) 0.910 ± 0.062 (12) 0.856 ± 0.068 (14) 0.946 ± 0.039 (1) 0.941 ± 0.042 (4)
0.940 ± 0.039 (11) 0.947 ± 0.037 (6) 0.774 ± 0.128 (15) 0.917 ± 0.063 (14) 0.920 ± 0.051 (13) 0.963 ± 0.029 (1) 0.958 ± 0.032 (2)
Median Unsup.
AutoEncoder
Reconstruction
0.708 ± 0.118
Table 34: The average ± standard deviation and ranking of AUCROC under different 𝑁𝑙𝑎 (=1, 3, 5, 10, 15, 20, 50) settings on text datasets (preprocessed by RoBERTa). Label Count (𝑁𝑙𝑎 )
Type
Model 1
3
5
10
15
20
50
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL DDAE
0.651 ± 0.085 (11) 0.582 ± 0.074 (18) 0.597 ± 0.073 (17) 0.537 ± 0.120 (19) 0.665 ± 0.126 (9) 0.655 ± 0.094 (10) 0.635 ± 0.088 (12) 0.691 ± 0.097 (4) 0.600 ± 0.094 (16) 0.722 ± 0.097 (2) 0.676 ± 0.119 (7) 0.721 ± 0.123 (3)
0.719 ± 0.079 (9) 0.612 ± 0.078 (16) 0.640 ± 0.069 (15) 0.577 ± 0.099 (19) 0.668 ± 0.121 (14) 0.770 ± 0.106 (5) 0.694 ± 0.090 (11) 0.771 ± 0.110 (4) 0.684 ± 0.093 (12) 0.785 ± 0.095 (3) 0.675 ± 0.105 (13) 0.721 ± 0.122 (8)
0.755 ± 0.088 (8) 0.636 ± 0.084 (16) 0.675 ± 0.067 (13) 0.544 ± 0.103 (18) 0.666 ± 0.124 (14) 0.819 ± 0.093 (3) 0.732 ± 0.108 (10) 0.814 ± 0.089 (5) 0.722 ± 0.096 (11) 0.817 ± 0.085 (4) 0.660 ± 0.097 (15) 0.722 ± 0.122 (12)
0.802 ± 0.074 (8) 0.682 ± 0.083 (15) 0.730 ± 0.076 (12) 0.530 ± 0.118 (18) 0.663 ± 0.128 (16) 0.874 ± 0.072 (1) 0.792 ± 0.092 (9) 0.864 ± 0.074 (3) 0.784 ± 0.081 (11) 0.859 ± 0.076 (5) 0.644 ± 0.092 (17) 0.724 ± 0.123 (13)
0.833 ± 0.073 (9) 0.708 ± 0.089 (15) 0.769 ± 0.077 (12) 0.539 ± 0.121 (18) 0.667 ± 0.128 (16) 0.901 ± 0.058 (1) 0.836 ± 0.082 (7) 0.894 ± 0.060 (2) 0.821 ± 0.079 (10) 0.885 ± 0.065 (5) 0.658 ± 0.095 (17) 0.727 ± 0.124 (13)
0.854 ± 0.070 (9) 0.738 ± 0.090 (13) 0.801 ± 0.080 (12) 0.548 ± 0.125 (17) 0.658 ± 0.123 (16) 0.915 ± 0.052 (1) 0.863 ± 0.074 (7) 0.906 ± 0.057 (3) 0.850 ± 0.068 (10) 0.906 ± 0.060 (2) 0.526 ± 0.103 (18) 0.730 ± 0.125 (14)
0.909 ± 0.054 (8) 0.823 ± 0.093 (13) 0.891 ± 0.063 (11) 0.610 ± 0.110 (17) 0.656 ± 0.127 (16) 0.943 ± 0.038 (1) 0.916 ± 0.054 (7) 0.938 ± 0.043 (4) 0.890 ± 0.060 (12) 0.936 ± 0.042 (5) 0.525 ± 0.128 (19) 0.740 ± 0.129 (15)
Supervised
XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
0.605 ± 0.065 (15) 0.679 ± 0.090 (6) 0.606 ± 0.103 (14) 0.685 ± 0.089 (5) 0.629 ± 0.116 (13) 0.675 ± 0.107 (8) 0.738 ± 0.109 (1)
0.703 ± 0.078 (10) 0.722 ± 0.091 (7) 0.603 ± 0.157 (17) 0.785 ± 0.092 (2) 0.593 ± 0.106 (18) 0.745 ± 0.107 (6) 0.802 ± 0.098 (1)
0.751 ± 0.081 (9) 0.760 ± 0.094 (7) 0.543 ± 0.184 (19) 0.834 ± 0.075 (1) 0.636 ± 0.101 (17) 0.798 ± 0.086 (6) 0.828 ± 0.087 (2)
0.804 ± 0.073 (7) 0.787 ± 0.089 (10) 0.436 ± 0.216 (19) 0.870 ± 0.069 (2) 0.684 ± 0.123 (14) 0.847 ± 0.070 (6) 0.862 ± 0.076 (4)
0.835 ± 0.069 (8) 0.793 ± 0.089 (11) 0.397 ± 0.234 (19) 0.888 ± 0.065 (3) 0.710 ± 0.132 (14) 0.874 ± 0.067 (6) 0.886 ± 0.070 (4)
0.854 ± 0.066 (8) 0.821 ± 0.081 (11) 0.394 ± 0.274 (19) 0.905 ± 0.058 (5) 0.722 ± 0.133 (15) 0.896 ± 0.058 (6) 0.906 ± 0.062 (4)
0.905 ± 0.050 (9) 0.900 ± 0.059 (10) 0.539 ± 0.341 (18) 0.928 ± 0.053 (6) 0.813 ± 0.104 (14) 0.940 ± 0.039 (2) 0.939 ± 0.041 (3)
Median Unsup.
VAE
0.659 ± 0.100
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 35: The average AUCPR performance on individual datasets under the 𝛾𝑙𝑎 = 1% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.467 (7) 0.542 (14) 0.467 (12) 0.414 (7) 0.393 (7) 0.454 (8) 0.837 (9) 0.381 (9) 0.652 (10) 0.520 (1) 0.056 (4) 0.209 (4) 0.900 (4) 0.645 (9) 0.280 (11) 0.512 (7) 0.607 (9) 0.469 (10) 0.520 (7) 0.456 (9) 0.512 (3) 0.647 (1) 0.702 (8) 0.734 (8) 0.942 (10) 0.792 (9) 0.504 (6) 0.640 (6) 0.021 (9) 0.315 (11) 0.567 (7) 0.187 (13) 0.304 (13) 0.403 (8) 0.078 (9) 0.574 (10) 0.792 (7) 0.141 (8) 0.611 (9) 0.332 (10) 0.428 (5) 0.844 (9) 0.346 (5) 0.537 (6) 0.519 (10) 0.114 (7) 0.260 (2)
0.211 (12) 0.713 (11) 0.510 (3) 0.213 (15) 0.233 (12) 0.359 (13) 0.757 (10) 0.389 (7) 0.713 (9) 0.437 (9) 0.065 (3) 0.206 (5) 0.390 (15) 0.635 (10) 0.328 (9) 0.347 (14) 0.373 (15) 0.210 (13) 0.556 (5) 0.372 (11) 0.444 (11) 0.264 (15) 0.669 (10) 0.399 (13) 0.928 (11) 0.836 (7) 0.582 (4) 0.442 (14) 0.022 (6) 0.175 (16) 0.409 (10) 0.157 (14) 0.467 (8) 0.144 (15) 0.161 (3) 0.572 (11) 0.407 (14) 0.080 (12) 0.150 (15) 0.268 (14) 0.340 (15) 0.822 (10) 0.189 (14) 0.328 (14) 0.400 (16) 0.047 (14) 0.091 (15)
0.269 (10) 0.860 (6) 0.469 (10) 0.269 (13) 0.307 (10) 0.396 (12) 0.933 (6) 0.390 (6) 0.765 (6) 0.297 (14) 0.052 (6) 0.114 (15) 0.662 (9) 0.659 (8) 0.444 (2) 0.352 (13) 0.490 (11) 0.209 (14) 0.601 (4) 0.446 (10) 0.506 (4) 0.337 (11) 0.659 (11) 0.324 (15) 0.972 (4) 0.549 (14) 0.440 (10) 0.576 (9) 0.021 (7) 0.401 (8) 0.665 (4) 0.264 (4) 0.307 (12) 0.223 (14) 0.036 (16) 0.349 (14) 0.375 (15) 0.083 (11) 0.366 (12) 0.325 (11) 0.386 (11) 0.744 (13) 0.217 (10) 0.374 (11) 0.448 (14) 0.081 (12) 0.177 (7)
0.091 (15) 0.116 (16) 0.373 (17) 0.402 (8) 0.073 (16) 0.295 (15) 0.280 (15) 0.374 (10) 0.746 (7) 0.160 (17) 0.054 (5) 0.153 (9) 0.963 (1) 0.520 (15) 0.060 (15) 0.419 (10) 0.356 (16) 0.029 (15) 0.355 (14) 0.124 (16) 0.376 (17) 0.152 (16) 0.574 (15) 0.768 (7) 0.714 (16) 0.229 (16) 0.216 (14) 0.363 (17) 0.024 (5) 0.259 (13) 0.325 (14) 0.125 (15) 0.050 (16) 0.096 (16) 0.038 (15) 0.841 (2) 0.457 (13) 0.053 (15) 0.168 (14) 0.195 (17) 0.317 (16) 0.774 (12) 0.220 (9) 0.507 (7) 0.550 (7) 0.070 (13) 0.071 (17)
0.009 (17) 0.097 (17) 0.493 (7) 0.556 (1) 0.171 (15) 0.240 (16) 0.450 (14) 0.441 (4) 0.860 (2) 0.204 (16) 0.041 (13) 0.164 (7) 0.621 (10) 0.519 (16) 0.122 (14) 0.204 (17) 0.718 (7) 0.026 (17) 0.389 (13) 0.175 (15) 0.456 (10) 0.309 (13) 0.644 (12) 0.470 (12) 0.375 (17) 0.212 (17) 0.106 (15) 0.402 (16) 0.020 (11) 0.284 (12) 0.496 (8) 0.098 (17) 0.309 (11) 0.309 (12) 0.043 (14) 0.418 (13) 0.634 (11) 0.041 (17) 0.132 (16) 0.232 (16) 0.317 (17) 0.848 (8) 0.230 (8) 0.358 (13) 0.351 (17) 0.095 (9) 0.078 (16)
0.644 (4) 0.990 (1) 0.506 (4) 0.530 (4) 0.209 (13) 0.527 (5) 1.000 (1) 0.396 (5) 0.559 (13) 0.372 (11) 0.041 (12) 0.125 (13) 0.696 (8) 0.704 (5) 0.470 (1) 0.559 (4) 0.974 (2) 0.970 (1) 0.512 (8) 0.720 (1) 0.479 (6) 0.328 (12) 0.793 (1) 0.893 (4) 0.971 (5) 0.657 (12) 0.600 (1) 0.696 (3) 0.027 (4) 0.500 (5) 0.809 (3) 0.242 (8) 0.701 (5) 0.435 (7) 0.072 (10) 0.790 (5) 0.980 (1) 0.088 (10) 0.999 (2) 0.376 (3) 0.437 (3) 0.859 (6) 0.260 (6) 0.630 (3) 0.656 (1) 0.141 (1) 0.174 (8)
0.653 (3) 0.989 (2) 0.469 (11) 0.537 (3) 0.611 (2) 0.442 (10) 1.000 (2) 0.367 (11) 0.513 (15) 0.442 (8) 0.047 (8) 0.172 (6) 0.529 (13) 0.670 (7) 0.426 (3) 0.471 (9) 0.688 (8) 0.962 (3) 0.469 (10) 0.666 (3) 0.438 (14) 0.429 (8) 0.759 (7) 0.911 (1) 0.972 (3) 0.798 (8) 0.501 (7) 0.566 (10) 0.018 (15) 0.626 (2) 0.457 (9) 0.276 (2) 0.712 (3) 0.455 (6) 0.101 (6) 0.529 (12) 0.936 (3) 0.163 (5) 0.997 (3) 0.343 (7) 0.411 (7) 0.649 (15) 0.195 (12) 0.452 (10) 0.511 (11) 0.138 (3) 0.143 (11)
0.661 (2) 0.969 (3) 0.489 (8) 0.530 (5) 0.311 (9) 0.480 (7) 1.000 (3) 0.366 (12) 0.552 (14) 0.381 (10) 0.043 (10) 0.127 (11) 0.596 (11) 0.720 (3) 0.420 (4) 0.529 (6) 0.918 (4) 0.967 (2) 0.496 (9) 0.703 (2) 0.474 (9) 0.363 (10) 0.767 (6) 0.903 (3) 0.971 (6) 0.670 (11) 0.600 (3) 0.674 (5) 0.029 (2) 0.514 (4) 0.828 (2) 0.259 (5) 0.707 (4) 0.456 (5) 0.079 (8) 0.737 (8) 0.979 (2) 0.089 (9) 1.000 (1) 0.378 (2) 0.436 (4) 0.785 (11) 0.250 (7) 0.595 (4) 0.601 (5) 0.138 (2) 0.166 (9)
0.132 (14) 0.795 (8) 0.393 (16) 0.235 (14) 0.567 (3) 0.342 (14) 0.918 (7) 0.224 (15) 0.426 (17) 0.368 (12) 0.036 (16) 0.155 (8) 0.404 (14) 0.574 (13) 0.236 (13) 0.313 (15) 0.382 (14) 0.285 (12) 0.319 (16) 0.367 (12) 0.403 (16) 0.489 (6) 0.511 (16) 0.094 (16) 0.905 (13) 0.918 (4) 0.500 (8) 0.438 (15) 0.021 (8) 0.220 (15) 0.235 (16) 0.333 (1) 0.414 (10) 0.389 (9) 0.092 (7) 0.333 (15) 0.650 (10) 0.642 (1) 0.427 (11) 0.367 (4) 0.417 (6) 0.457 (17) 0.193 (13) 0.235 (16) 0.410 (15) 0.082 (11) 0.107 (13)
0.727 (1) 0.750 (10) 0.476 (9) 0.539 (2) 0.473 (5) 0.448 (9) 0.554 (13) 0.309 (13) 0.793 (5) 0.320 (13) 0.050 (7) 0.150 (10) 0.953 (2) 0.699 (6) 0.353 (7) 0.574 (3) 0.849 (6) 0.614 (7) 0.546 (6) 0.604 (4) 0.441 (12) 0.404 (9) 0.783 (2) 0.904 (2) 0.972 (2) 0.621 (13) 0.304 (12) 0.530 (13) 0.018 (13) 0.562 (3) 0.839 (1) 0.216 (10) 0.798 (1) 0.471 (4) 0.065 (12) 0.792 (4) 0.466 (12) 0.065 (13) 0.902 (5) 0.315 (12) 0.388 (10) 0.870 (4) 0.184 (15) 0.757 (1) 0.646 (4) 0.085 (10) 0.148 (10)
0.184 (13) 0.487 (15) 0.536 (1) 0.380 (11) 0.288 (11) 0.398 (11) 0.041 (16) 0.256 (14) 0.863 (1) 0.498 (6) 0.039 (14) 0.293 (3) 0.232 (16) 0.467 (17) 0.034 (16) 0.498 (8) 0.957 (3) 0.754 (5) 0.282 (17) 0.460 (8) 0.475 (8) 0.126 (17) 0.587 (14) 0.480 (11) 0.927 (12) 0.410 (15) 0.312 (11) 0.624 (7) 0.028 (3) 0.239 (14) 0.367 (12) 0.202 (12) 0.676 (6) 0.482 (3) 0.316 (1) 0.177 (16) 0.304 (16) 0.051 (16) 0.209 (13) 0.340 (8) 0.359 (14) 0.649 (16) 0.212 (11) 0.477 (9) 0.499 (12) 0.116 (6) 0.178 (6)
0.010 (16) 0.569 (13) 0.427 (15) 0.002 (17) 0.041 (17) 0.160 (17) 0.004 (17) 0.186 (16) 0.630 (11) 0.466 (7) 0.034 (17) 0.062 (17) 0.040 (17) 0.560 (14) 0.023 (17) 0.408 (11) 0.032 (17) 0.029 (16) 0.455 (12) 0.023 (17) 0.475 (7) 0.543 (5) 0.598 (13) 0.012 (17) 0.886 (14) 0.758 (10) 0.000 (17) 0.622 (8) 0.016 (17) 0.095 (17) 0.025 (17) 0.121 (16) 0.026 (17) 0.034 (17) 0.029 (17) 0.045 (17) 0.029 (17) 0.053 (14) 0.081 (17) 0.233 (15) 0.363 (13) 0.850 (7) 0.347 (4) 0.096 (17) 0.525 (9) 0.022 (17) 0.262 (1)
0.409 (8) 0.753 (9) 0.500 (5) 0.393 (10) 0.380 (8) 0.542 (4) 0.860 (8) 0.561 (2) 0.817 (3) 0.519 (2) 0.043 (11) 0.123 (14) 0.944 (3) 0.595 (12) 0.388 (6) 0.556 (5) 0.855 (5) 0.592 (8) 0.676 (3) 0.364 (13) 0.550 (1) 0.601 (3) 0.673 (9) 0.872 (5) 0.959 (9) 0.958 (1) 0.269 (13) 0.693 (4) 0.021 (10) 0.399 (9) 0.642 (5) 0.219 (9) 0.265 (14) 0.383 (10) 0.069 (11) 0.778 (7) 0.813 (6) 0.193 (4) 0.535 (10) 0.289 (13) 0.384 (12) 0.937 (1) 0.407 (2) 0.692 (2) 0.647 (3) 0.120 (5) 0.259 (3)
0.492 (5) 0.905 (4) 0.461 (13) 0.372 (12) 0.636 (1) 0.543 (3) 1.000 (4) 0.186 (17) 0.445 (16) 0.510 (5) 0.036 (15) 0.125 (12) 0.549 (12) 0.707 (4) 0.271 (12) 0.387 (12) 0.578 (10) 0.676 (6) 0.338 (15) 0.502 (6) 0.426 (15) 0.289 (14) 0.779 (3) 0.603 (10) 0.960 (7) 0.955 (2) 0.500 (9) 0.562 (11) 0.020 (12) 0.670 (1) 0.281 (15) 0.249 (6) 0.463 (9) 0.279 (13) 0.134 (4) 0.662 (9) 0.859 (4) 0.146 (7) 0.792 (6) 0.379 (1) 0.439 (2) 0.718 (14) 0.138 (17) 0.258 (15) 0.530 (8) 0.098 (8) 0.110 (12)
0.231 (11) 0.676 (12) 0.429 (14) 0.194 (16) 0.196 (14) 0.560 (1) 0.575 (12) 0.487 (3) 0.581 (12) 0.247 (15) 0.046 (9) 0.083 (16) 0.873 (5) 0.608 (11) 0.316 (10) 0.299 (16) 0.480 (12) 0.302 (11) 0.459 (11) 0.263 (14) 0.440 (13) 0.435 (7) 0.396 (17) 0.360 (14) 0.754 (15) 0.927 (3) 0.074 (16) 0.531 (12) 0.030 (1) 0.424 (6) 0.377 (11) 0.207 (11) 0.252 (15) 0.318 (11) 0.046 (13) 0.896 (1) 0.711 (9) 0.152 (6) 0.924 (4) 0.355 (6) 0.448 (1) 0.918 (2) 0.168 (16) 0.363 (12) 0.484 (13) 0.046 (15) 0.094 (14)
0.338 (9) 0.884 (5) 0.499 (6) 0.400 (9) 0.534 (4) 0.512 (6) 0.720 (11) 0.389 (8) 0.720 (8) 0.516 (4) 0.070 (2) 0.360 (2) 0.783 (7) 0.757 (2) 0.340 (8) 0.636 (2) 0.414 (13) 0.554 (9) 0.686 (2) 0.478 (7) 0.517 (2) 0.608 (2) 0.772 (4) 0.630 (9) 0.986 (1) 0.849 (6) 0.600 (2) 0.757 (2) 0.018 (16) 0.405 (7) 0.365 (13) 0.267 (3) 0.485 (7) 0.630 (1) 0.103 (5) 0.809 (3) 0.834 (5) 0.477 (3) 0.688 (7) 0.356 (5) 0.406 (8) 0.880 (3) 0.459 (1) 0.483 (8) 0.593 (6) 0.034 (16) 0.196 (5)
0.478 (6) 0.855 (7) 0.515 (2) 0.478 (6) 0.398 (6) 0.556 (2) 0.947 (5) 0.624 (1) 0.809 (4) 0.517 (3) 0.086 (1) 0.374 (1) 0.826 (6) 0.766 (1) 0.409 (5) 0.822 (1) 0.977 (1) 0.830 (4) 0.694 (1) 0.571 (5) 0.497 (5) 0.553 (4) 0.769 (5) 0.819 (6) 0.959 (8) 0.899 (5) 0.516 (5) 0.824 (1) 0.018 (14) 0.388 (10) 0.601 (6) 0.242 (7) 0.778 (2) 0.580 (2) 0.208 (2) 0.788 (6) 0.784 (8) 0.537 (2) 0.623 (8) 0.336 (9) 0.397 (9) 0.865 (5) 0.388 (3) 0.564 (5) 0.649 (2) 0.137 (4) 0.228 (4)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.374 (10) 0.443 (9) 0.507 (8) 0.445 (12) 0.069 (8)
0.267 (11) 0.254 (14) 0.287 (13) 0.393 (16) 0.062 (13)
0.198 (14) 0.253 (16) 0.300 (12) 0.507 (9) 0.068 (10)
0.417 (6) 0.395 (11) 0.218 (15) 0.563 (5) 0.059 (16)
0.114 (16) 0.271 (13) 0.139 (17) 0.687 (1) 0.056 (17)
0.388 (8) 0.513 (4) 0.711 (2) 0.455 (11) 0.078 (2)
0.382 (9) 0.492 (6) 0.713 (1) 0.435 (14) 0.073 (4)
0.438 (5) 0.558 (3) 0.707 (3) 0.437 (13) 0.073 (6)
0.166 (15) 0.231 (17) 0.407 (11) 0.408 (15) 0.079 (1)
0.498 (2) 0.588 (2) 0.687 (5) 0.609 (3) 0.071 (7)
0.260 (12) 0.373 (12) 0.187 (16) 0.525 (6) 0.062 (14)
0.050 (17) 0.399 (10) 0.486 (9) 0.293 (17) 0.062 (15)
0.456 (4) 0.487 (7) 0.440 (10) 0.680 (2) 0.064 (12)
0.400 (7) 0.481 (8) 0.700 (4) 0.490 (10) 0.077 (3)
0.238 (13) 0.253 (15) 0.259 (14) 0.519 (7) 0.066 (11)
0.462 (3) 0.496 (5) 0.554 (7) 0.585 (4) 0.073 (5)
0.571 (1) 0.668 (1) 0.676 (6) 0.516 (8) 0.068 (9)
20news agnews amazon imdb yelp
0.160 (8) 0.207 (8) 0.092 (12) 0.104 (9) 0.172 (10)
0.099 (15) 0.139 (14) 0.067 (14) 0.065 (14) 0.085 (14)
0.139 (11) 0.155 (12) 0.096 (9) 0.083 (11) 0.110 (12)
0.062 (17) 0.087 (17) 0.052 (15) 0.048 (17) 0.053 (15)
0.113 (14) 0.152 (13) 0.051 (16) 0.048 (16) 0.049 (16)
0.204 (2) 0.392 (3) 0.166 (3) 0.174 (2) 0.408 (1)
0.176 (6) 0.309 (6) 0.140 (6) 0.142 (7) 0.304 (6)
0.208 (1) 0.359 (5) 0.172 (2) 0.172 (3) 0.367 (3)
0.154 (10) 0.177 (11) 0.136 (7) 0.152 (5) 0.208 (7)
0.194 (4) 0.371 (4) 0.164 (4) 0.150 (6) 0.334 (4)
0.093 (16) 0.105 (16) 0.049 (17) 0.053 (15) 0.047 (17)
0.130 (13) 0.200 (9) 0.093 (11) 0.107 (8) 0.182 (9)
0.135 (12) 0.191 (10) 0.090 (13) 0.073 (13) 0.123 (11)
0.193 (5) 0.418 (1) 0.182 (1) 0.194 (1) 0.407 (2)
0.159 (9) 0.112 (15) 0.094 (10) 0.082 (12) 0.106 (13)
0.167 (7) 0.244 (7) 0.114 (8) 0.099 (10) 0.204 (8)
0.201 (3) 0.407 (2) 0.146 (5) 0.169 (4) 0.305 (5)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 36: The average AUCPR performance on individual datasets under the 𝛾𝑙𝑎 = 5% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.765 (10) 0.933 (11) 0.544 (5) 0.414 (8) 0.706 (7) 0.709 (8) 0.949 (8) 0.559 (4) 0.842 (9) 0.670 (4) 0.071 (3) 0.435 (3) 0.609 (12) 0.746 (7) 0.394 (12) 0.756 (6) 0.990 (7) 0.731 (10) 0.688 (8) 0.736 (12) 0.600 (6) 0.763 (3) 0.808 (7) 0.786 (11) 0.982 (2) 0.972 (6) 0.504 (6) 0.835 (6) 0.038 (12) 0.620 (10) 0.766 (10) 0.336 (10) 0.627 (8) 0.393 (8) 0.214 (6) 0.533 (13) 0.789 (8) 0.324 (7) 0.864 (11) 0.483 (9) 0.421 (9) 0.938 (8) 0.471 (3) 0.779 (8) 0.657 (8) 0.164 (8) 0.360 (3)
0.678 (12) 0.979 (6) 0.522 (10) 0.213 (16) 0.277 (14) 0.504 (14) 0.805 (12) 0.454 (11) 0.773 (11) 0.632 (5) 0.063 (4) 0.221 (11) 0.417 (15) 0.701 (13) 0.509 (8) 0.552 (14) 0.761 (13) 0.595 (12) 0.718 (6) 0.762 (10) 0.501 (15) 0.386 (14) 0.773 (9) 0.827 (9) 0.969 (11) 0.951 (8) 0.582 (4) 0.586 (13) 0.024 (17) 0.244 (17) 0.579 (12) 0.201 (15) 0.616 (9) 0.157 (16) 0.259 (5) 0.606 (12) 0.409 (14) 0.173 (10) 0.292 (15) 0.327 (15) 0.368 (15) 0.859 (14) 0.244 (15) 0.471 (16) 0.593 (12) 0.095 (15) 0.147 (13)
0.794 (8) 0.984 (4) 0.510 (13) 0.269 (14) 0.460 (11) 0.739 (6) 0.933 (9) 0.496 (8) 0.915 (2) 0.385 (11) 0.059 (5) 0.229 (8) 0.605 (13) 0.740 (8) 0.534 (7) 0.593 (12) 0.799 (11) 0.313 (15) 0.772 (3) 0.688 (13) 0.586 (7) 0.562 (9) 0.758 (11) 0.585 (14) 0.969 (10) 0.547 (14) 0.440 (11) 0.704 (10) 0.059 (2) 0.646 (6) 0.815 (5) 0.427 (7) 0.592 (10) 0.217 (15) 0.113 (15) 0.469 (14) 0.398 (15) 0.109 (11) 0.830 (12) 0.419 (13) 0.377 (12) 0.922 (11) 0.269 (11) 0.651 (10) 0.656 (9) 0.143 (10) 0.221 (8)
0.310 (16) 0.057 (18) 0.439 (18) 0.402 (9) 0.113 (17) 0.280 (16) 0.474 (14) 0.305 (17) 0.677 (15) 0.161 (18) 0.055 (7) 0.079 (18) 0.765 (9) 0.524 (16) 0.030 (18) 0.638 (10) 0.054 (18) 0.070 (16) 0.328 (17) 0.082 (18) 0.418 (18) 0.140 (16) 0.572 (16) 0.601 (13) 0.648 (16) 0.179 (18) 0.216 (15) 0.562 (14) 0.032 (14) 0.198 (18) 0.305 (17) 0.154 (16) 0.076 (18) 0.044 (17) 0.041 (18) 0.773 (8) 0.437 (13) 0.052 (17) 0.070 (18) 0.214 (18) 0.316 (16) 0.943 (6) 0.317 (9) 0.545 (13) 0.510 (15) 0.079 (16) 0.083 (16)
0.010 (18) 0.121 (15) 0.497 (14) 0.556 (1) 0.172 (15) 0.247 (18) 0.452 (15) 0.442 (13) 0.861 (7) 0.209 (17) 0.041 (16) 0.166 (17) 0.628 (11) 0.513 (17) 0.125 (15) 0.201 (18) 0.591 (16) 0.029 (18) 0.443 (16) 0.174 (16) 0.458 (17) 0.364 (15) 0.638 (14) 0.477 (15) 0.275 (17) 0.213 (17) 0.106 (16) 0.403 (18) 0.020 (18) 0.287 (16) 0.505 (13) 0.098 (17) 0.344 (17) 0.314 (12) 0.043 (17) 0.413 (15) 0.633 (10) 0.089 (13) 0.134 (17) 0.231 (16) 0.312 (17) 0.852 (15) 0.251 (14) 0.363 (18) 0.358 (18) 0.067 (17) 0.078 (17)
0.892 (4) 0.995 (2) 0.542 (6) 0.530 (4) 0.285 (13) 0.741 (5) 1.000 (1) 0.468 (10) 0.742 (12) 0.392 (10) 0.048 (14) 0.223 (9) 0.803 (6) 0.732 (9) 0.610 (1) 0.777 (4) 1.000 (1) 0.991 (1) 0.679 (11) 0.923 (4) 0.620 (3) 0.450 (12) 0.813 (5) 0.896 (4) 0.971 (8) 0.659 (12) 0.600 (1) 0.852 (5) 0.054 (4) 0.633 (8) 0.895 (2) 0.284 (14) 0.798 (6) 0.472 (7) 0.163 (13) 0.829 (3) 0.980 (1) 0.087 (14) 1.000 (1) 0.455 (12) 0.440 (4) 0.958 (5) 0.434 (5) 0.865 (2) 0.771 (2) 0.232 (2) 0.269 (6)
0.887 (5) 0.998 (1) 0.518 (11) 0.537 (3) 0.724 (6) 0.695 (10) 1.000 (2) 0.403 (14) 0.687 (14) 0.516 (8) 0.053 (8) 0.264 (4) 0.788 (7) 0.772 (4) 0.577 (3) 0.735 (7) 1.000 (2) 0.989 (3) 0.739 (5) 0.966 (1) 0.541 (12) 0.635 (8) 0.823 (3) 0.910 (3) 0.973 (6) 0.858 (10) 0.501 (7) 0.806 (7) 0.051 (7) 0.626 (9) 0.896 (1) 0.408 (8) 0.837 (5) 0.485 (6) 0.309 (4) 0.655 (10) 0.936 (3) 0.288 (8) 1.000 (2) 0.489 (6) 0.438 (5) 0.813 (16) 0.331 (8) 0.822 (7) 0.638 (10) 0.231 (3) 0.203 (9)
0.900 (3) 0.977 (7) 0.557 (4) 0.530 (5) 0.541 (9) 0.717 (7) 1.000 (3) 0.444 (12) 0.729 (13) 0.437 (9) 0.050 (11) 0.235 (6) 0.786 (8) 0.756 (5) 0.562 (4) 0.776 (5) 1.000 (4) 0.991 (2) 0.683 (9) 0.939 (2) 0.619 (4) 0.451 (11) 0.814 (4) 0.885 (5) 0.970 (9) 0.677 (11) 0.600 (3) 0.864 (4) 0.058 (3) 0.696 (3) 0.893 (3) 0.329 (11) 0.839 (4) 0.517 (3) 0.214 (7) 0.808 (7) 0.979 (2) 0.090 (12) 1.000 (3) 0.489 (7) 0.440 (3) 0.901 (12) 0.427 (6) 0.860 (3) 0.755 (4) 0.238 (1) 0.238 (7)
0.531 (13) 0.946 (9) 0.465 (17) 0.235 (15) 0.752 (4) 0.547 (13) 0.918 (10) 0.312 (16) 0.583 (18) 0.322 (14) 0.049 (12) 0.234 (7) 0.464 (14) 0.731 (10) 0.444 (11) 0.396 (16) 0.698 (14) 0.422 (13) 0.644 (13) 0.516 (14) 0.499 (16) 0.704 (5) 0.663 (13) 0.448 (17) 0.980 (5) 0.970 (7) 0.500 (8) 0.617 (12) 0.047 (10) 0.389 (13) 0.473 (14) 0.397 (9) 0.453 (14) 0.494 (4) 0.207 (9) 0.330 (16) 0.587 (11) 0.803 (1) 0.707 (13) 0.519 (3) 0.408 (10) 0.634 (18) 0.223 (17) 0.375 (17) 0.493 (16) 0.118 (13) 0.146 (14)
0.917 (2) 0.830 (13) 0.537 (7) 0.539 (2) 0.480 (10) 0.489 (15) 0.601 (13) 0.558 (5) 0.866 (6) 0.324 (12) 0.051 (10) 0.216 (12) 0.953 (1) 0.726 (11) 0.366 (13) 0.671 (8) 0.984 (8) 0.856 (7) 0.628 (14) 0.780 (9) 0.561 (10) 0.443 (13) 0.772 (10) 0.933 (1) 0.972 (7) 0.624 (13) 0.304 (13) 0.527 (15) 0.048 (8) 0.638 (7) 0.878 (4) 0.295 (13) 0.843 (3) 0.490 (5) 0.202 (11) 0.811 (5) 0.543 (12) 0.073 (16) 0.966 (6) 0.367 (14) 0.372 (14) 0.942 (7) 0.265 (12) 0.872 (1) 0.770 (3) 0.187 (7) 0.174 (10)
0.358 (15) 0.068 (17) 0.486 (15) 0.380 (12) 0.337 (12) 0.549 (12) 0.184 (17) 0.478 (9) 0.855 (8) 0.281 (15) 0.034 (18) 0.244 (5) 0.292 (16) 0.503 (18) 0.038 (17) 0.565 (13) 0.766 (12) 0.839 (8) 0.311 (18) 0.738 (11) 0.512 (14) 0.107 (18) 0.584 (15) 0.858 (6) 0.088 (18) 0.447 (15) 0.312 (12) 0.509 (16) 0.071 (1) 0.346 (14) 0.398 (16) 0.315 (12) 0.566 (11) 0.328 (11) 0.124 (14) 0.167 (17) 0.326 (17) 0.076 (15) 0.472 (14) 0.471 (11) 0.376 (13) 0.754 (17) 0.238 (16) 0.593 (11) 0.532 (14) 0.132 (11) 0.076 (18)
0.100 (17) 0.096 (16) 0.536 (8) 0.491 (6) 0.156 (16) 0.264 (17) 0.293 (16) 0.321 (15) 0.919 (1) 0.215 (16) 0.049 (13) 0.221 (10) 0.218 (17) 0.652 (14) 0.115 (16) 0.380 (17) 0.447 (17) 0.040 (17) 0.533 (15) 0.102 (17) 0.541 (11) 0.123 (17) 0.503 (18) 0.440 (18) 0.885 (15) 0.286 (16) 0.480 (10) 0.421 (17) 0.027 (16) 0.321 (15) 0.291 (18) 0.092 (18) 0.375 (16) 0.254 (14) 0.081 (16) 0.619 (11) 0.346 (16) 0.037 (18) 0.160 (16) 0.219 (17) 0.295 (18) 0.930 (9) 0.253 (13) 0.481 (15) 0.431 (17) 0.036 (18) 0.101 (15)
0.771 (9) 0.774 (14) 0.515 (12) 0.002 (18) 0.041 (18) 0.744 (4) 0.004 (18) 0.555 (6) 0.827 (10) 0.570 (6) 0.056 (6) 0.191 (15) 0.040 (18) 0.637 (15) 0.354 (14) 0.635 (11) 0.972 (9) 0.660 (11) 0.697 (7) 0.791 (8) 0.583 (8) 0.660 (6) 0.711 (12) 0.710 (12) 0.932 (13) 0.942 (9) 0.000 (18) 0.754 (9) 0.036 (13) 0.607 (11) 0.810 (6) 0.484 (1) 0.560 (12) 0.034 (18) 0.209 (8) 0.045 (18) 0.029 (18) 0.389 (6) 0.902 (8) 0.529 (2) 0.393 (11) 0.894 (13) 0.420 (7) 0.675 (9) 0.631 (11) 0.145 (9) 0.349 (4)
0.686 (11) 0.901 (12) 0.562 (3) 0.393 (11) 0.762 (3) 0.755 (3) 0.958 (7) 0.685 (2) 0.905 (3) 0.683 (3) 0.052 (9) 0.195 (13) 0.946 (2) 0.712 (12) 0.446 (10) 0.798 (3) 0.964 (10) 0.811 (9) 0.763 (4) 0.819 (7) 0.611 (5) 0.754 (4) 0.786 (8) 0.844 (7) 0.981 (4) 0.986 (2) 0.269 (14) 0.891 (3) 0.028 (15) 0.686 (4) 0.795 (7) 0.435 (6) 0.521 (13) 0.366 (9) 0.202 (10) 0.839 (2) 0.821 (6) 0.585 (5) 0.920 (7) 0.490 (5) 0.446 (1) 0.970 (1) 0.460 (4) 0.857 (4) 0.723 (6) 0.131 (12) 0.336 (5)
0.851 (7) 0.989 (3) 0.523 (9) 0.372 (13) 0.827 (1) 0.701 (9) 1.000 (4) 0.273 (18) 0.608 (17) 0.564 (7) 0.037 (17) 0.194 (14) 0.753 (10) 0.773 (3) 0.539 (6) 0.641 (9) 0.991 (5) 0.863 (6) 0.682 (10) 0.937 (3) 0.581 (9) 0.498 (10) 0.811 (6) 0.834 (8) 0.962 (12) 0.982 (4) 0.500 (9) 0.763 (8) 0.053 (5) 0.700 (1) 0.603 (11) 0.465 (3) 0.648 (7) 0.331 (10) 0.346 (3) 0.767 (9) 0.865 (4) 0.282 (9) 0.972 (5) 0.513 (4) 0.437 (7) 0.960 (4) 0.297 (10) 0.566 (12) 0.659 (7) 0.200 (6) 0.173 (11)
0.461 (14) 0.943 (10) 0.477 (16) 0.194 (17) 0.553 (8) 0.625 (11) 0.968 (6) 0.498 (7) 0.616 (16) 0.323 (13) 0.043 (15) 0.172 (16) 0.831 (5) 0.752 (6) 0.465 (9) 0.428 (15) 0.618 (15) 0.366 (14) 0.652 (12) 0.476 (15) 0.539 (13) 0.652 (7) 0.535 (17) 0.475 (16) 0.909 (14) 0.985 (3) 0.074 (17) 0.691 (11) 0.047 (9) 0.573 (12) 0.469 (15) 0.459 (4) 0.408 (15) 0.263 (13) 0.196 (12) 0.811 (4) 0.704 (9) 0.745 (2) 0.882 (10) 0.476 (10) 0.441 (2) 0.927 (10) 0.211 (18) 0.528 (14) 0.577 (13) 0.100 (14) 0.154 (12)
0.874 (6) 0.982 (5) 0.580 (1) 0.400 (10) 0.769 (2) 0.772 (1) 0.881 (11) 0.652 (3) 0.874 (5) 0.717 (2) 0.087 (2) 0.519 (2) 0.866 (4) 0.827 (1) 0.582 (2) 0.824 (2) 0.990 (6) 0.910 (5) 0.812 (1) 0.891 (6) 0.621 (2) 0.786 (1) 0.851 (1) 0.824 (10) 0.992 (1) 0.981 (5) 0.600 (2) 0.935 (2) 0.052 (6) 0.697 (2) 0.767 (9) 0.479 (2) 0.849 (2) 0.670 (1) 0.361 (2) 0.811 (6) 0.837 (5) 0.739 (3) 0.984 (4) 0.531 (1) 0.438 (6) 0.969 (2) 0.569 (1) 0.834 (6) 0.784 (1) 0.216 (5) 0.363 (2)
0.932 (1) 0.970 (8) 0.569 (2) 0.478 (7) 0.728 (5) 0.755 (2) 0.975 (5) 0.767 (1) 0.899 (4) 0.721 (1) 0.114 (1) 0.578 (1) 0.875 (3) 0.824 (2) 0.553 (5) 0.927 (1) 1.000 (3) 0.931 (4) 0.808 (2) 0.905 (5) 0.633 (1) 0.763 (2) 0.844 (2) 0.924 (2) 0.982 (3) 0.993 (1) 0.516 (5) 0.944 (1) 0.044 (11) 0.680 (5) 0.786 (8) 0.449 (5) 0.889 (1) 0.606 (2) 0.406 (1) 0.852 (1) 0.795 (7) 0.737 (4) 0.897 (9) 0.485 (8) 0.434 (8) 0.967 (3) 0.534 (2) 0.835 (5) 0.751 (5) 0.230 (4) 0.365 (1)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.576 (9) 0.608 (9) 0.722 (8) 0.704 (4) 0.085 (11)
0.380 (14) 0.385 (15) 0.520 (13) 0.494 (16) 0.072 (14)
0.408 (11) 0.453 (11) 0.646 (12) 0.640 (10) 0.083 (12)
0.396 (13) 0.431 (13) 0.215 (15) 0.472 (17) 0.059 (17)
0.114 (18) 0.284 (18) 0.138 (18) 0.686 (6) 0.056 (18)
0.666 (4) 0.723 (4) 0.831 (2) 0.668 (7) 0.112 (2)
0.618 (7) 0.700 (6) 0.831 (1) 0.608 (12) 0.106 (4)
0.655 (5) 0.734 (2) 0.825 (3) 0.657 (8) 0.109 (3)
0.350 (16) 0.396 (14) 0.704 (9) 0.540 (14) 0.105 (5)
0.670 (2) 0.733 (3) 0.811 (6) 0.716 (3) 0.101 (6)
0.282 (17) 0.382 (16) 0.206 (16) 0.464 (18) 0.061 (16)
0.397 (12) 0.440 (12) 0.183 (17) 0.525 (15) 0.066 (15)
0.511 (10) 0.541 (10) 0.667 (11) 0.702 (5) 0.086 (10)
0.580 (8) 0.644 (8) 0.687 (10) 0.758 (1) 0.074 (13)
0.650 (6) 0.721 (5) 0.793 (7) 0.633 (11) 0.100 (7)
0.379 (15) 0.381 (17) 0.477 (14) 0.599 (13) 0.088 (9)
0.668 (3) 0.700 (7) 0.813 (5) 0.745 (2) 0.114 (1)
0.716 (1) 0.768 (1) 0.819 (4) 0.646 (9) 0.098 (8)
20news agnews amazon imdb yelp
0.219 (8) 0.427 (8) 0.174 (9) 0.197 (10) 0.300 (11)
0.128 (15) 0.273 (14) 0.089 (14) 0.103 (14) 0.169 (14)
0.191 (11) 0.373 (11) 0.131 (12) 0.153 (13) 0.302 (10)
0.073 (18) 0.085 (18) 0.051 (17) 0.050 (17) 0.056 (16)
0.113 (16) 0.161 (16) 0.052 (16) 0.047 (18) 0.051 (17)
0.320 (2) 0.617 (2) 0.269 (2) 0.365 (1) 0.565 (1)
0.263 (7) 0.572 (6) 0.246 (4) 0.319 (5) 0.489 (5)
0.316 (3) 0.610 (3) 0.269 (3) 0.359 (2) 0.542 (2)
0.209 (10) 0.385 (10) 0.209 (8) 0.232 (8) 0.367 (8)
0.293 (6) 0.601 (4) 0.246 (5) 0.327 (4) 0.518 (3)
0.108 (17) 0.128 (17) 0.049 (18) 0.053 (16) 0.046 (18)
0.147 (14) 0.298 (13) 0.054 (15) 0.054 (15) 0.061 (15)
0.211 (9) 0.421 (9) 0.174 (10) 0.216 (9) 0.303 (9)
0.189 (12) 0.325 (12) 0.107 (13) 0.163 (11) 0.208 (13)
0.296 (5) 0.629 (1) 0.226 (7) 0.311 (7) 0.454 (7)
0.169 (13) 0.230 (15) 0.133 (11) 0.159 (12) 0.237 (12)
0.310 (4) 0.565 (7) 0.235 (6) 0.314 (6) 0.455 (6)
0.323 (1) 0.599 (5) 0.282 (1) 0.339 (3) 0.499 (4)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 37: The average AUCPR performance on individual datasets under the 𝛾𝑙𝑎 = 10% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.832 (9) 0.975 (11) 0.583 (4) 0.391 (15) 0.765 (4) 0.875 (5) 0.949 (9) 0.573 (4) 0.904 (7) 0.739 (4) 0.093 (3) 0.415 (3) 0.788 (11) 0.787 (4) 0.488 (10) 0.854 (4) 0.999 (8) 0.869 (11) 0.765 (6) 0.814 (12) 0.617 (7) 0.811 (2) 0.853 (3) 0.785 (11) 0.990 (3) 0.986 (4) 0.504 (6) 0.919 (4) 0.034 (13) 0.754 (6) 0.864 (7) 0.584 (5) 0.701 (7) 0.586 (8) 0.281 (7) 0.650 (13) 0.819 (8) 0.520 (7) 0.897 (12) 0.649 (4) 0.455 (8) 0.940 (12) 0.501 (3) 0.832 (8) 0.747 (9) 0.205 (7) 0.384 (3)
0.794 (12) 0.998 (3) 0.557 (10) 0.321 (16) 0.314 (14) 0.676 (15) 0.859 (13) 0.504 (9) 0.821 (11) 0.725 (5) 0.065 (4) 0.223 (9) 0.544 (16) 0.755 (9) 0.570 (7) 0.679 (11) 0.876 (12) 0.886 (10) 0.764 (7) 0.893 (9) 0.527 (16) 0.506 (11) 0.831 (6) 0.872 (7) 0.974 (7) 0.971 (7) 0.582 (4) 0.698 (14) 0.024 (16) 0.350 (15) 0.676 (14) 0.258 (15) 0.649 (9) 0.213 (17) 0.293 (6) 0.668 (12) 0.461 (14) 0.357 (10) 0.394 (15) 0.399 (15) 0.406 (12) 0.888 (15) 0.273 (13) 0.588 (14) 0.702 (11) 0.114 (14) 0.179 (10)
0.866 (8) 0.997 (5) 0.565 (8) 0.437 (11) 0.478 (10) 0.862 (8) 0.933 (11) 0.466 (11) 0.932 (3) 0.448 (11) 0.058 (6) 0.260 (4) 0.705 (12) 0.768 (8) 0.563 (8) 0.701 (10) 0.831 (14) 0.514 (15) 0.788 (4) 0.910 (8) 0.598 (10) 0.663 (9) 0.785 (11) 0.708 (13) 0.970 (11) 0.539 (14) 0.440 (11) 0.762 (11) 0.110 (1) 0.769 (5) 0.845 (9) 0.481 (8) 0.632 (11) 0.474 (11) 0.115 (15) 0.641 (14) 0.592 (13) 0.117 (11) 0.952 (8) 0.564 (11) 0.388 (14) 0.948 (10) 0.293 (11) 0.708 (10) 0.728 (10) 0.162 (10) 0.238 (9)
0.190 (16) 0.059 (18) 0.448 (18) 0.434 (12) 0.130 (18) 0.376 (16) 0.244 (18) 0.485 (10) 0.703 (17) 0.160 (18) 0.048 (14) 0.069 (18) 0.599 (15) 0.560 (16) 0.037 (17) 0.604 (12) 0.233 (18) 0.065 (16) 0.432 (16) 0.066 (18) 0.383 (18) 0.138 (16) 0.569 (17) 0.599 (15) 0.669 (16) 0.165 (18) 0.216 (15) 0.741 (13) 0.015 (18) 0.212 (18) 0.277 (17) 0.120 (16) 0.252 (18) 0.058 (18) 0.048 (17) 0.671 (11) 0.425 (15) 0.057 (17) 0.063 (18) 0.215 (18) 0.299 (17) 0.958 (7) 0.340 (10) 0.533 (15) 0.472 (16) 0.089 (16) 0.059 (18)
0.008 (18) 0.105 (16) 0.500 (16) 0.556 (7) 0.172 (16) 0.249 (18) 0.453 (15) 0.445 (12) 0.865 (8) 0.216 (16) 0.041 (18) 0.166 (17) 0.628 (13) 0.472 (18) 0.110 (16) 0.206 (18) 0.715 (16) 0.029 (18) 0.337 (17) 0.177 (16) 0.458 (17) 0.346 (15) 0.631 (15) 0.481 (17) 0.313 (17) 0.197 (17) 0.106 (16) 0.406 (18) 0.020 (17) 0.290 (17) 0.506 (16) 0.098 (17) 0.352 (17) 0.317 (15) 0.042 (18) 0.417 (16) 0.636 (12) 0.059 (16) 0.140 (17) 0.231 (16) 0.311 (16) 0.861 (16) 0.285 (12) 0.367 (18) 0.360 (18) 0.072 (17) 0.075 (17)
0.929 (4) 0.999 (2) 0.560 (9) 0.639 (2) 0.207 (15) 0.840 (9) 1.000 (1) 0.527 (7) 0.823 (10) 0.419 (12) 0.052 (12) 0.174 (15) 0.932 (5) 0.729 (12) 0.630 (2) 0.835 (5) 1.000 (2) 0.992 (2) 0.667 (13) 0.935 (5) 0.613 (8) 0.469 (12) 0.821 (7) 0.899 (4) 0.971 (10) 0.656 (12) 0.600 (1) 0.866 (7) 0.106 (3) 0.688 (12) 0.906 (3) 0.345 (14) 0.885 (2) 0.720 (6) 0.188 (13) 0.867 (3) 1.000 (1) 0.088 (13) 1.000 (1) 0.530 (13) 0.449 (9) 0.985 (1) 0.460 (6) 0.907 (1) 0.819 (3) 0.265 (5) 0.297 (6)
0.921 (5) 0.999 (1) 0.547 (11) 0.626 (3) 0.741 (6) 0.864 (6) 1.000 (2) 0.404 (14) 0.759 (14) 0.568 (8) 0.044 (15) 0.228 (8) 0.832 (9) 0.777 (6) 0.620 (3) 0.796 (7) 1.000 (4) 0.992 (1) 0.767 (5) 0.969 (1) 0.605 (9) 0.694 (8) 0.835 (5) 0.910 (3) 0.973 (8) 0.861 (10) 0.501 (7) 0.882 (6) 0.082 (7) 0.734 (7) 0.909 (2) 0.463 (10) 0.736 (6) 0.764 (4) 0.337 (4) 0.735 (9) 0.988 (3) 0.360 (9) 1.000 (2) 0.640 (5) 0.497 (2) 0.948 (9) 0.388 (9) 0.880 (6) 0.766 (8) 0.271 (3) 0.243 (8)
0.938 (3) 0.984 (10) 0.573 (5) 0.644 (1) 0.402 (12) 0.825 (11) 1.000 (3) 0.520 (8) 0.801 (13) 0.477 (10) 0.056 (8) 0.188 (12) 0.937 (4) 0.755 (10) 0.591 (6) 0.831 (6) 1.000 (6) 0.991 (3) 0.677 (12) 0.929 (6) 0.628 (5) 0.468 (13) 0.820 (8) 0.880 (6) 0.971 (9) 0.678 (11) 0.600 (3) 0.883 (5) 0.106 (2) 0.733 (8) 0.895 (4) 0.406 (11) 0.879 (3) 0.790 (3) 0.219 (9) 0.851 (4) 1.000 (2) 0.090 (12) 1.000 (3) 0.561 (12) 0.461 (6) 0.972 (4) 0.475 (5) 0.900 (2) 0.827 (2) 0.267 (4) 0.295 (7)
0.650 (13) 0.993 (8) 0.506 (15) 0.406 (14) 0.731 (7) 0.775 (12) 0.918 (12) 0.318 (17) 0.650 (18) 0.320 (14) 0.055 (10) 0.231 (7) 0.615 (14) 0.782 (5) 0.437 (12) 0.480 (16) 0.907 (11) 0.546 (14) 0.760 (8) 0.658 (15) 0.549 (13) 0.786 (5) 0.783 (12) 0.576 (16) 0.980 (5) 0.969 (8) 0.500 (8) 0.757 (12) 0.089 (6) 0.706 (11) 0.715 (12) 0.477 (9) 0.495 (14) 0.484 (10) 0.195 (12) 0.375 (17) 0.692 (11) 0.899 (1) 0.717 (13) 0.597 (10) 0.417 (11) 0.731 (17) 0.247 (17) 0.446 (17) 0.587 (14) 0.129 (13) 0.165 (11)
0.941 (2) 0.849 (14) 0.571 (6) 0.572 (6) 0.439 (11) 0.704 (14) 0.626 (14) 0.550 (6) 0.864 (9) 0.344 (13) 0.053 (11) 0.178 (13) 0.987 (1) 0.710 (13) 0.385 (13) 0.578 (14) 1.000 (3) 0.901 (9) 0.638 (14) 0.817 (11) 0.548 (14) 0.408 (14) 0.792 (10) 0.927 (2) 0.975 (6) 0.632 (13) 0.304 (13) 0.526 (15) 0.064 (9) 0.656 (13) 0.887 (5) 0.404 (12) 0.839 (5) 0.750 (5) 0.201 (11) 0.831 (6) 0.812 (9) 0.075 (14) 0.964 (7) 0.413 (14) 0.401 (13) 0.955 (8) 0.254 (16) 0.891 (5) 0.791 (5) 0.176 (8) 0.159 (13)
0.389 (15) 0.136 (15) 0.496 (17) 0.428 (13) 0.392 (13) 0.732 (13) 0.303 (16) 0.240 (18) 0.811 (12) 0.260 (15) 0.043 (16) 0.191 (11) 0.477 (17) 0.548 (17) 0.027 (18) 0.589 (13) 0.841 (13) 0.946 (6) 0.252 (18) 0.786 (13) 0.543 (15) 0.125 (18) 0.630 (16) 0.885 (5) 0.269 (18) 0.353 (15) 0.312 (12) 0.478 (16) 0.102 (4) 0.541 (14) 0.507 (15) 0.401 (13) 0.408 (15) 0.459 (12) 0.150 (14) 0.177 (18) 0.398 (17) 0.070 (15) 0.679 (14) 0.614 (9) 0.369 (15) 0.516 (18) 0.232 (18) 0.643 (12) 0.548 (15) 0.136 (12) 0.076 (16)
0.095 (17) 0.100 (17) 0.547 (12) 0.484 (10) 0.159 (17) 0.296 (17) 0.297 (17) 0.329 (15) 0.919 (5) 0.207 (17) 0.049 (13) 0.236 (6) 0.297 (18) 0.656 (15) 0.114 (15) 0.390 (17) 0.425 (17) 0.041 (17) 0.537 (15) 0.115 (17) 0.554 (12) 0.125 (17) 0.509 (18) 0.435 (18) 0.925 (15) 0.297 (16) 0.480 (10) 0.419 (17) 0.026 (15) 0.322 (16) 0.274 (18) 0.096 (18) 0.375 (16) 0.274 (16) 0.080 (16) 0.623 (15) 0.407 (16) 0.037 (18) 0.220 (16) 0.225 (17) 0.295 (18) 0.931 (13) 0.256 (15) 0.492 (16) 0.434 (17) 0.035 (18) 0.103 (15)
0.803 (11) 0.886 (13) 0.543 (13) 0.002 (18) 0.637 (9) 0.892 (4) 0.983 (7) 0.568 (5) 0.908 (6) 0.647 (6) 0.057 (7) 0.218 (10) 0.813 (10) 0.676 (14) 0.377 (14) 0.749 (9) 0.988 (9) 0.799 (12) 0.749 (9) 0.840 (10) 0.628 (4) 0.768 (6) 0.777 (13) 0.767 (12) 0.954 (13) 0.971 (6) 0.000 (18) 0.825 (9) 0.035 (12) 0.716 (9) 0.836 (10) 0.591 (4) 0.685 (8) 0.456 (13) 0.274 (8) 0.721 (10) 0.221 (18) 0.563 (6) 0.936 (10) 0.652 (3) 0.431 (10) 0.920 (14) 0.451 (8) 0.734 (9) 0.695 (12) 0.165 (9) 0.344 (5)
0.811 (10) 0.954 (12) 0.595 (3) 0.604 (4) 0.862 (1) 0.901 (3) 0.956 (8) 0.741 (3) 0.926 (4) 0.753 (3) 0.059 (5) 0.246 (5) 0.954 (3) 0.768 (7) 0.464 (11) 0.868 (3) 1.000 (1) 0.909 (8) 0.821 (3) 0.926 (7) 0.643 (3) 0.790 (4) 0.849 (4) 0.854 (9) 0.987 (4) 0.991 (3) 0.269 (14) 0.923 (3) 0.029 (14) 0.804 (4) 0.853 (8) 0.606 (3) 0.636 (10) 0.561 (9) 0.305 (5) 0.871 (1) 0.984 (4) 0.732 (5) 0.973 (6) 0.679 (2) 0.500 (1) 0.962 (6) 0.493 (4) 0.897 (3) 0.789 (6) 0.149 (11) 0.377 (4)
0.879 (7) 0.996 (6) 0.566 (7) 0.595 (5) 0.815 (3) 0.862 (7) 1.000 (4) 0.320 (16) 0.750 (15) 0.575 (7) 0.043 (17) 0.177 (14) 0.877 (8) 0.746 (11) 0.600 (5) 0.768 (8) 0.971 (10) 0.918 (7) 0.733 (11) 0.963 (2) 0.625 (6) 0.634 (10) 0.814 (9) 0.831 (10) 0.967 (12) 0.964 (9) 0.500 (9) 0.851 (8) 0.093 (5) 0.820 (1) 0.733 (11) 0.584 (6) 0.627 (12) 0.635 (7) 0.380 (3) 0.823 (7) 0.912 (7) 0.369 (8) 1.000 (4) 0.615 (8) 0.459 (7) 0.971 (5) 0.456 (7) 0.705 (11) 0.767 (7) 0.257 (6) 0.162 (12)
0.614 (14) 0.992 (9) 0.515 (14) 0.303 (17) 0.665 (8) 0.826 (10) 0.987 (6) 0.433 (13) 0.732 (16) 0.515 (9) 0.055 (9) 0.171 (16) 0.958 (2) 0.813 (3) 0.521 (9) 0.492 (15) 0.720 (15) 0.648 (13) 0.745 (10) 0.666 (14) 0.563 (11) 0.758 (7) 0.718 (14) 0.602 (14) 0.948 (14) 0.979 (5) 0.074 (17) 0.778 (10) 0.050 (11) 0.712 (10) 0.683 (13) 0.506 (7) 0.563 (13) 0.342 (14) 0.213 (10) 0.804 (8) 0.799 (10) 0.807 (3) 0.907 (11) 0.620 (7) 0.490 (3) 0.945 (11) 0.268 (14) 0.610 (13) 0.664 (13) 0.111 (15) 0.131 (14)
0.908 (6) 0.997 (4) 0.624 (2) 0.524 (9) 0.841 (2) 0.909 (1) 0.948 (10) 0.768 (2) 0.934 (2) 0.790 (2) 0.114 (2) 0.579 (2) 0.924 (7) 0.854 (2) 0.640 (1) 0.905 (2) 1.000 (7) 0.975 (4) 0.855 (2) 0.947 (3) 0.655 (2) 0.824 (1) 0.894 (2) 0.866 (8) 0.998 (1) 0.993 (2) 0.600 (2) 0.958 (1) 0.075 (8) 0.815 (2) 0.926 (1) 0.619 (2) 0.862 (4) 0.845 (1) 0.452 (2) 0.850 (5) 0.946 (5) 0.825 (2) 0.993 (5) 0.681 (1) 0.490 (4) 0.982 (2) 0.596 (1) 0.894 (4) 0.833 (1) 0.272 (2) 0.401 (2)
0.967 (1) 0.994 (7) 0.639 (1) 0.537 (8) 0.764 (5) 0.907 (2) 0.998 (5) 0.806 (1) 0.942 (1) 0.795 (1) 0.158 (1) 0.618 (1) 0.931 (6) 0.863 (1) 0.618 (4) 0.956 (1) 1.000 (5) 0.964 (5) 0.867 (1) 0.944 (4) 0.683 (1) 0.800 (3) 0.899 (1) 0.939 (1) 0.995 (2) 0.997 (1) 0.516 (5) 0.956 (2) 0.058 (10) 0.812 (3) 0.871 (6) 0.627 (1) 0.930 (1) 0.837 (2) 0.478 (1) 0.871 (2) 0.927 (6) 0.803 (4) 0.941 (9) 0.640 (6) 0.488 (5) 0.974 (3) 0.571 (2) 0.872 (7) 0.803 (4) 0.273 (1) 0.403 (1)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.653 (9) 0.697 (9) 0.775 (10) 0.782 (5) 0.099 (12)
0.455 (13) 0.494 (13) 0.613 (14) 0.585 (15) 0.082 (13)
0.553 (11) 0.607 (11) 0.776 (9) 0.735 (9) 0.104 (8)
0.399 (16) 0.420 (16) 0.212 (16) 0.506 (17) 0.059 (17)
0.116 (18) 0.287 (18) 0.135 (18) 0.686 (12) 0.055 (18)
0.739 (3) 0.790 (3) 0.855 (2) 0.800 (4) 0.133 (5)
0.702 (7) 0.758 (7) 0.857 (1) 0.730 (10) 0.134 (2)
0.733 (5) 0.793 (2) 0.854 (3) 0.779 (6) 0.134 (3)
0.486 (12) 0.515 (12) 0.792 (8) 0.636 (14) 0.118 (7)
0.731 (6) 0.786 (4) 0.838 (6) 0.816 (2) 0.121 (6)
0.289 (17) 0.367 (17) 0.252 (15) 0.440 (18) 0.060 (16)
0.402 (15) 0.445 (14) 0.187 (17) 0.543 (16) 0.066 (15)
0.580 (10) 0.630 (10) 0.728 (12) 0.773 (7) 0.101 (10)
0.672 (8) 0.709 (8) 0.746 (11) 0.811 (3) 0.081 (14)
0.737 (4) 0.786 (5) 0.806 (7) 0.742 (8) 0.103 (9)
0.428 (14) 0.439 (15) 0.625 (13) 0.673 (13) 0.101 (11)
0.752 (2) 0.779 (6) 0.852 (5) 0.829 (1) 0.146 (1)
0.765 (1) 0.809 (1) 0.852 (4) 0.714 (11) 0.134 (4)
20news agnews amazon imdb yelp
0.275 (9) 0.545 (9) 0.238 (8) 0.257 (9) 0.435 (11)
0.167 (14) 0.360 (13) 0.122 (14) 0.128 (14) 0.231 (14)
0.227 (12) 0.549 (8) 0.205 (11) 0.243 (11) 0.460 (8)
0.077 (18) 0.100 (18) 0.054 (16) 0.048 (18) 0.051 (16)
0.113 (16) 0.140 (17) 0.048 (18) 0.051 (17) 0.050 (17)
0.405 (5) 0.699 (2) 0.338 (2) 0.429 (1) 0.659 (2)
0.337 (7) 0.655 (7) 0.336 (3) 0.376 (5) 0.642 (3)
0.410 (2) 0.699 (1) 0.327 (5) 0.426 (2) 0.660 (1)
0.277 (8) 0.489 (11) 0.242 (7) 0.312 (7) 0.458 (9)
0.408 (3) 0.687 (4) 0.287 (6) 0.370 (6) 0.622 (5)
0.090 (17) 0.163 (16) 0.052 (17) 0.065 (15) 0.045 (18)
0.151 (15) 0.304 (15) 0.055 (15) 0.056 (16) 0.060 (15)
0.272 (10) 0.534 (10) 0.224 (10) 0.255 (10) 0.441 (10)
0.248 (11) 0.467 (12) 0.161 (13) 0.226 (12) 0.381 (12)
0.395 (6) 0.684 (6) 0.227 (9) 0.294 (8) 0.581 (7)
0.208 (13) 0.320 (14) 0.198 (12) 0.197 (13) 0.313 (13)
0.415 (1) 0.686 (5) 0.334 (4) 0.379 (4) 0.601 (6)
0.406 (4) 0.689 (3) 0.342 (1) 0.417 (3) 0.630 (4)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 38: The average AUCPR performance on individual datasets under the 𝛾𝑙𝑎 = 25% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.890 (13) 0.988 (10) 0.617 (4) 0.633 (10) 0.864 (6) 0.983 (5) 0.967 (9) 0.758 (4) 0.976 (3) 0.795 (4) 0.108 (3) 0.511 (3) 0.823 (13) 0.833 (5) 0.570 (12) 0.917 (4) 1.000 (8) 0.934 (9) 0.807 (10) 0.954 (8) 0.721 (4) 0.816 (6) 0.899 (3) 0.884 (10) 0.993 (3) 0.992 (4) 0.504 (6) 0.949 (4) 0.048 (11) 0.845 (7) 0.905 (6) 0.801 (4) 0.759 (9) 0.652 (11) 0.386 (5) 0.881 (8) 0.946 (10) 0.782 (6) 0.976 (11) 0.832 (5) 0.541 (4) 0.973 (9) 0.563 (3) 0.907 (7) 0.833 (8) 0.268 (7) 0.483 (3)
0.935 (7) 1.000 (1) 0.584 (12) 0.495 (14) 0.412 (12) 0.909 (12) 0.883 (13) 0.633 (10) 0.913 (10) 0.774 (5) 0.073 (6) 0.273 (13) 0.785 (14) 0.817 (6) 0.621 (5) 0.858 (9) 1.000 (9) 0.976 (6) 0.815 (8) 0.990 (2) 0.578 (13) 0.652 (11) 0.887 (5) 0.937 (4) 0.979 (7) 0.986 (8) 0.582 (4) 0.831 (11) 0.033 (13) 0.536 (15) 0.836 (13) 0.396 (14) 0.720 (11) 0.357 (15) 0.353 (7) 0.702 (15) 0.654 (14) 0.652 (8) 0.666 (14) 0.586 (12) 0.472 (12) 0.935 (15) 0.327 (13) 0.791 (13) 0.813 (10) 0.206 (12) 0.279 (10)
0.950 (5) 0.993 (8) 0.606 (6) 0.674 (7) 0.623 (10) 0.965 (10) 0.933 (11) 0.641 (9) 0.972 (4) 0.464 (11) 0.059 (9) 0.486 (5) 0.901 (10) 0.787 (10) 0.603 (9) 0.849 (10) 0.970 (12) 0.798 (12) 0.809 (9) 0.988 (4) 0.667 (10) 0.683 (9) 0.794 (13) 0.887 (9) 0.974 (10) 0.544 (14) 0.440 (11) 0.805 (13) 0.155 (1) 0.851 (5) 0.894 (10) 0.552 (10) 0.699 (12) 0.752 (8) 0.150 (15) 0.829 (11) 0.791 (12) 0.116 (11) 0.989 (7) 0.684 (10) 0.393 (13) 0.964 (12) 0.340 (11) 0.849 (10) 0.824 (9) 0.209 (11) 0.326 (9)
0.159 (15) 0.201 (15) 0.437 (18) 0.434 (17) 0.102 (18) 0.273 (17) 0.280 (16) 0.580 (11) 0.763 (18) 0.161 (18) 0.046 (17) 0.078 (18) 0.338 (18) 0.575 (17) 0.045 (17) 0.685 (13) 0.076 (18) 0.057 (15) 0.361 (17) 0.161 (17) 0.426 (17) 0.180 (16) 0.584 (17) 0.495 (15) 0.764 (16) 0.166 (18) 0.216 (15) 0.628 (14) 0.027 (15) 0.254 (18) 0.242 (18) 0.135 (16) 0.177 (18) 0.053 (18) 0.054 (17) 0.710 (14) 0.369 (17) 0.060 (16) 0.065 (18) 0.227 (18) 0.330 (16) 0.971 (10) 0.272 (16) 0.575 (15) 0.469 (17) 0.097 (15) 0.045 (18)
0.009 (18) 0.137 (16) 0.508 (16) 0.559 (12) 0.188 (16) 0.264 (18) 0.245 (17) 0.455 (15) 0.866 (16) 0.233 (16) 0.041 (18) 0.168 (17) 0.660 (15) 0.446 (18) 0.133 (15) 0.211 (18) 0.697 (16) 0.029 (18) 0.389 (16) 0.184 (16) 0.466 (16) 0.304 (15) 0.622 (16) 0.491 (16) 0.379 (18) 0.200 (17) 0.106 (16) 0.414 (17) 0.020 (18) 0.296 (17) 0.510 (15) 0.100 (17) 0.375 (17) 0.321 (16) 0.044 (18) 0.418 (17) 0.634 (15) 0.098 (12) 0.157 (17) 0.235 (16) 0.313 (17) 0.869 (17) 0.281 (15) 0.382 (18) 0.374 (18) 0.056 (16) 0.083 (17)
0.961 (2) 0.999 (4) 0.558 (15) 0.712 (1) 0.286 (14) 0.882 (13) 1.000 (1) 0.717 (6) 0.891 (13) 0.452 (12) 0.055 (10) 0.268 (14) 0.957 (7) 0.734 (14) 0.623 (4) 0.876 (7) 1.000 (2) 0.994 (4) 0.683 (13) 0.936 (11) 0.679 (9) 0.458 (13) 0.827 (9) 0.923 (7) 0.971 (12) 0.656 (12) 0.600 (1) 0.873 (9) 0.119 (2) 0.711 (12) 0.922 (3) 0.401 (13) 0.885 (4) 0.834 (5) 0.199 (14) 0.908 (5) 1.000 (2) 0.088 (14) 1.000 (1) 0.584 (13) 0.476 (10) 0.988 (2) 0.504 (7) 0.911 (6) 0.857 (4) 0.299 (5) 0.402 (7)
0.955 (3) 0.999 (3) 0.600 (8) 0.711 (2) 0.858 (8) 0.974 (7) 1.000 (2) 0.540 (12) 0.928 (9) 0.593 (9) 0.046 (16) 0.367 (9) 0.954 (8) 0.788 (9) 0.601 (10) 0.849 (11) 1.000 (4) 0.998 (2) 0.816 (7) 0.972 (7) 0.684 (7) 0.716 (8) 0.833 (8) 0.929 (5) 0.976 (8) 0.884 (10) 0.501 (7) 0.892 (7) 0.096 (7) 0.818 (8) 0.911 (5) 0.639 (9) 0.884 (5) 0.849 (4) 0.301 (9) 0.875 (9) 1.000 (3) 0.386 (10) 1.000 (2) 0.730 (8) 0.518 (6) 0.986 (4) 0.438 (9) 0.921 (3) 0.846 (7) 0.297 (6) 0.364 (8)
0.954 (4) 0.983 (12) 0.597 (9) 0.658 (9) 0.410 (13) 0.919 (11) 1.000 (4) 0.687 (7) 0.890 (14) 0.498 (10) 0.053 (12) 0.291 (12) 0.962 (5) 0.752 (13) 0.604 (8) 0.897 (6) 1.000 (6) 0.994 (3) 0.691 (12) 0.934 (12) 0.682 (8) 0.463 (12) 0.821 (10) 0.919 (8) 0.970 (13) 0.682 (11) 0.600 (3) 0.889 (8) 0.118 (3) 0.765 (11) 0.913 (4) 0.450 (11) 0.886 (3) 0.873 (3) 0.269 (10) 0.895 (6) 1.000 (4) 0.089 (13) 1.000 (3) 0.596 (11) 0.481 (9) 0.985 (5) 0.513 (6) 0.919 (5) 0.851 (5) 0.300 (4) 0.402 (6)
0.898 (12) 0.998 (7) 0.590 (11) 0.534 (13) 0.903 (5) 0.965 (9) 0.899 (12) 0.443 (16) 0.857 (17) 0.338 (14) 0.071 (7) 0.407 (6) 0.834 (12) 0.807 (8) 0.606 (7) 0.681 (14) 0.910 (14) 0.747 (13) 0.844 (4) 0.952 (9) 0.648 (11) 0.832 (3) 0.819 (11) 0.815 (12) 0.983 (6) 0.974 (9) 0.500 (8) 0.811 (12) 0.118 (4) 0.813 (9) 0.899 (8) 0.750 (6) 0.680 (13) 0.719 (9) 0.227 (13) 0.769 (13) 0.683 (13) 0.907 (1) 0.890 (13) 0.775 (6) 0.476 (11) 0.885 (16) 0.336 (12) 0.619 (14) 0.716 (14) 0.184 (13) 0.251 (11)
0.949 (6) 0.858 (14) 0.559 (14) 0.679 (6) 0.482 (11) 0.792 (14) 0.636 (14) 0.645 (8) 0.881 (15) 0.369 (13) 0.054 (11) 0.248 (15) 0.984 (1) 0.753 (12) 0.410 (14) 0.593 (15) 1.000 (3) 0.893 (11) 0.643 (14) 0.857 (14) 0.570 (14) 0.419 (14) 0.805 (12) 0.940 (3) 0.971 (11) 0.652 (13) 0.304 (13) 0.570 (15) 0.067 (9) 0.685 (13) 0.887 (11) 0.421 (12) 0.865 (7) 0.828 (6) 0.232 (12) 0.780 (12) 0.880 (11) 0.080 (15) 0.992 (5) 0.481 (14) 0.387 (14) 0.974 (8) 0.261 (17) 0.900 (8) 0.792 (12) 0.215 (10) 0.166 (14)
0.108 (16) 0.098 (18) 0.494 (17) 0.391 (18) 0.263 (15) 0.364 (16) 0.182 (18) 0.304 (18) 0.904 (11) 0.322 (15) 0.048 (14) 0.312 (11) 0.370 (17) 0.647 (16) 0.042 (18) 0.483 (16) 0.822 (15) 0.043 (17) 0.327 (18) 0.281 (15) 0.417 (18) 0.139 (17) 0.635 (15) 0.447 (17) 0.636 (17) 0.353 (15) 0.321 (12) 0.393 (18) 0.021 (17) 0.585 (14) 0.457 (16) 0.218 (15) 0.448 (15) 0.415 (14) 0.325 (8) 0.306 (18) 0.191 (18) 0.041 (17) 0.332 (15) 0.400 (15) 0.334 (15) 0.774 (18) 0.322 (14) 0.476 (17) 0.473 (15) 0.051 (17) 0.094 (16)
0.098 (17) 0.105 (17) 0.561 (13) 0.494 (15) 0.180 (17) 0.461 (15) 0.293 (15) 0.363 (17) 0.936 (8) 0.211 (17) 0.049 (13) 0.233 (16) 0.391 (16) 0.664 (15) 0.119 (16) 0.416 (17) 0.545 (17) 0.050 (16) 0.538 (15) 0.128 (18) 0.556 (15) 0.131 (18) 0.527 (18) 0.445 (18) 0.920 (15) 0.289 (16) 0.480 (10) 0.429 (16) 0.027 (16) 0.348 (16) 0.275 (17) 0.095 (18) 0.400 (16) 0.269 (17) 0.077 (16) 0.690 (16) 0.402 (16) 0.038 (18) 0.192 (16) 0.230 (17) 0.295 (18) 0.945 (14) 0.259 (18) 0.528 (16) 0.469 (16) 0.035 (18) 0.105 (15)
0.901 (10) 0.963 (13) 0.609 (5) 0.602 (11) 0.859 (7) 0.988 (2) 0.976 (8) 0.727 (5) 0.965 (6) 0.743 (6) 0.091 (5) 0.407 (7) 0.929 (9) 0.756 (11) 0.511 (13) 0.870 (8) 0.999 (10) 0.919 (10) 0.822 (6) 0.948 (10) 0.709 (5) 0.830 (4) 0.857 (6) 0.814 (13) 0.989 (5) 0.987 (7) 0.000 (18) 0.898 (6) 0.032 (14) 0.847 (6) 0.904 (7) 0.823 (3) 0.858 (8) 0.592 (12) 0.360 (6) 0.866 (10) 0.960 (9) 0.720 (7) 0.964 (12) 0.837 (4) 0.506 (7) 0.970 (11) 0.476 (8) 0.841 (11) 0.792 (11) 0.232 (8) 0.417 (5)
0.899 (11) 0.988 (11) 0.652 (3) 0.683 (5) 0.952 (1) 0.987 (4) 0.965 (10) 0.849 (3) 0.968 (5) 0.815 (3) 0.093 (4) 0.495 (4) 0.979 (2) 0.833 (4) 0.590 (11) 0.933 (3) 1.000 (1) 0.964 (8) 0.854 (3) 0.983 (6) 0.730 (3) 0.820 (5) 0.890 (4) 0.848 (11) 0.993 (4) 0.995 (3) 0.269 (14) 0.952 (3) 0.035 (12) 0.904 (2) 0.895 (9) 0.854 (2) 0.872 (6) 0.668 (10) 0.434 (4) 0.930 (2) 1.000 (1) 0.814 (5) 0.987 (9) 0.846 (3) 0.577 (1) 0.979 (6) 0.528 (5) 0.920 (4) 0.866 (3) 0.226 (9) 0.481 (4)
0.933 (9) 0.998 (6) 0.601 (7) 0.658 (8) 0.919 (4) 0.978 (6) 1.000 (5) 0.462 (14) 0.938 (7) 0.631 (8) 0.047 (15) 0.367 (8) 0.972 (3) 0.815 (7) 0.633 (3) 0.897 (5) 0.997 (11) 0.967 (7) 0.773 (11) 0.986 (5) 0.693 (6) 0.681 (10) 0.848 (7) 0.925 (6) 0.969 (14) 0.988 (6) 0.500 (9) 0.914 (5) 0.103 (5) 0.896 (4) 0.871 (12) 0.709 (7) 0.739 (10) 0.805 (7) 0.485 (3) 0.895 (7) 0.993 (7) 0.402 (9) 1.000 (4) 0.691 (9) 0.482 (8) 0.976 (7) 0.528 (4) 0.867 (9) 0.850 (6) 0.326 (3) 0.249 (12)
0.839 (14) 0.992 (9) 0.590 (10) 0.447 (16) 0.836 (9) 0.973 (8) 1.000 (6) 0.514 (13) 0.898 (12) 0.687 (7) 0.064 (8) 0.317 (10) 0.881 (11) 0.842 (3) 0.610 (6) 0.727 (12) 0.920 (13) 0.627 (14) 0.824 (5) 0.865 (13) 0.646 (12) 0.803 (7) 0.790 (14) 0.674 (14) 0.975 (9) 0.990 (5) 0.074 (17) 0.871 (10) 0.077 (8) 0.786 (10) 0.735 (14) 0.687 (8) 0.670 (14) 0.540 (13) 0.240 (11) 0.912 (3) 0.971 (8) 0.854 (4) 0.987 (8) 0.757 (7) 0.539 (5) 0.962 (13) 0.364 (10) 0.794 (12) 0.751 (13) 0.142 (14) 0.190 (13)
0.934 (8) 1.000 (2) 0.689 (2) 0.704 (3) 0.941 (2) 0.990 (1) 0.987 (7) 0.874 (2) 0.982 (1) 0.839 (2) 0.153 (2) 0.819 (1) 0.960 (6) 0.894 (1) 0.704 (1) 0.947 (2) 1.000 (7) 0.998 (1) 0.896 (2) 0.996 (1) 0.736 (2) 0.843 (2) 0.927 (2) 0.944 (2) 1.000 (1) 0.997 (2) 0.600 (2) 0.967 (1) 0.099 (6) 0.904 (3) 0.952 (1) 0.791 (5) 0.942 (2) 0.932 (1) 0.605 (2) 0.912 (4) 1.000 (5) 0.886 (2) 0.991 (6) 0.849 (2) 0.574 (2) 0.987 (3) 0.618 (1) 0.943 (1) 0.885 (1) 0.333 (1) 0.548 (1)
0.975 (1) 0.998 (5) 0.728 (1) 0.695 (4) 0.921 (3) 0.988 (3) 1.000 (3) 0.885 (1) 0.982 (2) 0.846 (1) 0.185 (1) 0.804 (2) 0.968 (4) 0.893 (2) 0.690 (2) 0.979 (1) 1.000 (5) 0.993 (5) 0.899 (1) 0.989 (3) 0.778 (1) 0.849 (1) 0.933 (1) 0.959 (1) 0.999 (2) 0.999 (1) 0.516 (5) 0.964 (2) 0.058 (10) 0.912 (1) 0.933 (2) 0.885 (1) 0.967 (1) 0.900 (2) 0.637 (1) 0.958 (1) 0.997 (6) 0.860 (3) 0.977 (10) 0.865 (1) 0.567 (3) 0.990 (1) 0.603 (2) 0.927 (2) 0.872 (2) 0.326 (2) 0.513 (2)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.760 (9) 0.795 (9) 0.827 (9) 0.885 (7) 0.132 (10)
0.600 (13) 0.654 (13) 0.753 (13) 0.749 (14) 0.113 (13)
0.742 (10) 0.795 (10) 0.862 (6) 0.851 (10) 0.144 (8)
0.394 (16) 0.429 (16) 0.226 (16) 0.491 (18) 0.059 (17)
0.116 (18) 0.301 (18) 0.142 (18) 0.686 (15) 0.056 (18)
0.821 (5) 0.847 (4) 0.873 (4) 0.908 (2) 0.181 (5)
0.807 (6) 0.836 (6) 0.879 (2) 0.896 (4) 0.184 (4)
0.824 (3) 0.847 (3) 0.872 (5) 0.887 (6) 0.188 (3)
0.620 (12) 0.692 (12) 0.846 (8) 0.802 (13) 0.152 (6)
0.798 (7) 0.833 (7) 0.858 (7) 0.907 (3) 0.151 (7)
0.362 (17) 0.383 (17) 0.317 (15) 0.498 (17) 0.061 (16)
0.420 (15) 0.467 (15) 0.198 (17) 0.609 (16) 0.067 (15)
0.690 (11) 0.736 (11) 0.795 (12) 0.882 (8) 0.129 (11)
0.768 (8) 0.813 (8) 0.824 (10) 0.891 (5) 0.115 (12)
0.824 (4) 0.845 (5) 0.809 (11) 0.877 (9) 0.113 (14)
0.535 (14) 0.547 (14) 0.744 (14) 0.834 (12) 0.141 (9)
0.836 (2) 0.858 (2) 0.890 (1) 0.922 (1) 0.218 (1)
0.841 (1) 0.861 (1) 0.879 (3) 0.841 (11) 0.200 (2)
20news agnews amazon imdb yelp
0.381 (9) 0.668 (9) 0.346 (9) 0.431 (8) 0.565 (10)
0.256 (14) 0.500 (13) 0.232 (14) 0.286 (14) 0.410 (13)
0.357 (11) 0.727 (7) 0.375 (7) 0.421 (10) 0.647 (7)
0.069 (18) 0.084 (18) 0.053 (17) 0.050 (17) 0.056 (16)
0.115 (16) 0.155 (16) 0.058 (15) 0.050 (18) 0.048 (18)
0.521 (4) 0.763 (3) 0.438 (4) 0.551 (4) 0.727 (1)
0.463 (7) 0.744 (6) 0.461 (3) 0.557 (3) 0.719 (3)
0.529 (2) 0.766 (2) 0.426 (5) 0.545 (5) 0.720 (2)
0.378 (10) 0.624 (12) 0.398 (6) 0.429 (9) 0.566 (9)
0.488 (5) 0.748 (5) 0.371 (8) 0.499 (6) 0.673 (6)
0.095 (17) 0.154 (17) 0.050 (18) 0.055 (16) 0.052 (17)
0.158 (15) 0.315 (15) 0.055 (16) 0.055 (15) 0.061 (15)
0.384 (8) 0.653 (11) 0.340 (10) 0.409 (11) 0.552 (11)
0.340 (12) 0.663 (10) 0.337 (11) 0.438 (7) 0.587 (8)
0.523 (3) 0.681 (8) 0.251 (13) 0.394 (12) 0.546 (12)
0.337 (13) 0.428 (14) 0.320 (12) 0.286 (13) 0.409 (14)
0.531 (1) 0.774 (1) 0.532 (1) 0.575 (1) 0.701 (4)
0.465 (6) 0.753 (4) 0.479 (2) 0.558 (2) 0.697 (5)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 39: The average AUCPR performance on individual datasets under the 𝛾𝑙𝑎 = 50% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.911 (14) 0.997 (11) 0.679 (4) 0.670 (9) 0.969 (5) 0.994 (10) 0.967 (9) 0.853 (5) 0.987 (3) 0.840 (4) 0.144 (4) 0.593 (4) 1.000 (11) 0.861 (5) 0.627 (7) 0.943 (5) 1.000 (12) 0.977 (9) 0.857 (7) 0.981 (8) 0.797 (4) 0.865 (3) 0.922 (4) 0.909 (8) 0.997 (3) 0.996 (5) 0.701 (7) 0.960 (4) 0.065 (11) 0.960 (4) 0.961 (3) 0.917 (5) 0.768 (14) 0.770 (12) 0.444 (5) 0.977 (1) 0.962 (10) 0.837 (6) 1.000 (13) 0.917 (5) 0.548 (5) 0.971 (10) 0.588 (3) 0.941 (4) 0.886 (5) 0.300 (7) 0.551 (4)
0.956 (7) 1.000 (4) 0.634 (10) 0.575 (14) 0.554 (12) 0.992 (11) 0.957 (11) 0.799 (9) 0.967 (8) 0.822 (5) 0.074 (7) 0.364 (11) 0.972 (13) 0.854 (6) 0.638 (6) 0.935 (7) 1.000 (2) 0.992 (6) 0.854 (9) 0.997 (3) 0.646 (13) 0.718 (10) 0.907 (5) 0.936 (4) 0.983 (8) 0.995 (6) 0.769 (6) 0.895 (10) 0.042 (15) 0.708 (13) 0.872 (13) 0.594 (11) 0.872 (8) 0.608 (14) 0.418 (6) 0.813 (14) 0.927 (11) 0.765 (8) 0.941 (14) 0.746 (11) 0.501 (8) 0.963 (13) 0.397 (13) 0.884 (13) 0.865 (7) 0.277 (10) 0.356 (10)
0.967 (2) 1.000 (2) 0.667 (6) 0.680 (6) 0.693 (10) 1.000 (3) 0.933 (12) 0.820 (8) 0.981 (6) 0.481 (11) 0.062 (9) 0.559 (7) 1.000 (7) 0.795 (11) 0.622 (9) 0.921 (8) 0.990 (13) 0.954 (11) 0.822 (10) 0.996 (5) 0.683 (12) 0.699 (11) 0.798 (13) 0.916 (6) 0.976 (10) 0.543 (14) 0.700 (10) 0.850 (13) 0.241 (1) 0.922 (9) 0.905 (10) 0.643 (10) 0.870 (9) 0.901 (7) 0.172 (14) 0.900 (11) 0.925 (12) 0.118 (11) 1.000 (7) 0.761 (10) 0.418 (13) 0.978 (9) 0.442 (10) 0.903 (10) 0.857 (10) 0.271 (11) 0.381 (9)
0.258 (16) 0.135 (15) 0.479 (17) 0.373 (18) 0.098 (18) 0.361 (17) 0.472 (16) 0.701 (12) 0.781 (18) 0.160 (18) 0.052 (14) 0.070 (18) 0.204 (18) 0.564 (17) 0.039 (17) 0.596 (16) 0.099 (18) 0.026 (18) 0.391 (18) 0.058 (18) 0.458 (18) 0.154 (17) 0.589 (17) 0.633 (16) 0.737 (17) 0.181 (18) 0.330 (16) 0.688 (14) 0.019 (18) 0.223 (18) 0.281 (18) 0.118 (16) 0.359 (18) 0.040 (18) 0.054 (17) 0.759 (15) 0.345 (18) 0.071 (16) 0.079 (18) 0.230 (18) 0.308 (17) 0.965 (12) 0.314 (15) 0.521 (17) 0.520 (15) 0.094 (15) 0.102 (17)
0.010 (18) 0.076 (18) 0.522 (16) 0.567 (15) 0.191 (17) 0.331 (18) 0.547 (15) 0.473 (16) 0.873 (17) 0.248 (15) 0.041 (18) 0.169 (17) 0.708 (15) 0.448 (18) 0.135 (15) 0.225 (18) 0.765 (16) 0.029 (17) 0.426 (16) 0.192 (16) 0.478 (16) 0.314 (15) 0.631 (16) 0.520 (17) 0.315 (18) 0.206 (17) 0.130 (17) 0.429 (18) 0.020 (17) 0.303 (17) 0.517 (15) 0.103 (17) 0.431 (17) 0.331 (16) 0.043 (18) 0.425 (17) 0.643 (15) 0.093 (12) 0.191 (17) 0.245 (17) 0.313 (16) 0.880 (17) 0.245 (18) 0.416 (18) 0.398 (18) 0.061 (16) 0.080 (18)
0.960 (5) 0.999 (8) 0.588 (14) 0.714 (2) 0.286 (15) 0.881 (13) 1.000 (1) 0.850 (6) 0.927 (15) 0.445 (12) 0.053 (12) 0.271 (14) 1.000 (2) 0.743 (14) 0.626 (8) 0.890 (10) 1.000 (3) 0.995 (5) 0.697 (13) 0.929 (13) 0.694 (10) 0.458 (13) 0.830 (11) 0.908 (9) 0.971 (14) 0.660 (13) 0.800 (4) 0.877 (12) 0.192 (4) 0.715 (12) 0.918 (8) 0.440 (14) 0.888 (6) 0.906 (6) 0.182 (13) 0.925 (10) 1.000 (2) 0.088 (14) 1.000 (2) 0.628 (13) 0.496 (10) 0.990 (3) 0.518 (8) 0.931 (7) 0.857 (11) 0.314 (4) 0.459 (7)
0.957 (6) 1.000 (6) 0.649 (8) 0.707 (3) 0.909 (9) 1.000 (9) 1.000 (2) 0.760 (10) 0.963 (9) 0.626 (9) 0.059 (10) 0.407 (10) 1.000 (4) 0.795 (10) 0.605 (11) 0.860 (12) 1.000 (5) 0.997 (3) 0.806 (11) 0.968 (10) 0.722 (8) 0.733 (9) 0.835 (9) 0.915 (7) 0.977 (9) 0.868 (10) 0.800 (3) 0.899 (8) 0.157 (7) 0.845 (10) 0.921 (7) 0.705 (9) 0.904 (5) 0.934 (4) 0.314 (8) 0.940 (7) 1.000 (3) 0.397 (10) 1.000 (4) 0.789 (9) 0.541 (7) 0.990 (1) 0.469 (9) 0.935 (5) 0.859 (8) 0.307 (6) 0.437 (8)
0.954 (8) 0.984 (13) 0.628 (12) 0.672 (8) 0.325 (14) 0.926 (12) 1.000 (4) 0.854 (4) 0.935 (13) 0.518 (10) 0.052 (13) 0.307 (12) 1.000 (6) 0.758 (13) 0.613 (10) 0.910 (9) 1.000 (7) 0.996 (4) 0.700 (12) 0.929 (12) 0.703 (9) 0.459 (12) 0.827 (12) 0.896 (11) 0.971 (12) 0.679 (11) 0.800 (5) 0.899 (9) 0.196 (3) 0.764 (11) 0.914 (9) 0.472 (12) 0.880 (7) 0.913 (5) 0.226 (12) 0.868 (13) 1.000 (5) 0.090 (13) 1.000 (6) 0.665 (12) 0.500 (9) 0.986 (6) 0.526 (7) 0.935 (6) 0.858 (9) 0.312 (5) 0.479 (6)
0.944 (10) 1.000 (5) 0.647 (9) 0.622 (12) 0.973 (4) 1.000 (4) 0.899 (13) 0.660 (13) 0.906 (16) 0.227 (17) 0.097 (6) 0.583 (5) 0.965 (14) 0.840 (7) 0.600 (12) 0.829 (13) 1.000 (8) 0.931 (13) 0.879 (5) 0.994 (6) 0.693 (11) 0.842 (7) 0.832 (10) 0.880 (12) 0.983 (7) 0.980 (9) 0.672 (11) 0.883 (11) 0.210 (2) 0.952 (7) 0.934 (5) 0.821 (6) 0.820 (13) 0.901 (8) 0.283 (10) 0.886 (12) 0.891 (14) 0.909 (2) 1.000 (8) 0.867 (7) 0.471 (12) 0.969 (11) 0.424 (12) 0.747 (14) 0.828 (12) 0.236 (12) 0.326 (12)
0.951 (9) 0.838 (14) 0.617 (13) 0.689 (4) 0.614 (11) 0.786 (14) 0.601 (14) 0.704 (11) 0.928 (14) 0.389 (13) 0.052 (15) 0.224 (16) 1.000 (3) 0.772 (12) 0.407 (14) 0.662 (14) 1.000 (4) 0.934 (12) 0.658 (14) 0.855 (14) 0.592 (14) 0.444 (14) 0.791 (14) 0.942 (2) 0.974 (11) 0.663 (12) 0.403 (13) 0.614 (15) 0.132 (9) 0.669 (14) 0.888 (11) 0.445 (13) 0.861 (10) 0.864 (10) 0.229 (11) 0.926 (9) 0.924 (13) 0.080 (15) 1.000 (3) 0.487 (14) 0.380 (14) 0.980 (7) 0.302 (16) 0.909 (9) 0.822 (14) 0.203 (14) 0.195 (14)
0.267 (15) 0.089 (17) 0.477 (18) 0.451 (17) 0.354 (13) 0.405 (16) 0.426 (17) 0.423 (18) 0.935 (12) 0.363 (14) 0.048 (17) 0.281 (13) 0.647 (17) 0.607 (16) 0.026 (18) 0.628 (15) 0.878 (15) 0.208 (15) 0.392 (17) 0.442 (15) 0.470 (17) 0.160 (16) 0.645 (15) 0.836 (14) 0.943 (16) 0.387 (15) 0.367 (14) 0.432 (17) 0.051 (14) 0.553 (15) 0.442 (16) 0.242 (15) 0.521 (15) 0.491 (15) 0.168 (15) 0.419 (18) 0.584 (17) 0.045 (17) 0.380 (15) 0.407 (15) 0.344 (15) 0.782 (18) 0.344 (14) 0.572 (15) 0.460 (17) 0.047 (17) 0.130 (15)
0.120 (17) 0.133 (16) 0.584 (15) 0.495 (16) 0.200 (16) 0.677 (15) 0.293 (18) 0.438 (17) 0.943 (11) 0.237 (16) 0.049 (16) 0.232 (15) 0.701 (16) 0.681 (15) 0.125 (16) 0.461 (17) 0.764 (17) 0.064 (16) 0.564 (15) 0.191 (17) 0.562 (15) 0.134 (18) 0.588 (18) 0.494 (18) 0.951 (15) 0.326 (16) 0.462 (12) 0.441 (16) 0.027 (16) 0.353 (16) 0.298 (17) 0.099 (18) 0.439 (16) 0.293 (17) 0.083 (16) 0.732 (16) 0.594 (16) 0.037 (18) 0.289 (16) 0.246 (16) 0.299 (18) 0.949 (16) 0.272 (17) 0.552 (16) 0.509 (16) 0.038 (18) 0.108 (16)
0.934 (11) 0.993 (12) 0.669 (5) 0.673 (7) 0.958 (7) 1.000 (8) 0.968 (8) 0.838 (7) 0.985 (4) 0.821 (6) 0.125 (5) 0.566 (6) 1.000 (10) 0.817 (8) 0.590 (13) 0.937 (6) 1.000 (11) 0.964 (10) 0.866 (6) 0.976 (9) 0.813 (3) 0.863 (4) 0.902 (6) 0.866 (13) 0.995 (5) 0.994 (8) 0.000 (18) 0.943 (5) 0.061 (12) 0.965 (3) 0.933 (6) 0.941 (3) 0.934 (4) 0.745 (13) 0.417 (7) 0.936 (8) 0.976 (9) 0.821 (7) 1.000 (12) 0.934 (4) 0.541 (6) 0.980 (8) 0.528 (6) 0.921 (8) 0.887 (4) 0.284 (9) 0.482 (5)
0.923 (13) 0.997 (10) 0.729 (3) 0.760 (1) 0.985 (1) 1.000 (1) 0.965 (10) 0.899 (3) 0.984 (5) 0.863 (3) 0.155 (3) 0.653 (3) 1.000 (1) 0.876 (4) 0.681 (3) 0.965 (3) 1.000 (1) 0.981 (7) 0.888 (3) 0.997 (4) 0.820 (2) 0.857 (5) 0.933 (3) 0.908 (10) 0.997 (4) 0.997 (4) 0.900 (1) 0.966 (2) 0.054 (13) 0.971 (1) 0.965 (1) 0.943 (2) 0.956 (3) 0.865 (9) 0.484 (4) 0.967 (2) 1.000 (1) 0.874 (5) 1.000 (1) 0.936 (3) 0.603 (3) 0.987 (5) 0.567 (4) 0.946 (2) 0.906 (3) 0.299 (8) 0.554 (3)
0.963 (4) 1.000 (7) 0.633 (11) 0.634 (10) 0.967 (6) 1.000 (5) 1.000 (5) 0.620 (15) 0.975 (7) 0.659 (8) 0.057 (11) 0.537 (8) 1.000 (8) 0.817 (9) 0.660 (5) 0.949 (4) 1.000 (9) 0.981 (8) 0.857 (8) 0.992 (7) 0.722 (7) 0.831 (8) 0.853 (8) 0.918 (5) 0.971 (13) 0.997 (3) 0.700 (9) 0.929 (6) 0.179 (5) 0.956 (6) 0.874 (12) 0.775 (7) 0.838 (11) 0.953 (3) 0.534 (3) 0.961 (5) 1.000 (6) 0.398 (9) 1.000 (9) 0.813 (8) 0.481 (11) 0.962 (14) 0.559 (5) 0.897 (11) 0.883 (6) 0.349 (2) 0.334 (11)
0.924 (12) 0.999 (9) 0.657 (7) 0.618 (13) 0.917 (8) 1.000 (7) 1.000 (7) 0.652 (14) 0.945 (10) 0.771 (7) 0.068 (8) 0.510 (9) 0.986 (12) 0.879 (3) 0.674 (4) 0.864 (11) 0.985 (14) 0.923 (14) 0.881 (4) 0.964 (11) 0.724 (6) 0.848 (6) 0.858 (7) 0.786 (15) 0.985 (6) 0.994 (7) 0.339 (15) 0.914 (7) 0.102 (10) 0.925 (8) 0.742 (14) 0.771 (8) 0.822 (12) 0.830 (11) 0.314 (9) 0.966 (3) 1.000 (8) 0.891 (4) 1.000 (11) 0.903 (6) 0.585 (4) 0.959 (15) 0.438 (11) 0.887 (12) 0.824 (13) 0.207 (13) 0.255 (13)
0.966 (3) 1.000 (3) 0.774 (2) 0.688 (5) 0.984 (3) 1.000 (6) 1.000 (6) 0.918 (1) 0.990 (2) 0.893 (2) 0.234 (2) 0.910 (1) 1.000 (9) 0.910 (1) 0.731 (1) 0.977 (2) 1.000 (10) 0.999 (2) 0.928 (1) 0.999 (1) 0.775 (5) 0.891 (2) 0.953 (2) 0.940 (3) 1.000 (1) 0.998 (2) 0.701 (8) 0.972 (1) 0.161 (6) 0.959 (5) 0.962 (2) 0.927 (4) 0.964 (2) 0.984 (1) 0.636 (1) 0.952 (6) 1.000 (7) 0.921 (1) 1.000 (10) 0.938 (2) 0.615 (2) 0.989 (4) 0.641 (1) 0.962 (1) 0.924 (1) 0.345 (3) 0.603 (1)
0.985 (1) 1.000 (1) 0.787 (1) 0.625 (11) 0.985 (2) 1.000 (2) 1.000 (3) 0.914 (2) 0.990 (1) 0.899 (1) 0.312 (1) 0.878 (2) 1.000 (5) 0.905 (2) 0.716 (2) 0.990 (1) 1.000 (6) 1.000 (1) 0.921 (2) 0.998 (2) 0.838 (1) 0.900 (1) 0.955 (1) 0.955 (1) 1.000 (2) 0.999 (1) 0.900 (2) 0.961 (3) 0.135 (8) 0.969 (2) 0.956 (4) 0.958 (1) 0.985 (1) 0.962 (2) 0.635 (2) 0.963 (4) 1.000 (4) 0.905 (3) 1.000 (5) 0.944 (1) 0.615 (1) 0.990 (2) 0.627 (2) 0.943 (3) 0.920 (2) 0.367 (1) 0.572 (2)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.813 (10) 0.849 (10) 0.856 (10) 0.948 (6) 0.174 (11)
0.731 (13) 0.792 (13) 0.851 (11) 0.895 (14) 0.174 (10)
0.831 (7) 0.872 (3) 0.880 (5) 0.937 (11) 0.198 (7)
0.388 (16) 0.436 (16) 0.223 (17) 0.534 (18) 0.059 (17)
0.117 (18) 0.291 (18) 0.142 (18) 0.698 (16) 0.056 (18)
0.849 (5) 0.870 (4) 0.880 (6) 0.949 (5) 0.224 (5)
0.846 (6) 0.869 (5) 0.886 (4) 0.966 (2) 0.238 (3)
0.851 (4) 0.866 (6) 0.879 (7) 0.939 (9) 0.229 (4)
0.745 (12) 0.805 (12) 0.891 (3) 0.916 (13) 0.211 (6)
0.815 (9) 0.856 (9) 0.865 (8) 0.951 (4) 0.175 (9)
0.353 (17) 0.415 (17) 0.421 (15) 0.556 (17) 0.060 (16)
0.454 (15) 0.520 (15) 0.224 (16) 0.711 (15) 0.067 (15)
0.776 (11) 0.821 (11) 0.846 (12) 0.946 (7) 0.166 (13)
0.831 (8) 0.857 (8) 0.860 (9) 0.952 (3) 0.167 (12)
0.854 (3) 0.865 (7) 0.794 (14) 0.945 (8) 0.116 (14)
0.642 (14) 0.678 (14) 0.837 (13) 0.933 (12) 0.189 (8)
0.879 (1) 0.898 (1) 0.905 (1) 0.972 (1) 0.300 (1)
0.878 (2) 0.892 (2) 0.891 (2) 0.937 (10) 0.260 (2)
20news agnews amazon imdb yelp
0.491 (9) 0.728 (9) 0.403 (9) 0.511 (10) 0.649 (11)
0.426 (14) 0.650 (12) 0.366 (13) 0.433 (12) 0.590 (12)
0.512 (8) 0.794 (2) 0.454 (6) 0.563 (7) 0.765 (5)
0.071 (18) 0.085 (18) 0.053 (17) 0.051 (17) 0.056 (17)
0.118 (17) 0.152 (17) 0.061 (15) 0.050 (18) 0.051 (18)
0.612 (5) 0.772 (6) 0.468 (5) 0.594 (4) 0.766 (4)
0.569 (6) 0.772 (5) 0.515 (3) 0.606 (3) 0.772 (2)
0.630 (2) 0.775 (4) 0.453 (7) 0.587 (5) 0.758 (6)
0.478 (10) 0.725 (10) 0.513 (4) 0.564 (6) 0.701 (8)
0.617 (3) 0.762 (7) 0.402 (10) 0.543 (8) 0.708 (7)
0.142 (16) 0.179 (16) 0.053 (18) 0.079 (15) 0.057 (16)
0.165 (15) 0.339 (15) 0.055 (16) 0.056 (16) 0.064 (15)
0.472 (11) 0.716 (11) 0.401 (11) 0.483 (11) 0.657 (10)
0.467 (12) 0.733 (8) 0.433 (8) 0.531 (9) 0.685 (9)
0.615 (4) 0.552 (14) 0.183 (14) 0.222 (14) 0.464 (14)
0.443 (13) 0.580 (13) 0.377 (12) 0.405 (13) 0.524 (13)
0.639 (1) 0.816 (1) 0.591 (1) 0.650 (1) 0.800 (1)
0.564 (7) 0.791 (3) 0.524 (2) 0.610 (2) 0.766 (3)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 40: The average AUCPR performance on individual datasets under the 𝛾𝑙𝑎 = 100% setting. Dataset
XGBOD
DeepSAD
REPEN
AA- BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.974 (10) 1.000 (10) 0.773 (6) 0.722 (4) 0.990 (8) 0.996 (10) 1.000 (9) 0.903 (10) 0.998 (3) 0.886 (6) 0.186 (5) 0.768 (5) 1.000 (15) 0.888 (6) 0.705 (6) 0.978 (9) 1.000 (16) 0.992 (10) 0.918 (7) 0.995 (9) 0.865 (5) 0.939 (3) 0.945 (6) 0.952 (4) 1.000 (5) 0.999 (3) 0.701 (11) 0.974 (5) 0.157 (11) 0.995 (5) 0.980 (2) 0.991 (5) 0.876 (13) 0.905 (12) 0.559 (6) 1.000 (9) 1.000 (14) 0.919 (5) 1.000 (13) 0.995 (6) 0.618 (5) 0.997 (6) 0.624 (3) 0.985 (2) 0.954 (4) 0.380 (4) 0.619 (4)
0.984 (6) 1.000 (8) 0.729 (10) 0.690 (10) 0.708 (10) 1.000 (2) 0.999 (10) 0.932 (4) 0.991 (9) 0.861 (7) 0.094 (8) 0.444 (10) 1.000 (3) 0.882 (7) 0.659 (8) 0.971 (10) 1.000 (3) 0.999 (4) 0.902 (8) 0.997 (7) 0.726 (10) 0.789 (9) 0.941 (7) 0.940 (7) 0.992 (7) 0.991 (7) 0.803 (3) 0.952 (9) 0.072 (14) 0.833 (11) 0.898 (13) 0.735 (9) 0.958 (8) 0.819 (14) 0.538 (7) 0.896 (14) 1.000 (2) 0.825 (8) 0.992 (14) 0.879 (10) 0.536 (10) 0.994 (9) 0.504 (12) 0.954 (11) 0.918 (10) 0.360 (5) 0.515 (7)
0.983 (7) 1.000 (3) 0.730 (9) 0.688 (11) 0.666 (11) 1.000 (4) 0.949 (12) 0.937 (3) 0.994 (6) 0.501 (12) 0.073 (10) 0.653 (9) 1.000 (9) 0.815 (10) 0.644 (9) 0.981 (7) 1.000 (10) 0.985 (12) 0.850 (10) 0.998 (6) 0.694 (13) 0.753 (10) 0.831 (11) 0.941 (6) 0.978 (9) 0.517 (14) 0.800 (8) 0.914 (11) 0.409 (1) 0.911 (9) 0.924 (9) 0.668 (11) 0.958 (7) 0.947 (8) 0.209 (13) 1.000 (4) 1.000 (7) 0.123 (11) 1.000 (7) 0.881 (9) 0.465 (13) 0.996 (8) 0.532 (9) 0.981 (5) 0.920 (9) 0.313 (11) 0.516 (6)
0.133 (17) 0.064 (18) 0.518 (18) 0.467 (18) 0.077 (18) 0.465 (17) 0.243 (18) 0.694 (16) 0.743 (18) 0.164 (18) 0.054 (12) 0.089 (18) 0.519 (18) 0.599 (16) 0.022 (18) 0.507 (17) 0.417 (18) 0.177 (17) 0.427 (18) 0.128 (18) 0.495 (18) 0.199 (16) 0.620 (18) 0.720 (17) 0.928 (17) 0.205 (18) 0.305 (16) 0.418 (18) 0.036 (15) 0.229 (18) 0.267 (18) 0.133 (16) 0.624 (16) 0.054 (18) 0.070 (17) 0.825 (16) 0.526 (18) 0.089 (14) 0.239 (18) 0.243 (18) 0.307 (17) 0.974 (16) 0.341 (16) 0.541 (17) 0.474 (17) 0.078 (16) 0.066 (18)
0.009 (18) 0.089 (17) 0.553 (16) 0.572 (14) 0.194 (17) 0.423 (18) 0.447 (15) 0.520 (17) 0.891 (17) 0.260 (17) 0.041 (18) 0.172 (17) 0.743 (17) 0.566 (17) 0.134 (16) 0.284 (18) 1.000 (7) 0.033 (18) 0.573 (16) 0.209 (17) 0.499 (17) 0.402 (15) 0.771 (15) 0.600 (18) 0.660 (18) 0.244 (17) 0.130 (18) 0.474 (17) 0.020 (18) 0.333 (17) 0.548 (15) 0.109 (18) 0.484 (18) 0.357 (16) 0.044 (18) 0.439 (18) 0.709 (17) 0.092 (12) 0.340 (17) 0.260 (17) 0.315 (16) 0.931 (17) 0.346 (15) 0.510 (18) 0.469 (18) 0.046 (17) 0.088 (16)
0.966 (13) 0.999 (12) 0.622 (15) 0.718 (6) 0.306 (15) 0.870 (14) 1.000 (2) 0.919 (8) 0.940 (15) 0.474 (13) 0.051 (15) 0.306 (14) 1.000 (4) 0.759 (13) 0.638 (10) 0.894 (12) 1.000 (4) 0.995 (8) 0.712 (13) 0.923 (14) 0.699 (12) 0.460 (12) 0.820 (13) 0.907 (16) 0.971 (13) 0.663 (13) 0.800 (6) 0.887 (13) 0.172 (7) 0.731 (14) 0.918 (11) 0.429 (13) 0.888 (11) 0.922 (10) 0.209 (15) 0.941 (11) 1.000 (3) 0.089 (15) 1.000 (2) 0.676 (13) 0.487 (12) 0.992 (12) 0.529 (10) 0.938 (13) 0.856 (13) 0.330 (8) 0.508 (9)
0.969 (12) 1.000 (9) 0.709 (11) 0.721 (5) 0.875 (9) 0.996 (11) 1.000 (3) 0.920 (7) 0.985 (11) 0.629 (9) 0.066 (11) 0.420 (11) 1.000 (6) 0.815 (9) 0.613 (13) 0.877 (13) 1.000 (6) 0.997 (6) 0.825 (11) 0.964 (11) 0.739 (9) 0.713 (11) 0.847 (9) 0.910 (13) 0.975 (11) 0.830 (10) 0.801 (5) 0.934 (10) 0.169 (8) 0.845 (10) 0.923 (10) 0.703 (10) 0.934 (10) 0.948 (7) 0.327 (10) 1.000 (2) 1.000 (4) 0.392 (10) 1.000 (4) 0.864 (11) 0.539 (8) 0.993 (11) 0.501 (13) 0.967 (8) 0.891 (11) 0.324 (9) 0.491 (10)
0.972 (11) 0.989 (13) 0.664 (12) 0.712 (7) 0.340 (14) 0.949 (13) 1.000 (5) 0.927 (6) 0.946 (14) 0.523 (11) 0.048 (17) 0.339 (12) 1.000 (8) 0.763 (12) 0.620 (12) 0.916 (11) 1.000 (9) 0.996 (7) 0.712 (12) 0.932 (12) 0.720 (11) 0.456 (14) 0.824 (12) 0.907 (14) 0.970 (14) 0.678 (12) 0.800 (9) 0.902 (12) 0.165 (9) 0.766 (12) 0.914 (12) 0.472 (12) 0.882 (12) 0.910 (11) 0.234 (12) 0.870 (15) 1.000 (6) 0.090 (13) 1.000 (6) 0.703 (12) 0.490 (11) 0.991 (14) 0.541 (7) 0.943 (12) 0.871 (12) 0.320 (10) 0.508 (8)
0.979 (9) 1.000 (4) 0.756 (7) 0.542 (15) 1.000 (3) 1.000 (5) 0.899 (13) 0.893 (11) 0.992 (7) 0.617 (10) 0.122 (7) 0.776 (4) 1.000 (10) 0.874 (8) 0.621 (11) 0.984 (5) 1.000 (11) 0.999 (3) 0.890 (9) 0.999 (4) 0.781 (7) 0.858 (8) 0.838 (10) 0.926 (12) 0.987 (8) 0.988 (8) 0.531 (13) 0.956 (8) 0.277 (2) 0.979 (7) 0.943 (8) 0.920 (7) 0.944 (9) 0.983 (4) 0.433 (9) 1.000 (5) 1.000 (8) 0.917 (6) 1.000 (8) 0.923 (8) 0.561 (7) 0.980 (15) 0.506 (11) 0.967 (9) 0.933 (8) 0.281 (12) 0.464 (12)
0.959 (14) 0.942 (14) 0.645 (13) 0.677 (12) 0.541 (12) 0.866 (15) 0.666 (14) 0.776 (14) 0.959 (13) 0.364 (15) 0.051 (14) 0.239 (15) 1.000 (5) 0.792 (11) 0.390 (14) 0.794 (14) 1.000 (5) 0.956 (14) 0.667 (14) 0.925 (13) 0.633 (14) 0.460 (13) 0.804 (14) 0.938 (9) 0.973 (12) 0.703 (11) 0.305 (17) 0.713 (14) 0.128 (13) 0.744 (13) 0.888 (14) 0.417 (14) 0.855 (14) 0.857 (13) 0.246 (11) 0.933 (12) 0.986 (15) 0.077 (16) 1.000 (3) 0.524 (14) 0.415 (14) 0.993 (10) 0.423 (14) 0.928 (14) 0.837 (14) 0.231 (14) 0.325 (14)
0.266 (15) 0.172 (16) 0.553 (17) 0.508 (16) 0.440 (13) 0.609 (16) 0.391 (16) 0.430 (18) 0.931 (16) 0.414 (14) 0.053 (13) 0.307 (13) 0.866 (16) 0.621 (15) 0.042 (17) 0.626 (16) 0.953 (17) 0.440 (16) 0.466 (17) 0.524 (16) 0.570 (16) 0.198 (17) 0.720 (17) 0.942 (5) 0.968 (15) 0.448 (15) 0.600 (12) 0.517 (15) 0.032 (16) 0.702 (15) 0.451 (16) 0.241 (15) 0.622 (17) 0.511 (15) 0.209 (14) 0.735 (17) 1.000 (9) 0.043 (17) 0.456 (16) 0.490 (15) 0.330 (15) 0.828 (18) 0.333 (17) 0.663 (16) 0.599 (16) 0.088 (15) 0.084 (17)
0.145 (16) 0.545 (15) 0.637 (14) 0.501 (17) 0.237 (16) 0.982 (12) 0.296 (17) 0.763 (15) 0.973 (12) 0.314 (16) 0.049 (16) 0.239 (16) 1.000 (1) 0.732 (14) 0.135 (15) 0.719 (15) 1.000 (1) 0.564 (15) 0.621 (15) 0.758 (15) 0.611 (15) 0.149 (18) 0.747 (16) 0.967 (1) 0.965 (16) 0.443 (16) 0.457 (14) 0.482 (16) 0.027 (17) 0.434 (16) 0.321 (17) 0.111 (17) 0.735 (15) 0.341 (17) 0.096 (16) 0.912 (13) 0.931 (16) 0.037 (18) 0.859 (15) 0.314 (16) 0.306 (18) 0.992 (13) 0.312 (18) 0.725 (15) 0.641 (15) 0.038 (18) 0.116 (15)
0.980 (8) 1.000 (11) 0.823 (4) 0.697 (8) 1.000 (7) 1.000 (9) 0.980 (11) 0.912 (9) 0.996 (5) 0.907 (4) 0.212 (4) 0.749 (6) 1.000 (14) 0.898 (5) 0.708 (5) 0.985 (4) 1.000 (15) 0.984 (13) 0.922 (6) 0.990 (10) 0.899 (4) 0.923 (5) 0.956 (4) 0.907 (15) 1.000 (4) 0.999 (6) 0.331 (15) 0.976 (4) 0.158 (10) 0.999 (4) 0.971 (5) 0.992 (4) 0.977 (4) 0.938 (9) 0.536 (8) 1.000 (8) 1.000 (13) 0.915 (7) 1.000 (12) 0.996 (5) 0.620 (4) 0.997 (7) 0.595 (5) 0.975 (7) 0.947 (6) 0.339 (7) 0.618 (5)
0.985 (5) 1.000 (1) 0.834 (3) 0.741 (2) 1.000 (1) 1.000 (1) 1.000 (1) 0.930 (5) 0.997 (4) 0.910 (3) 0.236 (3) 0.819 (3) 1.000 (2) 0.908 (3) 0.762 (3) 0.988 (3) 1.000 (2) 0.995 (9) 0.938 (3) 0.996 (8) 0.900 (3) 0.926 (4) 0.958 (3) 0.939 (8) 1.000 (1) 0.999 (4) 0.900 (1) 0.976 (3) 0.196 (5) 1.000 (1) 0.972 (4) 0.994 (3) 0.985 (3) 0.965 (6) 0.605 (4) 1.000 (1) 1.000 (1) 0.926 (4) 1.000 (1) 0.999 (2) 0.662 (2) 0.997 (4) 0.622 (4) 0.980 (6) 0.962 (3) 0.355 (6) 0.646 (2)
0.988 (4) 1.000 (5) 0.736 (8) 0.737 (3) 1.000 (4) 1.000 (6) 1.000 (6) 0.887 (12) 0.991 (8) 0.815 (8) 0.091 (9) 0.710 (8) 1.000 (11) 0.352 (18) 0.692 (7) 0.981 (8) 1.000 (12) 0.992 (11) 0.924 (4) 0.998 (5) 0.756 (8) 0.860 (7) 0.904 (8) 0.930 (11) 0.976 (10) 0.841 (9) 0.800 (7) 0.966 (6) 0.190 (6) 0.978 (8) 0.952 (6) 0.868 (8) 0.958 (6) 0.991 (3) 0.666 (3) 1.000 (6) 1.000 (10) 0.623 (9) 1.000 (9) 0.953 (7) 0.537 (9) 0.998 (3) 0.588 (6) 0.982 (4) 0.940 (7) 0.383 (3) 0.481 (11)
0.991 (3) 1.000 (7) 0.781 (5) 0.670 (13) 1.000 (6) 1.000 (8) 1.000 (8) 0.886 (13) 0.990 (10) 0.896 (5) 0.134 (6) 0.716 (7) 1.000 (13) 0.906 (4) 0.722 (4) 0.983 (6) 1.000 (14) 0.998 (5) 0.922 (5) 1.000 (3) 0.855 (6) 0.871 (6) 0.952 (5) 0.935 (10) 0.996 (6) 0.999 (5) 0.703 (10) 0.960 (7) 0.139 (12) 0.988 (6) 0.951 (7) 0.942 (6) 0.964 (5) 0.972 (5) 0.560 (5) 0.951 (10) 1.000 (12) 0.945 (3) 1.000 (11) 0.997 (4) 0.612 (6) 0.997 (5) 0.535 (8) 0.960 (10) 0.947 (5) 0.250 (13) 0.414 (13)
0.996 (2) 1.000 (6) 0.895 (2) 0.742 (1) 1.000 (5) 1.000 (7) 1.000 (7) 0.947 (1) 0.999 (1) 0.946 (2) 0.376 (2) 0.961 (2) 1.000 (12) 0.927 (1) 0.785 (1) 0.993 (2) 1.000 (13) 1.000 (2) 0.948 (1) 1.000 (1) 0.931 (2) 0.940 (2) 0.975 (2) 0.964 (2) 1.000 (3) 1.000 (2) 0.802 (4) 0.983 (2) 0.267 (3) 1.000 (3) 0.979 (3) 0.997 (1) 0.991 (2) 0.999 (1) 0.706 (1) 1.000 (7) 1.000 (11) 0.952 (1) 1.000 (10) 0.999 (3) 0.649 (3) 0.999 (2) 0.658 (1) 0.989 (1) 0.970 (2) 0.399 (2) 0.647 (1)
0.997 (1) 1.000 (2) 0.895 (1) 0.693 (9) 1.000 (2) 1.000 (3) 1.000 (4) 0.938 (2) 0.999 (2) 0.954 (1) 0.457 (1) 0.965 (1) 1.000 (7) 0.926 (2) 0.784 (2) 0.996 (1) 1.000 (8) 1.000 (1) 0.947 (2) 1.000 (2) 0.942 (1) 0.948 (1) 0.978 (1) 0.964 (3) 1.000 (2) 1.000 (1) 0.900 (2) 0.985 (1) 0.247 (4) 1.000 (2) 0.984 (1) 0.996 (2) 0.997 (1) 0.998 (2) 0.684 (2) 1.000 (3) 1.000 (5) 0.950 (2) 1.000 (5) 1.000 (1) 0.662 (1) 0.999 (1) 0.658 (2) 0.984 (3) 0.970 (1) 0.415 (1) 0.641 (3)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.872 (9) 0.888 (9) 0.876 (12) 0.991 (8) 0.220 (11)
0.891 (5) 0.906 (4) 0.914 (1) 0.994 (3) 0.288 (4)
0.894 (4) 0.915 (3) 0.890 (6) 0.993 (4) 0.253 (9)
0.395 (17) 0.438 (17) 0.242 (17) 0.633 (18) 0.059 (17)
0.120 (18) 0.338 (18) 0.207 (18) 0.702 (16) 0.057 (18)
0.867 (10) 0.885 (10) 0.883 (8) 0.966 (13) 0.257 (8)
0.875 (8) 0.893 (6) 0.890 (7) 0.985 (10) 0.285 (5)
0.862 (11) 0.874 (13) 0.882 (9) 0.961 (14) 0.259 (7)
0.876 (7) 0.890 (7) 0.906 (4) 0.993 (5) 0.299 (3)
0.826 (14) 0.864 (14) 0.870 (13) 0.974 (12) 0.191 (13)
0.411 (16) 0.538 (16) 0.484 (16) 0.638 (17) 0.061 (16)
0.629 (15) 0.718 (15) 0.504 (15) 0.927 (15) 0.068 (15)
0.861 (12) 0.881 (11) 0.877 (11) 0.991 (7) 0.215 (12)
0.876 (6) 0.888 (8) 0.882 (10) 0.992 (6) 0.230 (10)
0.898 (3) 0.894 (5) 0.788 (14) 0.980 (11) 0.124 (14)
0.841 (13) 0.879 (12) 0.909 (3) 0.987 (9) 0.273 (6)
0.923 (1) 0.927 (1) 0.914 (2) 0.998 (1) 0.368 (1)
0.909 (2) 0.919 (2) 0.900 (5) 0.996 (2) 0.309 (2)
20news agnews amazon imdb yelp
0.631 (11) 0.794 (9) 0.459 (12) 0.562 (11) 0.720 (11)
0.697 (5) 0.826 (4) 0.571 (4) 0.621 (6) 0.761 (8)
0.719 (3) 0.837 (2) 0.534 (6) 0.641 (4) 0.801 (2)
0.072 (18) 0.087 (18) 0.053 (17) 0.052 (18) 0.056 (18)
0.118 (17) 0.224 (16) 0.052 (18) 0.054 (17) 0.058 (17)
0.691 (8) 0.786 (11) 0.487 (8) 0.615 (7) 0.775 (6)
0.695 (6) 0.796 (7) 0.536 (5) 0.637 (5) 0.788 (5)
0.692 (7) 0.785 (12) 0.468 (11) 0.593 (8) 0.766 (7)
0.627 (12) 0.810 (5) 0.601 (2) 0.653 (3) 0.793 (4)
0.685 (9) 0.774 (13) 0.415 (13) 0.567 (10) 0.718 (12)
0.160 (16) 0.180 (17) 0.054 (16) 0.057 (15) 0.082 (15)
0.195 (15) 0.389 (15) 0.057 (15) 0.057 (16) 0.066 (16)
0.623 (14) 0.795 (8) 0.468 (10) 0.561 (12) 0.728 (10)
0.648 (10) 0.799 (6) 0.497 (7) 0.588 (9) 0.754 (9)
0.709 (4) 0.511 (14) 0.162 (14) 0.262 (14) 0.453 (14)
0.623 (13) 0.788 (10) 0.479 (9) 0.539 (13) 0.679 (13)
0.762 (1) 0.854 (1) 0.644 (1) 0.697 (1) 0.835 (1)
0.723 (2) 0.827 (3) 0.582 (3) 0.665 (2) 0.800 (3)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 41: The average AUCROC performance on individual datasets under the 𝛾𝑙𝑎 = 1% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.942 (8) 0.945 (11) 0.588 (12) 0.924 (7) 0.822 (12) 0.688 (9) 0.991 (14) 0.653 (9) 0.743 (11) 0.786 (5) 0.583 (3) 0.686 (8) 0.992 (3) 0.777 (7) 0.812 (12) 0.876 (6) 0.937 (8) 0.931 (9) 0.879 (6) 0.911 (7) 0.643 (6) 0.913 (5) 0.808 (8) 0.987 (6) 0.984 (6) 0.929 (12) 0.834 (9) 0.733 (6) 0.499 (14) 0.807 (11) 0.953 (8) 0.584 (11) 0.850 (11) 0.869 (6) 0.657 (11) 0.919 (11) 0.980 (8) 0.664 (10) 0.885 (10) 0.569 (9) 0.577 (6) 0.891 (9) 0.774 (5) 0.859 (9) 0.744 (9) 0.756 (4) 0.780 (4)
0.835 (13) 0.963 (9) 0.685 (2) 0.895 (11) 0.801 (13) 0.657 (12) 0.999 (8) 0.689 (3) 0.774 (10) 0.739 (8) 0.575 (4) 0.729 (5) 0.816 (13) 0.759 (9) 0.857 (9) 0.756 (14) 0.841 (15) 0.833 (11) 0.866 (7) 0.877 (10) 0.583 (12) 0.741 (13) 0.785 (10) 0.955 (12) 0.978 (11) 0.971 (5) 0.882 (6) 0.567 (14) 0.527 (5) 0.648 (16) 0.925 (9) 0.559 (14) 0.901 (6) 0.720 (15) 0.659 (9) 0.938 (10) 0.897 (14) 0.612 (12) 0.587 (17) 0.511 (13) 0.489 (15) 0.894 (8) 0.625 (12) 0.724 (14) 0.657 (15) 0.585 (14) 0.604 (12)
0.788 (15) 0.964 (8) 0.614 (9) 0.840 (14) 0.873 (8) 0.654 (13) 1.000 (5) 0.658 (8) 0.844 (7) 0.567 (14) 0.547 (6) 0.597 (12) 0.937 (8) 0.760 (8) 0.888 (6) 0.773 (12) 0.858 (13) 0.722 (13) 0.884 (5) 0.801 (14) 0.656 (4) 0.795 (10) 0.743 (11) 0.911 (16) 0.990 (4) 0.913 (14) 0.907 (4) 0.705 (8) 0.480 (18) 0.743 (14) 0.978 (6) 0.605 (8) 0.815 (13) 0.664 (16) 0.530 (17) 0.843 (13) 0.796 (16) 0.619 (11) 0.698 (12) 0.536 (12) 0.532 (12) 0.843 (12) 0.673 (10) 0.779 (12) 0.675 (14) 0.737 (5) 0.738 (7)
0.829 (14) 0.591 (17) 0.478 (17) 0.806 (15) 0.404 (18) 0.696 (7) 0.844 (15) 0.608 (10) 0.782 (9) 0.350 (18) 0.523 (11) 0.690 (7) 0.998 (1) 0.624 (18) 0.510 (16) 0.849 (7) 0.910 (9) 0.474 (17) 0.621 (16) 0.836 (12) 0.479 (18) 0.543 (18) 0.578 (17) 0.949 (14) 0.946 (15) 0.516 (17) 0.925 (1) 0.432 (18) 0.517 (6) 0.761 (12) 0.872 (13) 0.408 (16) 0.556 (17) 0.561 (17) 0.573 (14) 0.987 (2) 0.952 (12) 0.483 (15) 0.643 (13) 0.391 (18) 0.451 (17) 0.798 (14) 0.597 (14) 0.868 (7) 0.792 (7) 0.704 (9) 0.465 (18)
0.479 (18) 0.573 (18) 0.642 (5) 0.964 (1) 0.719 (16) 0.605 (15) 0.807 (17) 0.670 (7) 0.900 (2) 0.496 (17) 0.547 (7) 0.700 (6) 0.936 (9) 0.639 (16) 0.769 (13) 0.716 (17) 0.843 (14) 0.432 (18) 0.721 (14) 0.699 (16) 0.605 (10) 0.739 (14) 0.701 (14) 0.969 (11) 0.571 (18) 0.508 (18) 0.559 (17) 0.528 (16) 0.506 (12) 0.748 (13) 0.916 (11) 0.372 (17) 0.883 (9) 0.781 (11) 0.537 (16) 0.915 (12) 0.980 (7) 0.363 (17) 0.615 (16) 0.486 (17) 0.458 (16) 0.921 (6) 0.633 (11) 0.785 (11) 0.559 (18) 0.718 (8) 0.599 (13)
0.985 (4) 0.999 (1) 0.624 (7) 0.910 (9) 0.860 (10) 0.715 (6) 1.000 (1) 0.567 (12) 0.563 (14) 0.713 (10) 0.482 (17) 0.533 (16) 0.840 (10) 0.812 (3) 0.896 (4) 0.823 (8) 0.995 (4) 0.997 (1) 0.749 (10) 0.922 (6) 0.632 (7) 0.750 (12) 0.829 (4) 0.987 (7) 0.980 (9) 0.950 (10) 0.693 (14) 0.781 (3) 0.536 (4) 0.865 (7) 0.981 (5) 0.635 (5) 0.924 (4) 0.879 (5) 0.732 (5) 0.971 (7) 0.999 (1) 0.678 (9) 1.000 (2) 0.624 (2) 0.608 (3) 0.846 (11) 0.716 (6) 0.889 (5) 0.840 (4) 0.727 (7) 0.652 (10)
0.985 (3) 0.999 (2) 0.563 (15) 0.891 (12) 0.872 (9) 0.578 (16) 1.000 (2) 0.563 (13) 0.544 (15) 0.731 (9) 0.534 (8) 0.636 (10) 0.707 (15) 0.757 (10) 0.832 (11) 0.754 (15) 0.756 (16) 0.986 (4) 0.723 (13) 0.842 (11) 0.543 (15) 0.851 (7) 0.809 (7) 0.997 (2) 0.981 (7) 0.959 (7) 0.878 (7) 0.585 (13) 0.488 (17) 0.849 (9) 0.830 (15) 0.578 (12) 0.854 (10) 0.810 (10) 0.590 (13) 0.713 (16) 0.985 (5) 0.703 (6) 1.000 (3) 0.547 (10) 0.552 (10) 0.617 (17) 0.577 (16) 0.687 (15) 0.635 (16) 0.654 (12) 0.575 (14)
0.986 (2) 0.997 (3) 0.603 (10) 0.928 (6) 0.907 (6) 0.686 (10) 1.000 (3) 0.546 (14) 0.543 (16) 0.740 (7) 0.477 (18) 0.538 (13) 0.816 (12) 0.807 (4) 0.876 (7) 0.790 (11) 0.997 (2) 0.994 (2) 0.748 (11) 0.888 (9) 0.617 (8) 0.771 (11) 0.814 (6) 0.989 (5) 0.981 (8) 0.952 (9) 0.652 (16) 0.757 (5) 0.555 (2) 0.866 (6) 0.985 (4) 0.668 (3) 0.923 (5) 0.894 (4) 0.734 (4) 0.960 (9) 0.999 (2) 0.682 (8) 1.000 (1) 0.635 (1) 0.607 (4) 0.764 (15) 0.700 (8) 0.861 (8) 0.804 (6) 0.728 (6) 0.646 (11)
0.532 (16) 0.893 (13) 0.469 (18) 0.654 (17) 0.744 (15) 0.532 (17) 1.000 (6) 0.497 (17) 0.459 (18) 0.662 (13) 0.488 (16) 0.537 (14) 0.629 (16) 0.670 (15) 0.591 (15) 0.605 (18) 0.689 (17) 0.597 (14) 0.542 (18) 0.602 (17) 0.481 (17) 0.844 (8) 0.590 (16) 0.426 (18) 0.919 (16) 0.972 (4) 0.785 (12) 0.511 (17) 0.488 (16) 0.579 (17) 0.694 (17) 0.676 (2) 0.560 (16) 0.734 (14) 0.658 (10) 0.627 (17) 0.812 (15) 0.952 (1) 0.626 (14) 0.574 (8) 0.579 (5) 0.486 (18) 0.578 (15) 0.480 (18) 0.605 (17) 0.572 (16) 0.551 (16)
0.993 (1) 0.990 (6) 0.586 (13) 0.933 (4) 0.930 (3) 0.679 (11) 0.997 (11) 0.607 (11) 0.846 (6) 0.667 (12) 0.526 (10) 0.609 (11) 0.991 (4) 0.807 (5) 0.890 (5) 0.913 (3) 0.988 (6) 0.966 (8) 0.823 (9) 0.940 (3) 0.579 (13) 0.798 (9) 0.821 (5) 0.989 (4) 0.989 (5) 0.938 (11) 0.790 (11) 0.662 (11) 0.509 (10) 0.927 (1) 0.991 (1) 0.609 (7) 0.892 (7) 0.863 (7) 0.653 (12) 0.980 (6) 0.938 (13) 0.579 (13) 0.987 (5) 0.580 (6) 0.576 (7) 0.903 (7) 0.611 (13) 0.956 (1) 0.844 (3) 0.683 (11) 0.715 (8)
0.865 (11) 0.890 (14) 0.669 (3) 0.920 (8) 0.911 (5) 0.689 (8) 0.838 (16) 0.525 (15) 0.894 (3) 0.808 (2) 0.499 (13) 0.878 (2) 0.539 (17) 0.628 (17) 0.319 (18) 0.907 (5) 0.996 (3) 0.985 (5) 0.712 (15) 0.953 (2) 0.615 (9) 0.574 (17) 0.729 (13) 0.977 (9) 0.976 (13) 0.804 (15) 0.842 (8) 0.729 (7) 0.604 (1) 0.689 (15) 0.840 (14) 0.597 (9) 0.929 (3) 0.917 (3) 0.777 (2) 0.758 (15) 0.767 (17) 0.459 (16) 0.615 (15) 0.581 (5) 0.499 (14) 0.827 (13) 0.692 (9) 0.828 (10) 0.731 (11) 0.769 (3) 0.764 (6)
0.922 (10) 0.691 (16) 0.705 (1) 0.947 (2) 0.745 (14) 0.635 (14) 0.995 (12) 0.671 (6) 0.934 (1) 0.514 (15) 0.562 (5) 0.790 (4) 0.793 (14) 0.733 (11) 0.852 (10) 0.806 (10) 0.869 (11) 0.569 (15) 0.902 (4) 0.891 (8) 0.728 (1) 0.592 (16) 0.609 (15) 0.976 (10) 0.980 (10) 0.732 (16) 0.889 (5) 0.554 (15) 0.509 (9) 0.874 (4) 0.916 (10) 0.364 (18) 0.884 (8) 0.820 (9) 0.721 (7) 0.969 (8) 0.956 (11) 0.325 (18) 0.732 (11) 0.511 (14) 0.390 (18) 0.978 (1) 0.706 (7) 0.911 (3) 0.691 (13) 0.627 (13) 0.686 (9)
0.500 (17) 0.938 (12) 0.542 (16) 0.500 (18) 0.500 (17) 0.500 (18) 0.500 (18) 0.500 (16) 0.706 (12) 0.699 (11) 0.509 (12) 0.500 (18) 0.500 (18) 0.688 (13) 0.500 (17) 0.814 (9) 0.500 (18) 0.500 (16) 0.835 (8) 0.500 (18) 0.587 (11) 0.928 (4) 0.731 (12) 0.500 (17) 0.971 (14) 0.916 (13) 0.500 (18) 0.699 (9) 0.500 (13) 0.500 (18) 0.500 (18) 0.500 (15) 0.500 (18) 0.500 (18) 0.500 (18) 0.500 (18) 0.500 (18) 0.500 (14) 0.500 (18) 0.500 (16) 0.510 (13) 0.862 (10) 0.790 (4) 0.500 (17) 0.786 (8) 0.500 (17) 0.819 (2)
0.969 (6) 0.976 (7) 0.618 (8) 0.943 (3) 0.918 (4) 0.841 (1) 0.999 (9) 0.782 (1) 0.851 (5) 0.807 (3) 0.526 (9) 0.648 (9) 0.997 (2) 0.728 (12) 0.904 (2) 0.913 (4) 0.992 (5) 0.971 (7) 0.942 (3) 0.929 (5) 0.716 (2) 0.955 (2) 0.805 (9) 0.982 (8) 0.991 (3) 0.990 (1) 0.802 (10) 0.774 (4) 0.514 (7) 0.879 (3) 0.988 (2) 0.562 (13) 0.823 (12) 0.830 (8) 0.697 (8) 0.986 (3) 0.989 (4) 0.682 (7) 0.908 (9) 0.507 (15) 0.542 (11) 0.976 (2) 0.839 (2) 0.939 (2) 0.847 (2) 0.820 (2) 0.829 (1)
0.930 (9) 0.953 (10) 0.602 (11) 0.848 (13) 0.905 (7) 0.783 (4) 1.000 (4) 0.447 (18) 0.505 (17) 0.778 (6) 0.490 (15) 0.535 (15) 0.839 (11) 0.790 (6) 0.731 (14) 0.765 (13) 0.941 (7) 0.930 (10) 0.590 (17) 0.815 (13) 0.520 (16) 0.698 (15) 0.843 (3) 0.950 (13) 0.977 (12) 0.989 (2) 0.681 (15) 0.679 (10) 0.510 (8) 0.852 (8) 0.767 (16) 0.696 (1) 0.790 (14) 0.744 (13) 0.758 (3) 0.812 (14) 0.972 (10) 0.737 (5) 0.956 (7) 0.579 (7) 0.610 (2) 0.669 (16) 0.539 (18) 0.570 (16) 0.737 (10) 0.693 (10) 0.564 (15)
0.853 (12) 0.885 (15) 0.566 (14) 0.740 (16) 0.823 (11) 0.824 (3) 0.994 (13) 0.673 (4) 0.685 (13) 0.509 (16) 0.499 (14) 0.502 (17) 0.973 (7) 0.684 (14) 0.869 (8) 0.718 (16) 0.882 (10) 0.773 (12) 0.730 (12) 0.790 (15) 0.566 (14) 0.852 (6) 0.532 (18) 0.944 (15) 0.880 (17) 0.953 (8) 0.708 (13) 0.629 (12) 0.553 (3) 0.817 (10) 0.902 (12) 0.613 (6) 0.687 (15) 0.777 (12) 0.567 (15) 0.993 (1) 0.982 (6) 0.777 (4) 0.992 (4) 0.609 (3) 0.619 (1) 0.953 (3) 0.572 (17) 0.729 (13) 0.709 (12) 0.487 (18) 0.512 (17)
0.982 (5) 0.995 (4) 0.651 (4) 0.904 (10) 0.950 (1) 0.776 (5) 0.999 (10) 0.673 (5) 0.804 (8) 0.806 (4) 0.609 (2) 0.876 (3) 0.976 (6) 0.853 (2) 0.900 (3) 0.930 (2) 0.859 (12) 0.976 (6) 0.954 (2) 0.930 (4) 0.685 (3) 0.967 (1) 0.850 (2) 0.993 (3) 0.998 (1) 0.966 (6) 0.922 (2) 0.853 (2) 0.506 (11) 0.886 (2) 0.968 (7) 0.644 (4) 0.934 (2) 0.971 (1) 0.731 (6) 0.985 (4) 0.994 (3) 0.949 (2) 0.967 (6) 0.581 (4) 0.566 (9) 0.948 (4) 0.866 (1) 0.876 (6) 0.833 (5) 0.582 (15) 0.765 (5)
0.951 (7) 0.991 (5) 0.641 (6) 0.930 (5) 0.934 (2) 0.838 (2) 1.000 (7) 0.756 (2) 0.868 (4) 0.814 (1) 0.716 (1) 0.881 (1) 0.988 (5) 0.862 (1) 0.919 (1) 0.980 (1) 0.999 (1) 0.987 (3) 0.960 (1) 0.953 (1) 0.652 (5) 0.941 (3) 0.858 (1) 0.997 (1) 0.996 (2) 0.973 (3) 0.920 (3) 0.873 (1) 0.491 (15) 0.867 (5) 0.986 (3) 0.589 (10) 0.962 (1) 0.959 (2) 0.833 (1) 0.982 (5) 0.976 (9) 0.942 (3) 0.943 (8) 0.546 (11) 0.566 (8) 0.927 (5) 0.805 (3) 0.909 (4) 0.856 (1) 0.863 (1) 0.816 (3)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.841 (7) 0.858 (8) 0.852 (9) 0.650 (12) 0.547 (11)
0.774 (11) 0.785 (14) 0.756 (12) 0.594 (15) 0.535 (14)
0.715 (15) 0.745 (17) 0.745 (13) 0.702 (8) 0.536 (13)
0.888 (4) 0.883 (6) 0.718 (15) 0.733 (5) 0.525 (17)
0.744 (14) 0.821 (12) 0.652 (18) 0.841 (1) 0.511 (18)
0.761 (12) 0.806 (13) 0.891 (3) 0.625 (14) 0.579 (1)
0.688 (16) 0.751 (15) 0.888 (5) 0.583 (16) 0.559 (6)
0.781 (10) 0.836 (11) 0.885 (6) 0.627 (13) 0.568 (2)
0.544 (17) 0.633 (18) 0.784 (11) 0.582 (17) 0.559 (7)
0.881 (5) 0.893 (5) 0.890 (4) 0.773 (4) 0.552 (8)
0.836 (8) 0.862 (7) 0.706 (17) 0.729 (6) 0.549 (9)
0.880 (6) 0.893 (4) 0.716 (16) 0.659 (11) 0.567 (3)
0.500 (18) 0.840 (10) 0.852 (8) 0.547 (18) 0.532 (15)
0.897 (3) 0.902 (2) 0.849 (10) 0.836 (2) 0.545 (12)
0.818 (9) 0.846 (9) 0.891 (2) 0.664 (10) 0.561 (4)
0.756 (13) 0.748 (16) 0.738 (14) 0.700 (9) 0.530 (16)
0.903 (2) 0.900 (3) 0.879 (7) 0.783 (3) 0.561 (5)
0.926 (1) 0.929 (1) 0.897 (1) 0.713 (7) 0.548 (10)
20news agnews amazon imdb yelp
0.662 (11) 0.762 (9) 0.633 (11) 0.657 (7) 0.757 (9)
0.594 (17) 0.677 (15) 0.542 (14) 0.539 (13) 0.599 (13)
0.622 (13) 0.669 (16) 0.609 (12) 0.580 (12) 0.634 (12)
0.499 (18) 0.662 (17) 0.512 (16) 0.465 (17) 0.555 (16)
0.698 (8) 0.750 (11) 0.501 (17) 0.451 (18) 0.500 (17)
0.666 (10) 0.842 (4) 0.718 (2) 0.738 (4) 0.869 (2)
0.657 (12) 0.732 (12) 0.654 (7) 0.643 (10) 0.791 (6)
0.700 (5) 0.803 (7) 0.715 (3) 0.752 (3) 0.851 (3)
0.618 (15) 0.681 (14) 0.706 (5) 0.726 (5) 0.763 (7)
0.734 (2) 0.837 (5) 0.712 (4) 0.706 (6) 0.834 (4)
0.615 (16) 0.701 (13) 0.471 (18) 0.511 (15) 0.481 (18)
0.727 (3) 0.849 (3) 0.530 (15) 0.525 (14) 0.565 (15)
0.621 (14) 0.753 (10) 0.634 (10) 0.655 (8) 0.760 (8)
0.699 (7) 0.795 (8) 0.641 (8) 0.604 (11) 0.731 (11)
0.700 (6) 0.868 (2) 0.745 (1) 0.777 (1) 0.877 (1)
0.683 (9) 0.588 (18) 0.580 (13) 0.477 (16) 0.570 (14)
0.710 (4) 0.812 (6) 0.639 (9) 0.650 (9) 0.742 (10)
0.750 (1) 0.881 (1) 0.706 (6) 0.752 (2) 0.829 (5)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 42: The average AUCROC performance on individual datasets under the 𝛾𝑙𝑎 = 5% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.965 (10) 0.996 (8) 0.706 (4) 0.924 (7) 0.957 (5) 0.836 (9) 0.974 (15) 0.766 (5) 0.885 (8) 0.872 (4) 0.660 (2) 0.864 (3) 0.754 (15) 0.847 (5) 0.885 (11) 0.961 (4) 0.999 (6) 0.979 (10) 0.951 (5) 0.977 (11) 0.719 (8) 0.966 (4) 0.891 (4) 0.982 (6) 0.997 (3) 0.993 (8) 0.834 (9) 0.891 (6) 0.557 (9) 0.901 (8) 0.970 (10) 0.692 (13) 0.929 (8) 0.857 (7) 0.789 (10) 0.726 (16) 0.970 (10) 0.836 (8) 0.929 (12) 0.661 (10) 0.577 (9) 0.973 (6) 0.880 (4) 0.955 (7) 0.849 (7) 0.864 (6) 0.876 (3)
0.934 (12) 0.999 (6) 0.702 (5) 0.895 (11) 0.877 (13) 0.763 (13) 0.999 (11) 0.721 (8) 0.833 (11) 0.862 (5) 0.566 (5) 0.737 (6) 0.850 (12) 0.813 (11) 0.918 (7) 0.867 (12) 0.952 (13) 0.949 (12) 0.929 (7) 0.970 (12) 0.640 (14) 0.812 (14) 0.882 (5) 0.982 (5) 0.991 (5) 0.995 (7) 0.882 (6) 0.695 (12) 0.537 (11) 0.742 (16) 0.958 (11) 0.633 (15) 0.946 (6) 0.731 (14) 0.720 (14) 0.947 (11) 0.896 (14) 0.801 (9) 0.755 (15) 0.575 (15) 0.539 (13) 0.922 (14) 0.672 (15) 0.821 (15) 0.798 (11) 0.758 (14) 0.690 (12)
0.956 (11) 0.999 (5) 0.677 (9) 0.840 (14) 0.913 (9) 0.848 (7) 1.000 (8) 0.768 (4) 0.948 (1) 0.640 (12) 0.552 (9) 0.701 (8) 0.861 (10) 0.813 (12) 0.929 (6) 0.887 (11) 0.953 (12) 0.771 (13) 0.958 (4) 0.883 (14) 0.743 (6) 0.920 (9) 0.801 (12) 0.958 (15) 0.990 (6) 0.912 (14) 0.907 (4) 0.816 (10) 0.523 (12) 0.907 (7) 0.974 (9) 0.773 (6) 0.913 (11) 0.789 (11) 0.621 (16) 0.808 (13) 0.813 (15) 0.731 (11) 0.945 (11) 0.594 (14) 0.529 (14) 0.965 (10) 0.726 (9) 0.886 (12) 0.836 (9) 0.830 (12) 0.800 (6)
0.917 (14) 0.388 (18) 0.576 (17) 0.806 (15) 0.484 (18) 0.724 (14) 0.988 (14) 0.572 (16) 0.736 (12) 0.363 (18) 0.524 (12) 0.417 (18) 0.974 (5) 0.627 (18) 0.386 (18) 0.931 (8) 0.691 (18) 0.529 (17) 0.592 (18) 0.820 (16) 0.536 (18) 0.602 (16) 0.577 (18) 0.941 (17) 0.941 (16) 0.428 (18) 0.925 (1) 0.615 (15) 0.466 (18) 0.617 (18) 0.848 (17) 0.571 (16) 0.489 (18) 0.361 (18) 0.571 (17) 0.981 (6) 0.924 (13) 0.505 (16) 0.355 (18) 0.401 (18) 0.453 (16) 0.960 (11) 0.747 (8) 0.871 (14) 0.769 (14) 0.775 (13) 0.480 (18)
0.489 (18) 0.608 (16) 0.644 (14) 0.964 (1) 0.721 (16) 0.614 (18) 0.843 (17) 0.671 (11) 0.901 (7) 0.505 (17) 0.548 (11) 0.701 (7) 0.939 (7) 0.645 (17) 0.773 (16) 0.714 (17) 0.818 (16) 0.435 (18) 0.790 (16) 0.641 (18) 0.608 (16) 0.779 (15) 0.704 (15) 0.970 (11) 0.640 (17) 0.509 (17) 0.559 (17) 0.530 (18) 0.506 (16) 0.752 (15) 0.918 (14) 0.373 (17) 0.885 (15) 0.783 (12) 0.534 (18) 0.911 (12) 0.980 (7) 0.463 (17) 0.621 (17) 0.486 (17) 0.446 (17) 0.923 (13) 0.665 (16) 0.789 (17) 0.566 (18) 0.639 (17) 0.605 (16)
0.996 (7) 1.000 (2) 0.689 (7) 0.910 (9) 0.884 (11) 0.893 (3) 1.000 (1) 0.623 (13) 0.718 (13) 0.754 (10) 0.486 (16) 0.611 (13) 0.851 (11) 0.833 (8) 0.929 (5) 0.952 (5) 1.000 (1) 0.999 (1) 0.883 (13) 0.996 (4) 0.759 (2) 0.826 (13) 0.847 (8) 0.976 (10) 0.980 (12) 0.951 (12) 0.693 (14) 0.905 (5) 0.601 (3) 0.926 (6) 0.995 (3) 0.744 (10) 0.954 (4) 0.878 (5) 0.866 (6) 0.986 (2) 0.999 (1) 0.693 (13) 1.000 (1) 0.725 (2) 0.632 (1) 0.973 (7) 0.848 (6) 0.976 (4) 0.911 (3) 0.906 (4) 0.797 (7)
0.997 (4) 1.000 (1) 0.634 (15) 0.891 (12) 0.878 (12) 0.787 (11) 1.000 (2) 0.574 (15) 0.701 (14) 0.805 (7) 0.555 (8) 0.695 (9) 0.908 (8) 0.850 (3) 0.933 (3) 0.925 (9) 1.000 (2) 0.995 (5) 0.914 (8) 0.993 (6) 0.639 (15) 0.932 (8) 0.864 (6) 0.979 (7) 0.984 (9) 0.982 (10) 0.878 (7) 0.834 (7) 0.498 (17) 0.821 (13) 0.997 (1) 0.762 (9) 0.958 (3) 0.798 (10) 0.771 (11) 0.753 (15) 0.985 (5) 0.769 (10) 1.000 (2) 0.624 (13) 0.586 (8) 0.772 (17) 0.707 (11) 0.920 (8) 0.769 (13) 0.838 (11) 0.676 (14)
0.996 (6) 0.999 (4) 0.678 (8) 0.928 (6) 0.954 (6) 0.851 (6) 1.000 (3) 0.605 (14) 0.700 (15) 0.773 (9) 0.500 (14) 0.607 (14) 0.832 (13) 0.836 (7) 0.930 (4) 0.944 (6) 1.000 (4) 0.999 (2) 0.889 (12) 0.997 (1) 0.759 (3) 0.829 (12) 0.843 (9) 0.966 (12) 0.980 (13) 0.954 (11) 0.652 (16) 0.907 (4) 0.611 (2) 0.944 (4) 0.995 (4) 0.771 (7) 0.946 (5) 0.908 (4) 0.863 (7) 0.984 (3) 0.999 (2) 0.694 (12) 1.000 (3) 0.733 (1) 0.632 (2) 0.902 (15) 0.828 (7) 0.972 (6) 0.899 (4) 0.899 (5) 0.747 (9)
0.781 (17) 0.957 (14) 0.551 (18) 0.654 (17) 0.820 (14) 0.659 (17) 1.000 (9) 0.547 (18) 0.567 (18) 0.609 (13) 0.498 (15) 0.575 (16) 0.537 (17) 0.805 (13) 0.868 (13) 0.630 (18) 0.762 (17) 0.672 (15) 0.842 (15) 0.669 (17) 0.586 (17) 0.939 (6) 0.730 (13) 0.760 (18) 0.987 (8) 0.996 (5) 0.785 (12) 0.641 (14) 0.513 (14) 0.675 (17) 0.738 (18) 0.723 (11) 0.763 (16) 0.780 (13) 0.724 (13) 0.606 (17) 0.803 (16) 0.982 (3) 0.806 (14) 0.660 (11) 0.558 (10) 0.631 (18) 0.598 (18) 0.617 (18) 0.630 (17) 0.687 (16) 0.564 (17)
0.998 (3) 0.994 (10) 0.668 (12) 0.933 (4) 0.936 (8) 0.697 (15) 0.997 (12) 0.754 (6) 0.885 (9) 0.691 (11) 0.548 (10) 0.638 (11) 0.995 (2) 0.819 (9) 0.903 (9) 0.937 (7) 0.999 (8) 0.991 (6) 0.889 (11) 0.993 (5) 0.720 (7) 0.844 (10) 0.820 (11) 0.996 (2) 0.989 (7) 0.940 (13) 0.790 (11) 0.667 (13) 0.573 (7) 0.961 (1) 0.996 (2) 0.719 (12) 0.923 (10) 0.834 (8) 0.877 (4) 0.978 (7) 0.945 (12) 0.654 (14) 0.996 (6) 0.680 (7) 0.555 (11) 0.972 (8) 0.674 (13) 0.980 (1) 0.895 (6) 0.841 (10) 0.754 (8)
0.903 (15) 0.492 (17) 0.655 (13) 0.920 (8) 0.911 (10) 0.779 (12) 0.887 (16) 0.655 (12) 0.916 (5) 0.589 (15) 0.484 (17) 0.851 (4) 0.561 (16) 0.661 (16) 0.393 (17) 0.852 (13) 0.971 (11) 0.981 (8) 0.684 (17) 0.981 (10) 0.682 (12) 0.474 (18) 0.723 (14) 0.988 (4) 0.271 (18) 0.836 (15) 0.842 (8) 0.605 (16) 0.663 (1) 0.799 (14) 0.876 (16) 0.689 (14) 0.909 (12) 0.910 (3) 0.712 (15) 0.777 (14) 0.786 (17) 0.561 (15) 0.835 (13) 0.640 (12) 0.514 (15) 0.899 (16) 0.674 (14) 0.892 (11) 0.745 (15) 0.850 (8) 0.627 (15)
0.923 (13) 0.696 (15) 0.699 (6) 0.947 (2) 0.748 (15) 0.666 (16) 0.995 (13) 0.673 (10) 0.937 (2) 0.511 (16) 0.557 (7) 0.793 (5) 0.807 (14) 0.732 (15) 0.854 (15) 0.811 (16) 0.883 (14) 0.573 (16) 0.901 (10) 0.882 (15) 0.715 (9) 0.594 (17) 0.624 (17) 0.977 (9) 0.980 (11) 0.720 (16) 0.889 (5) 0.560 (17) 0.508 (15) 0.872 (10) 0.924 (13) 0.359 (18) 0.890 (14) 0.825 (9) 0.725 (12) 0.970 (9) 0.960 (11) 0.326 (18) 0.749 (16) 0.497 (16) 0.393 (18) 0.979 (5) 0.708 (10) 0.908 (10) 0.708 (16) 0.634 (18) 0.684 (13)
0.995 (8) 0.977 (13) 0.675 (11) 0.500 (18) 0.500 (17) 0.859 (5) 0.500 (18) 0.725 (7) 0.866 (10) 0.806 (6) 0.576 (4) 0.636 (12) 0.500 (18) 0.768 (14) 0.860 (14) 0.924 (10) 0.998 (9) 0.966 (11) 0.940 (6) 0.985 (9) 0.708 (10) 0.933 (7) 0.839 (10) 0.947 (16) 0.977 (14) 0.985 (9) 0.500 (18) 0.821 (9) 0.591 (4) 0.870 (11) 0.992 (7) 0.786 (4) 0.929 (9) 0.500 (17) 0.805 (9) 0.500 (18) 0.500 (18) 0.858 (6) 0.986 (9) 0.680 (6) 0.554 (12) 0.934 (12) 0.853 (5) 0.920 (9) 0.820 (10) 0.844 (9) 0.868 (5)
0.980 (9) 0.994 (9) 0.720 (3) 0.943 (3) 0.973 (2) 0.912 (1) 1.000 (7) 0.818 (3) 0.927 (4) 0.886 (3) 0.565 (6) 0.671 (10) 0.997 (1) 0.816 (10) 0.911 (8) 0.969 (3) 0.999 (7) 0.991 (7) 0.963 (3) 0.991 (8) 0.748 (5) 0.982 (3) 0.892 (3) 0.978 (8) 0.991 (4) 0.997 (2) 0.802 (10) 0.924 (3) 0.543 (10) 0.948 (2) 0.992 (8) 0.773 (5) 0.940 (7) 0.862 (6) 0.870 (5) 0.986 (1) 0.989 (4) 0.920 (5) 0.989 (7) 0.673 (8) 0.619 (3) 0.989 (1) 0.888 (3) 0.973 (5) 0.896 (5) 0.856 (7) 0.876 (4)
0.996 (5) 0.993 (11) 0.677 (10) 0.848 (13) 0.963 (4) 0.843 (8) 1.000 (4) 0.560 (17) 0.632 (17) 0.801 (8) 0.484 (18) 0.603 (15) 0.899 (9) 0.850 (4) 0.890 (10) 0.821 (14) 0.997 (10) 0.980 (9) 0.869 (14) 0.993 (7) 0.694 (11) 0.837 (11) 0.855 (7) 0.961 (13) 0.980 (10) 0.996 (4) 0.681 (15) 0.825 (8) 0.581 (6) 0.876 (9) 0.913 (15) 0.831 (2) 0.906 (13) 0.661 (15) 0.893 (2) 0.982 (5) 0.972 (9) 0.836 (7) 0.996 (5) 0.689 (5) 0.610 (7) 0.989 (2) 0.690 (12) 0.806 (16) 0.848 (8) 0.925 (1) 0.716 (10)
0.881 (16) 0.983 (12) 0.618 (16) 0.740 (16) 0.944 (7) 0.794 (10) 1.000 (6) 0.709 (9) 0.665 (16) 0.595 (14) 0.503 (13) 0.569 (17) 0.964 (6) 0.839 (6) 0.877 (12) 0.816 (15) 0.881 (15) 0.726 (14) 0.905 (9) 0.899 (13) 0.671 (13) 0.943 (5) 0.641 (16) 0.960 (14) 0.952 (15) 0.995 (6) 0.708 (13) 0.753 (11) 0.522 (13) 0.839 (12) 0.931 (12) 0.825 (3) 0.753 (17) 0.607 (16) 0.830 (8) 0.966 (10) 0.982 (6) 0.972 (4) 0.968 (10) 0.689 (4) 0.614 (5) 0.966 (9) 0.602 (17) 0.874 (13) 0.783 (12) 0.724 (15) 0.691 (11)
0.999 (2) 0.999 (3) 0.733 (2) 0.904 (10) 0.976 (1) 0.898 (2) 0.999 (10) 0.838 (2) 0.905 (6) 0.902 (2) 0.632 (3) 0.895 (2) 0.979 (4) 0.892 (2) 0.934 (2) 0.979 (2) 1.000 (5) 0.997 (4) 0.977 (1) 0.997 (3) 0.759 (4) 0.988 (1) 0.929 (1) 0.996 (3) 0.999 (1) 0.997 (3) 0.922 (2) 0.956 (2) 0.582 (5) 0.947 (3) 0.994 (5) 0.832 (1) 0.974 (2) 0.973 (1) 0.886 (3) 0.983 (4) 0.994 (3) 0.985 (1) 0.998 (4) 0.719 (3) 0.615 (4) 0.988 (3) 0.921 (1) 0.980 (2) 0.927 (1) 0.908 (3) 0.881 (2)
0.999 (1) 0.998 (7) 0.734 (1) 0.930 (5) 0.971 (3) 0.887 (4) 1.000 (5) 0.889 (1) 0.927 (3) 0.905 (1) 0.732 (1) 0.933 (1) 0.987 (3) 0.898 (1) 0.942 (1) 0.992 (1) 1.000 (3) 0.998 (3) 0.977 (2) 0.997 (2) 0.760 (1) 0.985 (2) 0.926 (2) 0.998 (1) 0.999 (2) 0.999 (1) 0.920 (3) 0.962 (1) 0.558 (8) 0.942 (5) 0.994 (6) 0.770 (8) 0.989 (1) 0.963 (2) 0.916 (1) 0.977 (8) 0.977 (8) 0.983 (2) 0.988 (8) 0.667 (9) 0.612 (6) 0.985 (4) 0.913 (2) 0.978 (3) 0.916 (2) 0.920 (2) 0.882 (1)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.924 (6) 0.920 (7) 0.913 (8) 0.835 (5) 0.592 (10)
0.833 (15) 0.835 (15) 0.848 (13) 0.661 (18) 0.558 (15)
0.833 (16) 0.850 (14) 0.885 (12) 0.789 (9) 0.589 (11)
0.877 (11) 0.891 (11) 0.715 (16) 0.670 (16) 0.523 (17)
0.745 (18) 0.830 (16) 0.658 (18) 0.841 (3) 0.511 (18)
0.921 (7) 0.919 (8) 0.935 (2) 0.791 (8) 0.639 (1)
0.851 (12) 0.890 (12) 0.936 (1) 0.703 (13) 0.618 (5)
0.908 (8) 0.921 (6) 0.931 (4) 0.780 (10) 0.630 (3)
0.766 (17) 0.813 (17) 0.908 (9) 0.681 (15) 0.613 (7)
0.941 (3) 0.938 (3) 0.929 (6) 0.839 (4) 0.626 (4)
0.836 (14) 0.856 (13) 0.709 (17) 0.682 (14) 0.546 (16)
0.882 (10) 0.896 (9) 0.719 (15) 0.670 (17) 0.567 (13)
0.902 (9) 0.895 (10) 0.902 (10) 0.829 (6) 0.598 (9)
0.928 (5) 0.927 (5) 0.890 (11) 0.877 (1) 0.565 (14)
0.941 (4) 0.938 (4) 0.922 (7) 0.776 (11) 0.611 (8)
0.840 (13) 0.809 (18) 0.829 (14) 0.764 (12) 0.575 (12)
0.951 (2) 0.942 (2) 0.931 (5) 0.871 (2) 0.638 (2)
0.957 (1) 0.953 (1) 0.934 (3) 0.801 (7) 0.616 (6)
20news agnews amazon imdb yelp
0.734 (8) 0.893 (8) 0.758 (10) 0.762 (10) 0.844 (10)
0.619 (17) 0.773 (14) 0.602 (14) 0.622 (14) 0.690 (14)
0.674 (14) 0.827 (12) 0.668 (11) 0.691 (12) 0.781 (11)
0.532 (18) 0.658 (18) 0.498 (17) 0.456 (17) 0.570 (15)
0.697 (12) 0.754 (15) 0.509 (16) 0.448 (18) 0.522 (17)
0.773 (6) 0.938 (4) 0.852 (1) 0.873 (1) 0.940 (1)
0.721 (11) 0.895 (7) 0.773 (8) 0.788 (8) 0.874 (8)
0.783 (4) 0.938 (3) 0.850 (2) 0.865 (2) 0.933 (2)
0.677 (13) 0.870 (10) 0.789 (7) 0.795 (7) 0.892 (7)
0.775 (5) 0.938 (5) 0.826 (4) 0.856 (4) 0.926 (3)
0.622 (16) 0.716 (16) 0.474 (18) 0.511 (16) 0.472 (18)
0.725 (9) 0.850 (11) 0.531 (15) 0.522 (15) 0.569 (16)
0.722 (10) 0.886 (9) 0.773 (9) 0.771 (9) 0.848 (9)
0.744 (7) 0.821 (13) 0.642 (13) 0.726 (11) 0.772 (12)
0.788 (3) 0.942 (2) 0.815 (5) 0.853 (5) 0.901 (6)
0.639 (15) 0.691 (17) 0.661 (12) 0.689 (13) 0.699 (13)
0.794 (2) 0.925 (6) 0.814 (6) 0.831 (6) 0.904 (5)
0.798 (1) 0.944 (1) 0.838 (3) 0.857 (3) 0.915 (4)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 43: The average AUCROC performance on individual datasets under the 𝛾𝑙𝑎 = 10% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.990 (9) 0.999 (8) 0.738 (4) 0.903 (12) 0.974 (4) 0.937 (6) 0.974 (15) 0.770 (5) 0.939 (5) 0.905 (4) 0.728 (2) 0.897 (3) 0.883 (14) 0.870 (4) 0.894 (13) 0.982 (3) 1.000 (8) 0.991 (7) 0.964 (4) 0.987 (10) 0.736 (9) 0.985 (4) 0.917 (4) 0.978 (9) 0.999 (3) 0.996 (6) 0.834 (9) 0.945 (4) 0.589 (12) 0.964 (3) 0.990 (9) 0.853 (5) 0.968 (5) 0.944 (8) 0.860 (8) 0.864 (16) 0.951 (12) 0.903 (7) 0.923 (12) 0.763 (10) 0.613 (9) 0.978 (11) 0.898 (3) 0.971 (7) 0.901 (8) 0.924 (6) 0.886 (3)
0.964 (12) 1.000 (3) 0.729 (5) 0.906 (10) 0.899 (13) 0.854 (14) 0.999 (12) 0.750 (8) 0.871 (11) 0.900 (5) 0.570 (6) 0.749 (8) 0.870 (15) 0.846 (7) 0.921 (7) 0.915 (11) 0.978 (11) 0.991 (9) 0.943 (8) 0.993 (9) 0.669 (15) 0.885 (11) 0.909 (5) 0.984 (5) 0.993 (5) 0.997 (4) 0.882 (6) 0.802 (12) 0.547 (14) 0.836 (15) 0.974 (12) 0.696 (15) 0.962 (7) 0.788 (14) 0.762 (12) 0.956 (11) 0.927 (13) 0.900 (8) 0.848 (14) 0.651 (15) 0.589 (12) 0.947 (14) 0.694 (14) 0.884 (15) 0.865 (11) 0.812 (13) 0.730 (11)
0.978 (11) 1.000 (6) 0.723 (6) 0.899 (13) 0.928 (8) 0.937 (7) 1.000 (10) 0.783 (4) 0.954 (4) 0.652 (13) 0.566 (7) 0.797 (5) 0.952 (10) 0.834 (10) 0.941 (3) 0.927 (9) 0.966 (12) 0.876 (14) 0.963 (5) 0.981 (12) 0.747 (7) 0.961 (8) 0.812 (12) 0.970 (12) 0.990 (6) 0.910 (14) 0.907 (4) 0.866 (10) 0.598 (11) 0.954 (7) 0.989 (10) 0.821 (9) 0.964 (6) 0.889 (11) 0.644 (16) 0.871 (15) 0.914 (14) 0.751 (11) 0.976 (10) 0.739 (11) 0.534 (14) 0.979 (10) 0.721 (11) 0.935 (10) 0.892 (9) 0.854 (11) 0.787 (8)
0.887 (16) 0.456 (18) 0.586 (18) 0.839 (16) 0.596 (18) 0.763 (16) 0.671 (18) 0.718 (9) 0.769 (15) 0.361 (18) 0.534 (13) 0.410 (18) 0.904 (13) 0.669 (17) 0.444 (17) 0.873 (14) 0.749 (18) 0.650 (16) 0.665 (16) 0.798 (16) 0.506 (18) 0.581 (17) 0.584 (18) 0.941 (16) 0.931 (16) 0.399 (18) 0.925 (1) 0.770 (14) 0.400 (18) 0.657 (18) 0.847 (18) 0.486 (16) 0.737 (18) 0.397 (18) 0.568 (17) 0.957 (10) 0.901 (16) 0.486 (16) 0.280 (18) 0.433 (18) 0.418 (17) 0.969 (12) 0.780 (9) 0.859 (16) 0.737 (15) 0.797 (14) 0.433 (18)
0.392 (18) 0.567 (17) 0.646 (16) 0.964 (1) 0.721 (17) 0.615 (18) 0.848 (16) 0.672 (14) 0.905 (8) 0.515 (16) 0.550 (9) 0.702 (12) 0.940 (11) 0.587 (18) 0.734 (16) 0.717 (17) 0.892 (15) 0.434 (18) 0.644 (18) 0.712 (18) 0.608 (17) 0.754 (15) 0.709 (16) 0.970 (13) 0.687 (17) 0.473 (17) 0.559 (17) 0.534 (18) 0.506 (16) 0.754 (17) 0.920 (14) 0.374 (18) 0.886 (15) 0.784 (15) 0.533 (18) 0.916 (13) 0.980 (9) 0.458 (17) 0.634 (17) 0.486 (17) 0.446 (16) 0.928 (16) 0.709 (13) 0.794 (17) 0.569 (18) 0.714 (16) 0.599 (16)
0.999 (4) 1.000 (1) 0.713 (8) 0.913 (8) 0.862 (14) 0.953 (3) 1.000 (1) 0.686 (11) 0.807 (12) 0.761 (10) 0.509 (16) 0.697 (14) 0.977 (8) 0.833 (11) 0.934 (5) 0.973 (5) 1.000 (2) 0.998 (5) 0.880 (15) 0.997 (4) 0.760 (5) 0.840 (13) 0.851 (9) 0.977 (10) 0.980 (13) 0.950 (12) 0.693 (14) 0.919 (7) 0.700 (2) 0.954 (8) 0.996 (5) 0.803 (11) 0.972 (4) 0.963 (4) 0.895 (7) 0.991 (1) 1.000 (1) 0.695 (13) 1.000 (1) 0.771 (9) 0.648 (7) 0.993 (1) 0.864 (6) 0.985 (3) 0.935 (3) 0.934 (5) 0.843 (6)
0.999 (6) 1.000 (2) 0.688 (13) 0.904 (11) 0.905 (12) 0.885 (12) 1.000 (2) 0.612 (16) 0.755 (17) 0.822 (7) 0.536 (12) 0.729 (9) 0.911 (12) 0.851 (5) 0.936 (4) 0.964 (7) 1.000 (4) 0.999 (4) 0.947 (7) 0.993 (8) 0.716 (11) 0.961 (9) 0.866 (7) 0.978 (8) 0.982 (12) 0.982 (10) 0.878 (7) 0.922 (6) 0.566 (13) 0.940 (10) 0.998 (2) 0.803 (12) 0.974 (3) 0.957 (6) 0.853 (10) 0.922 (12) 1.000 (3) 0.802 (10) 1.000 (2) 0.772 (7) 0.673 (2) 0.944 (15) 0.778 (10) 0.958 (8) 0.881 (10) 0.904 (7) 0.733 (9)
0.999 (3) 0.999 (7) 0.710 (10) 0.918 (6) 0.923 (9) 0.937 (8) 1.000 (3) 0.684 (12) 0.789 (13) 0.783 (9) 0.509 (17) 0.720 (10) 0.987 (6) 0.836 (9) 0.931 (6) 0.971 (6) 1.000 (6) 0.999 (3) 0.886 (13) 0.997 (3) 0.767 (3) 0.840 (12) 0.840 (10) 0.979 (7) 0.980 (14) 0.954 (11) 0.652 (16) 0.928 (5) 0.680 (4) 0.959 (5) 0.995 (7) 0.827 (8) 0.958 (9) 0.975 (3) 0.896 (6) 0.990 (2) 1.000 (2) 0.698 (12) 1.000 (3) 0.788 (2) 0.651 (6) 0.979 (8) 0.863 (7) 0.983 (6) 0.934 (4) 0.935 (4) 0.804 (7)
0.841 (17) 0.994 (13) 0.612 (17) 0.696 (17) 0.780 (15) 0.798 (15) 1.000 (11) 0.593 (17) 0.612 (18) 0.580 (15) 0.544 (10) 0.643 (16) 0.738 (17) 0.842 (8) 0.913 (8) 0.664 (18) 0.942 (14) 0.752 (15) 0.914 (11) 0.748 (17) 0.639 (16) 0.965 (7) 0.807 (13) 0.806 (18) 0.988 (8) 0.996 (7) 0.785 (12) 0.790 (13) 0.625 (8) 0.869 (14) 0.920 (15) 0.736 (14) 0.742 (17) 0.775 (16) 0.741 (13) 0.584 (18) 0.903 (15) 0.991 (1) 0.809 (15) 0.717 (14) 0.579 (13) 0.720 (17) 0.631 (18) 0.640 (18) 0.700 (17) 0.676 (17) 0.587 (17)
0.999 (5) 0.995 (12) 0.708 (11) 0.933 (4) 0.909 (11) 0.888 (11) 0.997 (13) 0.755 (7) 0.886 (9) 0.701 (12) 0.530 (15) 0.687 (15) 0.999 (1) 0.813 (13) 0.912 (10) 0.917 (10) 1.000 (3) 0.991 (8) 0.883 (14) 0.994 (7) 0.707 (13) 0.827 (14) 0.820 (11) 0.996 (3) 0.990 (7) 0.942 (13) 0.790 (11) 0.675 (15) 0.625 (9) 0.959 (6) 0.996 (4) 0.814 (10) 0.938 (11) 0.962 (5) 0.897 (5) 0.985 (7) 0.986 (8) 0.651 (14) 0.995 (7) 0.721 (13) 0.602 (11) 0.982 (6) 0.678 (16) 0.986 (2) 0.905 (7) 0.837 (12) 0.730 (10)
0.897 (15) 0.607 (16) 0.655 (15) 0.906 (9) 0.929 (7) 0.855 (13) 0.834 (17) 0.430 (18) 0.884 (10) 0.584 (14) 0.541 (11) 0.773 (7) 0.670 (18) 0.703 (16) 0.280 (18) 0.875 (13) 0.965 (13) 0.985 (11) 0.654 (17) 0.981 (13) 0.714 (12) 0.449 (18) 0.736 (15) 0.989 (4) 0.575 (18) 0.687 (16) 0.842 (8) 0.540 (17) 0.736 (1) 0.831 (16) 0.897 (17) 0.763 (13) 0.903 (13) 0.922 (9) 0.694 (15) 0.797 (17) 0.837 (17) 0.550 (15) 0.897 (13) 0.726 (12) 0.526 (15) 0.642 (18) 0.670 (17) 0.892 (14) 0.800 (14) 0.858 (10) 0.623 (14)
0.921 (14) 0.711 (15) 0.711 (9) 0.947 (2) 0.745 (16) 0.671 (17) 0.995 (14) 0.683 (13) 0.933 (6) 0.497 (17) 0.560 (8) 0.798 (4) 0.840 (16) 0.734 (15) 0.851 (15) 0.814 (16) 0.866 (17) 0.588 (17) 0.901 (12) 0.896 (15) 0.723 (10) 0.595 (16) 0.629 (17) 0.977 (11) 0.984 (10) 0.736 (15) 0.889 (5) 0.557 (16) 0.508 (15) 0.872 (13) 0.920 (16) 0.389 (17) 0.887 (14) 0.830 (13) 0.728 (14) 0.971 (8) 0.966 (11) 0.330 (18) 0.795 (16) 0.508 (16) 0.389 (18) 0.979 (9) 0.711 (12) 0.913 (11) 0.702 (16) 0.626 (18) 0.688 (13)
0.996 (8) 0.991 (14) 0.698 (12) 0.500 (18) 0.917 (10) 0.949 (4) 1.000 (6) 0.768 (6) 0.923 (7) 0.857 (6) 0.591 (4) 0.698 (13) 0.955 (9) 0.799 (14) 0.868 (14) 0.951 (8) 0.998 (9) 0.979 (12) 0.958 (6) 0.984 (11) 0.738 (8) 0.973 (6) 0.884 (6) 0.936 (17) 0.988 (9) 0.993 (8) 0.500 (18) 0.878 (9) 0.627 (7) 0.941 (9) 0.990 (8) 0.838 (7) 0.956 (10) 0.876 (12) 0.856 (9) 0.904 (14) 0.600 (18) 0.915 (6) 0.984 (9) 0.772 (8) 0.604 (10) 0.962 (13) 0.876 (5) 0.940 (9) 0.864 (12) 0.880 (9) 0.856 (5)
0.990 (10) 0.998 (9) 0.741 (3) 0.945 (3) 0.989 (1) 0.957 (1) 1.000 (8) 0.847 (3) 0.954 (3) 0.918 (3) 0.586 (5) 0.776 (6) 0.997 (2) 0.849 (6) 0.912 (9) 0.980 (4) 1.000 (1) 0.996 (6) 0.975 (3) 0.997 (5) 0.765 (4) 0.988 (3) 0.921 (3) 0.984 (6) 0.996 (4) 0.998 (3) 0.802 (10) 0.947 (3) 0.616 (10) 0.966 (2) 0.995 (6) 0.871 (4) 0.960 (8) 0.922 (10) 0.908 (4) 0.990 (3) 0.999 (4) 0.961 (5) 0.997 (6) 0.786 (3) 0.675 (1) 0.986 (5) 0.894 (4) 0.985 (4) 0.922 (5) 0.881 (8) 0.881 (4)
0.998 (7) 0.997 (10) 0.716 (7) 0.846 (15) 0.975 (3) 0.928 (9) 1.000 (4) 0.631 (15) 0.768 (16) 0.805 (8) 0.487 (18) 0.709 (11) 0.977 (7) 0.830 (12) 0.910 (11) 0.876 (12) 0.995 (10) 0.987 (10) 0.923 (10) 0.996 (6) 0.754 (6) 0.905 (10) 0.855 (8) 0.956 (14) 0.982 (11) 0.992 (9) 0.681 (15) 0.882 (8) 0.683 (3) 0.934 (11) 0.957 (13) 0.873 (3) 0.907 (12) 0.951 (7) 0.917 (3) 0.985 (6) 0.995 (7) 0.865 (9) 1.000 (4) 0.774 (6) 0.646 (8) 0.989 (4) 0.833 (8) 0.906 (13) 0.908 (6) 0.947 (1) 0.697 (12)
0.924 (13) 0.997 (11) 0.662 (14) 0.856 (14) 0.959 (6) 0.910 (10) 1.000 (7) 0.700 (10) 0.770 (14) 0.734 (11) 0.533 (14) 0.636 (17) 0.995 (3) 0.881 (3) 0.907 (12) 0.841 (15) 0.890 (16) 0.890 (13) 0.936 (9) 0.948 (14) 0.706 (14) 0.977 (5) 0.800 (14) 0.948 (15) 0.971 (15) 0.997 (5) 0.708 (13) 0.819 (11) 0.495 (17) 0.917 (12) 0.980 (11) 0.847 (6) 0.859 (16) 0.704 (17) 0.789 (11) 0.971 (9) 0.974 (10) 0.979 (4) 0.966 (11) 0.774 (5) 0.671 (3) 0.980 (7) 0.683 (15) 0.909 (12) 0.836 (13) 0.760 (15) 0.622 (15)
0.999 (2) 1.000 (4) 0.765 (2) 0.917 (7) 0.979 (2) 0.956 (2) 1.000 (9) 0.891 (2) 0.957 (2) 0.930 (2) 0.659 (3) 0.946 (2) 0.992 (5) 0.906 (2) 0.949 (1) 0.988 (2) 1.000 (7) 0.999 (1) 0.984 (2) 0.998 (1) 0.788 (2) 0.990 (1) 0.945 (2) 0.997 (2) 1.000 (1) 0.999 (2) 0.922 (2) 0.971 (2) 0.646 (6) 0.969 (1) 0.998 (1) 0.891 (1) 0.987 (2) 0.992 (1) 0.923 (2) 0.989 (5) 0.998 (5) 0.989 (2) 0.999 (5) 0.807 (1) 0.661 (5) 0.993 (2) 0.927 (1) 0.988 (1) 0.948 (1) 0.942 (2) 0.893 (2)
1.000 (1) 1.000 (5) 0.770 (1) 0.926 (5) 0.972 (5) 0.946 (5) 1.000 (5) 0.917 (1) 0.959 (1) 0.932 (1) 0.784 (1) 0.949 (1) 0.995 (4) 0.916 (1) 0.948 (2) 0.995 (1) 1.000 (5) 0.999 (2) 0.986 (1) 0.998 (2) 0.791 (1) 0.989 (2) 0.945 (1) 0.998 (1) 1.000 (2) 0.999 (1) 0.920 (3) 0.971 (1) 0.649 (5) 0.962 (4) 0.996 (3) 0.880 (2) 0.994 (1) 0.989 (2) 0.927 (1) 0.989 (4) 0.997 (6) 0.987 (3) 0.994 (8) 0.781 (4) 0.670 (4) 0.989 (3) 0.923 (2) 0.984 (5) 0.940 (2) 0.936 (3) 0.900 (1)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.948 (6) 0.940 (7) 0.929 (8) 0.887 (4) 0.623 (11)
0.856 (15) 0.871 (15) 0.879 (14) 0.725 (15) 0.575 (14)
0.894 (11) 0.902 (11) 0.924 (10) 0.845 (9) 0.623 (10)
0.880 (14) 0.890 (13) 0.717 (17) 0.697 (16) 0.524 (17)
0.747 (18) 0.833 (17) 0.661 (18) 0.841 (11) 0.509 (18)
0.949 (5) 0.950 (5) 0.943 (3) 0.883 (5) 0.679 (2)
0.902 (10) 0.923 (9) 0.942 (5) 0.789 (13) 0.653 (6)
0.943 (8) 0.947 (6) 0.943 (4) 0.858 (7) 0.669 (3)
0.884 (12) 0.885 (14) 0.936 (7) 0.750 (14) 0.638 (7)
0.959 (4) 0.954 (4) 0.939 (6) 0.897 (3) 0.665 (5)
0.825 (17) 0.850 (16) 0.720 (16) 0.667 (18) 0.543 (16)
0.884 (13) 0.897 (12) 0.722 (15) 0.688 (17) 0.567 (15)
0.924 (9) 0.922 (10) 0.921 (11) 0.875 (6) 0.629 (9)
0.943 (7) 0.933 (8) 0.910 (12) 0.904 (2) 0.580 (13)
0.964 (3) 0.955 (3) 0.928 (9) 0.851 (8) 0.632 (8)
0.841 (16) 0.804 (18) 0.897 (13) 0.809 (12) 0.594 (12)
0.967 (2) 0.957 (2) 0.944 (2) 0.916 (1) 0.684 (1)
0.968 (1) 0.962 (1) 0.945 (1) 0.842 (10) 0.667 (4)
20news agnews amazon imdb yelp
0.789 (7) 0.928 (7) 0.813 (8) 0.814 (10) 0.899 (10)
0.654 (16) 0.819 (14) 0.650 (14) 0.636 (14) 0.739 (14)
0.709 (13) 0.897 (10) 0.732 (11) 0.766 (12) 0.866 (11)
0.480 (18) 0.681 (18) 0.528 (16) 0.466 (18) 0.526 (16)
0.699 (14) 0.738 (16) 0.467 (18) 0.476 (17) 0.519 (17)
0.831 (6) 0.961 (1) 0.887 (1) 0.904 (1) 0.960 (1)
0.768 (10) 0.925 (8) 0.837 (6) 0.831 (8) 0.929 (7)
0.839 (4) 0.958 (3) 0.881 (2) 0.899 (2) 0.959 (2)
0.752 (11) 0.888 (12) 0.779 (10) 0.846 (7) 0.923 (8)
0.848 (1) 0.956 (4) 0.866 (4) 0.883 (4) 0.952 (4)
0.615 (17) 0.724 (17) 0.468 (17) 0.534 (15) 0.463 (18)
0.730 (12) 0.851 (13) 0.530 (15) 0.523 (16) 0.567 (15)
0.784 (8) 0.921 (9) 0.811 (9) 0.814 (9) 0.903 (9)
0.784 (9) 0.892 (11) 0.732 (12) 0.801 (11) 0.863 (12)
0.840 (2) 0.953 (6) 0.817 (7) 0.866 (6) 0.933 (6)
0.673 (15) 0.798 (15) 0.708 (13) 0.726 (13) 0.776 (13)
0.840 (3) 0.955 (5) 0.862 (5) 0.880 (5) 0.948 (5)
0.835 (5) 0.960 (2) 0.871 (3) 0.894 (3) 0.954 (3)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 44: The average AUCROC performance on individual datasets under the 𝛾𝑙𝑎 = 25% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.992 (12) 0.999 (7) 0.765 (4) 0.943 (5) 0.990 (4) 0.991 (6) 0.983 (15) 0.869 (5) 0.983 (3) 0.927 (4) 0.757 (2) 0.918 (3) 0.930 (14) 0.893 (4) 0.930 (10) 0.990 (4) 1.000 (8) 0.997 (8) 0.976 (4) 0.998 (7) 0.810 (5) 0.988 (5) 0.941 (4) 0.997 (2) 0.999 (3) 0.998 (7) 0.834 (8) 0.965 (4) 0.673 (6) 0.978 (8) 0.997 (7) 0.934 (5) 0.976 (7) 0.961 (9) 0.902 (8) 0.959 (13) 0.990 (10) 0.955 (8) 0.996 (11) 0.885 (3) 0.677 (6) 0.990 (8) 0.919 (3) 0.988 (4) 0.947 (7) 0.953 (4) 0.916 (3)
0.994 (11) 1.000 (1) 0.753 (6) 0.940 (7) 0.938 (11) 0.969 (13) 0.999 (12) 0.814 (8) 0.944 (9) 0.926 (5) 0.596 (6) 0.791 (11) 0.962 (12) 0.885 (6) 0.944 (6) 0.970 (10) 1.000 (9) 0.998 (7) 0.960 (9) 1.000 (3) 0.737 (13) 0.940 (11) 0.939 (5) 0.991 (9) 0.993 (6) 0.999 (5) 0.882 (6) 0.902 (11) 0.568 (14) 0.916 (14) 0.989 (12) 0.807 (14) 0.973 (9) 0.865 (14) 0.816 (11) 0.968 (10) 0.969 (13) 0.964 (6) 0.947 (13) 0.801 (12) 0.656 (11) 0.975 (15) 0.736 (12) 0.954 (12) 0.928 (11) 0.896 (11) 0.793 (10)
0.999 (9) 1.000 (6) 0.755 (5) 0.912 (15) 0.968 (10) 0.982 (9) 1.000 (10) 0.872 (4) 0.980 (4) 0.656 (13) 0.566 (8) 0.895 (4) 0.987 (10) 0.846 (11) 0.947 (3) 0.974 (9) 0.996 (12) 0.974 (12) 0.966 (7) 0.999 (5) 0.776 (11) 0.968 (9) 0.812 (14) 0.997 (4) 0.991 (7) 0.912 (14) 0.907 (4) 0.897 (12) 0.620 (11) 0.981 (6) 0.995 (11) 0.870 (10) 0.977 (6) 0.968 (7) 0.805 (12) 0.953 (14) 0.957 (15) 0.746 (11) 0.999 (6) 0.831 (9) 0.537 (14) 0.984 (13) 0.774 (11) 0.974 (9) 0.946 (8) 0.917 (10) 0.854 (8)
0.863 (16) 0.655 (16) 0.595 (18) 0.909 (18) 0.516 (18) 0.731 (16) 0.970 (16) 0.779 (11) 0.800 (18) 0.365 (18) 0.544 (13) 0.374 (18) 0.851 (17) 0.677 (17) 0.385 (17) 0.952 (12) 0.757 (18) 0.621 (16) 0.696 (16) 0.851 (17) 0.566 (18) 0.688 (16) 0.594 (18) 0.873 (18) 0.955 (16) 0.408 (18) 0.925 (1) 0.646 (15) 0.402 (18) 0.760 (18) 0.826 (18) 0.478 (16) 0.714 (18) 0.484 (18) 0.545 (17) 0.976 (9) 0.875 (16) 0.433 (16) 0.320 (18) 0.454 (18) 0.474 (16) 0.986 (11) 0.681 (18) 0.891 (15) 0.757 (15) 0.818 (14) 0.324 (18)
0.457 (18) 0.591 (18) 0.649 (16) 0.964 (1) 0.723 (17) 0.634 (18) 0.564 (18) 0.676 (16) 0.906 (14) 0.539 (16) 0.549 (11) 0.705 (16) 0.939 (13) 0.589 (18) 0.778 (16) 0.722 (18) 0.887 (17) 0.434 (18) 0.690 (17) 0.730 (18) 0.617 (16) 0.718 (15) 0.728 (16) 0.972 (13) 0.708 (18) 0.487 (17) 0.559 (17) 0.546 (17) 0.507 (17) 0.761 (17) 0.926 (15) 0.384 (17) 0.882 (15) 0.787 (17) 0.537 (18) 0.918 (16) 0.980 (12) 0.586 (15) 0.666 (17) 0.492 (17) 0.448 (17) 0.933 (16) 0.709 (16) 0.805 (17) 0.582 (18) 0.654 (16) 0.619 (16)
0.999 (3) 1.000 (5) 0.726 (12) 0.932 (8) 0.889 (14) 0.974 (12) 1.000 (1) 0.826 (7) 0.889 (16) 0.772 (11) 0.537 (14) 0.774 (14) 0.997 (6) 0.833 (14) 0.931 (9) 0.985 (6) 1.000 (2) 1.000 (4) 0.893 (15) 0.997 (9) 0.800 (6) 0.833 (13) 0.853 (10) 0.995 (6) 0.980 (14) 0.950 (12) 0.693 (14) 0.922 (9) 0.725 (1) 0.977 (9) 0.998 (3) 0.826 (13) 0.976 (8) 0.979 (5) 0.910 (7) 0.991 (6) 1.000 (2) 0.694 (13) 1.000 (1) 0.795 (13) 0.672 (9) 0.994 (3) 0.888 (8) 0.985 (7) 0.951 (4) 0.951 (6) 0.898 (5)
0.999 (6) 1.000 (3) 0.739 (10) 0.911 (17) 0.975 (9) 0.986 (8) 1.000 (2) 0.689 (15) 0.929 (10) 0.823 (9) 0.564 (9) 0.796 (10) 0.989 (9) 0.854 (8) 0.945 (4) 0.983 (7) 1.000 (4) 1.000 (2) 0.963 (8) 0.995 (12) 0.782 (9) 0.968 (8) 0.870 (9) 0.988 (10) 0.988 (10) 0.986 (10) 0.878 (7) 0.947 (5) 0.578 (13) 0.979 (7) 0.998 (4) 0.917 (8) 0.989 (3) 0.982 (4) 0.864 (10) 0.923 (15) 1.000 (3) 0.804 (10) 1.000 (2) 0.859 (6) 0.694 (5) 0.993 (4) 0.853 (9) 0.981 (8) 0.945 (9) 0.947 (7) 0.824 (9)
0.999 (4) 0.999 (9) 0.745 (9) 0.932 (9) 0.921 (13) 0.976 (11) 1.000 (4) 0.803 (10) 0.890 (15) 0.782 (10) 0.530 (17) 0.782 (13) 0.997 (4) 0.834 (13) 0.929 (11) 0.987 (5) 1.000 (6) 1.000 (5) 0.898 (13) 0.998 (8) 0.797 (7) 0.832 (14) 0.841 (11) 0.994 (7) 0.980 (15) 0.955 (11) 0.652 (16) 0.931 (8) 0.702 (3) 0.982 (5) 0.997 (5) 0.841 (11) 0.960 (10) 0.987 (3) 0.919 (4) 0.991 (4) 1.000 (4) 0.696 (12) 1.000 (3) 0.806 (11) 0.674 (8) 0.992 (5) 0.894 (6) 0.987 (5) 0.947 (6) 0.952 (5) 0.878 (7)
0.985 (14) 0.998 (12) 0.719 (13) 0.911 (16) 0.983 (7) 0.980 (10) 1.000 (11) 0.652 (17) 0.843 (17) 0.606 (15) 0.594 (7) 0.749 (15) 0.923 (15) 0.852 (9) 0.943 (7) 0.811 (17) 0.931 (15) 0.868 (13) 0.959 (10) 0.971 (14) 0.723 (15) 0.988 (4) 0.837 (12) 0.949 (15) 0.989 (9) 0.997 (9) 0.785 (12) 0.880 (13) 0.624 (10) 0.947 (13) 0.986 (14) 0.889 (9) 0.873 (17) 0.904 (12) 0.733 (15) 0.868 (17) 0.786 (17) 0.991 (2) 0.942 (14) 0.840 (8) 0.623 (12) 0.899 (18) 0.726 (13) 0.732 (18) 0.829 (14) 0.780 (15) 0.697 (13)
0.999 (7) 0.995 (14) 0.700 (15) 0.951 (2) 0.931 (12) 0.953 (14) 0.997 (13) 0.805 (9) 0.914 (12) 0.717 (12) 0.532 (16) 0.785 (12) 0.999 (1) 0.837 (12) 0.914 (13) 0.921 (14) 1.000 (3) 0.994 (10) 0.896 (14) 0.995 (10) 0.728 (14) 0.836 (12) 0.835 (13) 0.996 (5) 0.989 (8) 0.947 (13) 0.790 (11) 0.716 (14) 0.682 (4) 0.975 (10) 0.996 (9) 0.839 (12) 0.946 (11) 0.967 (8) 0.914 (6) 0.965 (12) 0.984 (11) 0.666 (14) 0.999 (5) 0.739 (14) 0.583 (13) 0.989 (9) 0.700 (17) 0.987 (6) 0.918 (12) 0.888 (12) 0.766 (11)
0.759 (17) 0.642 (17) 0.631 (17) 0.919 (14) 0.858 (15) 0.648 (17) 0.755 (17) 0.581 (18) 0.926 (11) 0.655 (14) 0.546 (12) 0.879 (6) 0.546 (18) 0.721 (16) 0.354 (18) 0.888 (15) 0.975 (14) 0.546 (17) 0.669 (18) 0.852 (16) 0.591 (17) 0.556 (18) 0.758 (15) 0.945 (17) 0.940 (17) 0.721 (16) 0.826 (9) 0.473 (18) 0.539 (15) 0.872 (16) 0.897 (17) 0.633 (15) 0.874 (16) 0.915 (11) 0.788 (13) 0.819 (18) 0.652 (18) 0.392 (17) 0.729 (16) 0.650 (15) 0.477 (15) 0.921 (17) 0.723 (14) 0.830 (16) 0.690 (17) 0.627 (17) 0.609 (17)
0.920 (15) 0.727 (15) 0.716 (14) 0.948 (3) 0.746 (16) 0.772 (15) 0.995 (14) 0.703 (14) 0.947 (8) 0.504 (17) 0.559 (10) 0.797 (9) 0.851 (16) 0.745 (15) 0.859 (15) 0.830 (16) 0.899 (16) 0.650 (15) 0.901 (12) 0.907 (15) 0.738 (12) 0.600 (17) 0.646 (17) 0.978 (12) 0.984 (13) 0.724 (15) 0.889 (5) 0.570 (16) 0.511 (16) 0.886 (15) 0.917 (16) 0.382 (18) 0.906 (13) 0.841 (15) 0.724 (16) 0.977 (8) 0.964 (14) 0.338 (18) 0.784 (15) 0.511 (16) 0.394 (18) 0.983 (14) 0.712 (15) 0.926 (14) 0.732 (16) 0.625 (18) 0.693 (14)
0.999 (8) 0.998 (13) 0.748 (8) 0.928 (12) 0.983 (8) 0.995 (3) 1.000 (8) 0.848 (6) 0.974 (6) 0.915 (6) 0.664 (4) 0.813 (7) 0.990 (8) 0.847 (10) 0.895 (14) 0.979 (8) 1.000 (10) 0.991 (11) 0.974 (5) 0.995 (11) 0.793 (8) 0.988 (6) 0.925 (6) 0.956 (14) 0.995 (5) 0.997 (8) 0.500 (18) 0.933 (7) 0.645 (8) 0.971 (11) 0.997 (8) 0.941 (4) 0.985 (4) 0.897 (13) 0.895 (9) 0.966 (11) 0.997 (9) 0.956 (7) 0.973 (12) 0.883 (5) 0.659 (10) 0.987 (10) 0.891 (7) 0.965 (10) 0.930 (10) 0.922 (9) 0.898 (6)
0.999 (10) 0.999 (8) 0.787 (3) 0.947 (4) 0.997 (1) 0.996 (2) 1.000 (9) 0.919 (3) 0.978 (5) 0.940 (3) 0.642 (5) 0.888 (5) 0.999 (2) 0.891 (5) 0.944 (5) 0.991 (3) 1.000 (1) 0.999 (6) 0.982 (3) 0.999 (4) 0.817 (3) 0.990 (3) 0.942 (3) 0.982 (11) 0.997 (4) 0.999 (3) 0.802 (10) 0.967 (3) 0.638 (9) 0.982 (4) 0.997 (6) 0.956 (2) 0.985 (5) 0.948 (10) 0.918 (5) 0.995 (2) 1.000 (1) 0.976 (5) 0.998 (8) 0.884 (4) 0.723 (2) 0.991 (6) 0.912 (4) 0.988 (3) 0.959 (3) 0.926 (8) 0.914 (4)
0.999 (5) 0.998 (10) 0.753 (7) 0.928 (13) 0.987 (6) 0.991 (5) 1.000 (5) 0.770 (12) 0.948 (7) 0.833 (8) 0.508 (18) 0.804 (8) 0.998 (3) 0.869 (7) 0.937 (8) 0.960 (11) 1.000 (11) 0.995 (9) 0.948 (11) 0.999 (6) 0.813 (4) 0.945 (10) 0.889 (8) 0.993 (8) 0.985 (12) 0.998 (6) 0.681 (15) 0.933 (6) 0.709 (2) 0.982 (3) 0.996 (10) 0.933 (6) 0.921 (12) 0.979 (6) 0.931 (3) 0.989 (7) 1.000 (7) 0.868 (9) 1.000 (4) 0.830 (10) 0.676 (7) 0.991 (7) 0.897 (5) 0.963 (11) 0.948 (5) 0.959 (3) 0.756 (12)
0.991 (13) 0.998 (11) 0.732 (11) 0.931 (11) 0.989 (5) 0.989 (7) 1.000 (6) 0.764 (13) 0.913 (13) 0.878 (7) 0.537 (15) 0.660 (17) 0.985 (11) 0.902 (3) 0.916 (12) 0.931 (13) 0.978 (13) 0.838 (14) 0.970 (6) 0.980 (13) 0.777 (10) 0.986 (7) 0.895 (7) 0.947 (16) 0.985 (11) 0.999 (4) 0.708 (13) 0.902 (10) 0.580 (12) 0.960 (12) 0.987 (13) 0.925 (7) 0.902 (14) 0.811 (16) 0.770 (14) 0.991 (5) 0.999 (8) 0.980 (4) 0.998 (9) 0.841 (7) 0.697 (4) 0.985 (12) 0.804 (10) 0.946 (13) 0.916 (13) 0.840 (13) 0.654 (15)
1.000 (2) 1.000 (2) 0.823 (2) 0.941 (6) 0.995 (2) 0.996 (1) 1.000 (7) 0.943 (2) 0.985 (2) 0.950 (2) 0.717 (3) 0.976 (1) 0.997 (5) 0.927 (2) 0.960 (1) 0.994 (2) 1.000 (7) 1.000 (1) 0.988 (2) 1.000 (1) 0.832 (2) 0.991 (2) 0.962 (2) 0.997 (3) 1.000 (1) 1.000 (2) 0.922 (2) 0.978 (2) 0.656 (7) 0.990 (1) 0.999 (1) 0.947 (3) 0.997 (2) 0.997 (1) 0.950 (1) 0.992 (3) 1.000 (5) 0.993 (1) 0.999 (7) 0.899 (2) 0.734 (1) 0.995 (2) 0.934 (2) 0.992 (1) 0.968 (1) 0.965 (1) 0.929 (1)
1.000 (1) 1.000 (4) 0.839 (1) 0.931 (10) 0.993 (3) 0.994 (4) 1.000 (3) 0.943 (1) 0.987 (1) 0.953 (1) 0.818 (1) 0.969 (2) 0.997 (7) 0.930 (1) 0.956 (2) 0.998 (1) 1.000 (5) 1.000 (3) 0.988 (1) 1.000 (2) 0.845 (1) 0.991 (1) 0.964 (1) 0.998 (1) 1.000 (2) 1.000 (1) 0.920 (3) 0.978 (1) 0.675 (5) 0.987 (2) 0.999 (2) 0.965 (1) 0.998 (1) 0.995 (2) 0.943 (2) 0.997 (1) 1.000 (6) 0.989 (3) 0.997 (10) 0.904 (1) 0.721 (3) 0.996 (1) 0.934 (1) 0.989 (2) 0.966 (2) 0.960 (2) 0.924 (2)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.969 (7) 0.962 (8) 0.942 (9) 0.940 (5) 0.676 (7)
0.911 (13) 0.916 (13) 0.918 (14) 0.834 (15) 0.621 (14)
0.955 (9) 0.956 (10) 0.944 (7) 0.912 (11) 0.669 (9)
0.878 (16) 0.886 (15) 0.730 (17) 0.696 (17) 0.523 (17)
0.750 (18) 0.831 (18) 0.662 (18) 0.841 (14) 0.513 (18)
0.973 (4) 0.970 (3) 0.947 (4) 0.951 (2) 0.728 (3)
0.953 (11) 0.961 (9) 0.950 (3) 0.927 (9) 0.712 (5)
0.971 (6) 0.964 (6) 0.946 (5) 0.928 (8) 0.722 (4)
0.941 (12) 0.932 (12) 0.944 (8) 0.870 (13) 0.676 (8)
0.972 (5) 0.968 (5) 0.944 (6) 0.951 (3) 0.708 (6)
0.843 (17) 0.856 (17) 0.742 (15) 0.693 (18) 0.544 (16)
0.891 (14) 0.902 (14) 0.734 (16) 0.744 (16) 0.567 (15)
0.955 (10) 0.953 (11) 0.936 (10) 0.937 (6) 0.667 (10)
0.967 (8) 0.962 (7) 0.935 (11) 0.944 (4) 0.643 (12)
0.978 (3) 0.969 (4) 0.930 (12) 0.933 (7) 0.658 (11)
0.887 (15) 0.869 (16) 0.927 (13) 0.908 (12) 0.636 (13)
0.981 (2) 0.973 (2) 0.954 (1) 0.963 (1) 0.744 (1)
0.981 (1) 0.974 (1) 0.951 (2) 0.913 (10) 0.736 (2)
20news agnews amazon imdb yelp
0.847 (7) 0.957 (7) 0.869 (9) 0.886 (9) 0.932 (10)
0.735 (15) 0.871 (13) 0.747 (14) 0.745 (13) 0.824 (14)
0.792 (12) 0.952 (9) 0.851 (11) 0.866 (12) 0.933 (9)
0.566 (18) 0.652 (18) 0.521 (17) 0.476 (17) 0.571 (16)
0.704 (16) 0.748 (16) 0.529 (16) 0.461 (18) 0.497 (18)
0.903 (1) 0.971 (2) 0.921 (2) 0.933 (1) 0.967 (2)
0.845 (9) 0.957 (6) 0.906 (5) 0.916 (5) 0.958 (6)
0.891 (3) 0.969 (4) 0.912 (4) 0.927 (4) 0.968 (1)
0.821 (11) 0.937 (12) 0.878 (7) 0.873 (11) 0.927 (12)
0.868 (6) 0.966 (5) 0.905 (6) 0.916 (6) 0.964 (4)
0.620 (17) 0.732 (17) 0.471 (18) 0.523 (16) 0.502 (17)
0.738 (14) 0.855 (14) 0.534 (15) 0.526 (15) 0.572 (15)
0.847 (8) 0.952 (10) 0.872 (8) 0.880 (10) 0.935 (7)
0.841 (10) 0.951 (11) 0.866 (10) 0.892 (8) 0.935 (8)
0.892 (2) 0.954 (8) 0.830 (13) 0.894 (7) 0.928 (11)
0.740 (13) 0.824 (15) 0.835 (12) 0.740 (14) 0.857 (13)
0.883 (4) 0.971 (1) 0.923 (1) 0.931 (2) 0.964 (5)
0.873 (5) 0.971 (3) 0.918 (3) 0.927 (3) 0.964 (3)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 45: The average AUCROC performance on individual datasets under the 𝛾𝑙𝑎 = 50% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.998 (12) 1.000 (11) 0.809 (4) 0.948 (6) 0.998 (4) 0.999 (10) 0.983 (15) 0.919 (6) 0.989 (4) 0.943 (4) 0.802 (2) 0.938 (3) 1.000 (11) 0.904 (6) 0.946 (7) 0.994 (4) 1.000 (12) 0.998 (8) 0.982 (6) 0.999 (7) 0.869 (4) 0.992 (4) 0.954 (4) 0.996 (5) 1.000 (3) 0.999 (7) 0.986 (5) 0.973 (4) 0.693 (8) 0.994 (5) 0.999 (4) 0.976 (5) 0.976 (9) 0.985 (10) 0.913 (7) 0.999 (1) 0.996 (13) 0.981 (7) 1.000 (13) 0.940 (5) 0.716 (6) 0.991 (9) 0.930 (3) 0.992 (4) 0.966 (5) 0.960 (4) 0.931 (4)
0.999 (11) 1.000 (4) 0.788 (8) 0.936 (12) 0.969 (11) 0.998 (11) 1.000 (8) 0.899 (9) 0.975 (8) 0.942 (5) 0.631 (7) 0.836 (11) 0.998 (13) 0.904 (5) 0.948 (5) 0.988 (8) 1.000 (2) 1.000 (6) 0.972 (8) 1.000 (3) 0.796 (11) 0.958 (11) 0.951 (5) 0.992 (6) 0.993 (7) 0.999 (6) 0.995 (4) 0.941 (9) 0.609 (15) 0.971 (13) 0.992 (13) 0.886 (11) 0.985 (7) 0.911 (15) 0.852 (12) 0.986 (12) 0.997 (10) 0.982 (6) 0.992 (14) 0.890 (10) 0.689 (10) 0.985 (13) 0.794 (13) 0.975 (12) 0.956 (8) 0.927 (11) 0.829 (11)
1.000 (4) 1.000 (2) 0.791 (7) 0.947 (7) 0.979 (10) 1.000 (3) 1.000 (10) 0.934 (4) 0.985 (6) 0.661 (14) 0.577 (9) 0.921 (5) 1.000 (7) 0.853 (11) 0.953 (3) 0.988 (10) 0.999 (13) 0.998 (10) 0.968 (10) 1.000 (6) 0.795 (12) 0.970 (10) 0.813 (14) 0.997 (2) 0.992 (8) 0.912 (14) 0.875 (15) 0.926 (12) 0.699 (7) 0.993 (6) 0.997 (10) 0.906 (10) 0.989 (5) 0.993 (5) 0.858 (11) 0.979 (14) 0.997 (11) 0.756 (11) 1.000 (7) 0.886 (11) 0.563 (14) 0.990 (10) 0.848 (11) 0.988 (9) 0.958 (7) 0.938 (10) 0.874 (9)
0.933 (16) 0.562 (16) 0.617 (18) 0.840 (18) 0.543 (18) 0.811 (16) 0.901 (18) 0.849 (12) 0.802 (18) 0.364 (18) 0.539 (16) 0.407 (18) 0.713 (18) 0.689 (17) 0.468 (17) 0.910 (15) 0.768 (18) 0.384 (18) 0.664 (18) 0.793 (17) 0.598 (18) 0.647 (16) 0.593 (18) 0.930 (18) 0.950 (17) 0.434 (18) 0.904 (12) 0.752 (15) 0.403 (18) 0.722 (18) 0.864 (18) 0.480 (16) 0.806 (18) 0.455 (18) 0.627 (17) 0.981 (13) 0.845 (17) 0.556 (15) 0.388 (18) 0.462 (18) 0.437 (17) 0.983 (15) 0.740 (14) 0.842 (17) 0.785 (15) 0.822 (15) 0.455 (18)
0.484 (18) 0.539 (18) 0.659 (16) 0.965 (2) 0.727 (17) 0.700 (17) 0.979 (16) 0.681 (17) 0.912 (16) 0.566 (15) 0.555 (12) 0.707 (17) 0.952 (15) 0.589 (18) 0.790 (16) 0.728 (18) 0.960 (17) 0.439 (17) 0.766 (16) 0.724 (18) 0.632 (16) 0.723 (15) 0.740 (16) 0.976 (15) 0.646 (18) 0.491 (17) 0.616 (17) 0.566 (17) 0.508 (17) 0.770 (17) 0.928 (16) 0.399 (17) 0.896 (16) 0.792 (17) 0.534 (18) 0.922 (17) 0.982 (14) 0.521 (16) 0.715 (17) 0.509 (17) 0.449 (16) 0.938 (17) 0.639 (18) 0.829 (18) 0.602 (18) 0.678 (16) 0.605 (17)
0.999 (6) 1.000 (8) 0.739 (13) 0.941 (9) 0.887 (15) 0.982 (13) 1.000 (1) 0.914 (8) 0.929 (15) 0.769 (11) 0.547 (15) 0.817 (13) 1.000 (2) 0.835 (14) 0.932 (12) 0.989 (7) 1.000 (3) 1.000 (5) 0.896 (15) 0.997 (10) 0.812 (10) 0.833 (14) 0.857 (10) 0.991 (7) 0.980 (16) 0.951 (12) 0.924 (9) 0.925 (13) 0.777 (1) 0.978 (12) 0.997 (8) 0.837 (13) 0.978 (8) 0.991 (8) 0.909 (8) 0.995 (7) 1.000 (2) 0.695 (13) 1.000 (2) 0.822 (13) 0.692 (9) 0.995 (4) 0.894 (8) 0.990 (6) 0.953 (10) 0.956 (5) 0.916 (5)
1.000 (5) 1.000 (6) 0.772 (10) 0.892 (17) 0.988 (8) 1.000 (9) 1.000 (2) 0.859 (11) 0.970 (9) 0.832 (9) 0.603 (8) 0.859 (10) 1.000 (4) 0.855 (10) 0.950 (4) 0.988 (9) 1.000 (5) 1.000 (3) 0.968 (11) 0.995 (14) 0.812 (9) 0.970 (9) 0.875 (9) 0.988 (10) 0.991 (9) 0.984 (10) 0.972 (7) 0.950 (7) 0.631 (11) 0.990 (9) 0.998 (7) 0.936 (9) 0.988 (6) 0.996 (4) 0.894 (10) 0.990 (11) 1.000 (3) 0.814 (10) 1.000 (4) 0.897 (9) 0.724 (5) 0.996 (1) 0.873 (9) 0.990 (5) 0.955 (9) 0.953 (7) 0.883 (8)
0.999 (7) 0.999 (13) 0.757 (12) 0.964 (3) 0.907 (13) 0.990 (12) 1.000 (4) 0.918 (7) 0.938 (14) 0.785 (10) 0.548 (14) 0.831 (12) 1.000 (6) 0.836 (13) 0.932 (11) 0.990 (6) 1.000 (7) 1.000 (4) 0.904 (13) 0.997 (11) 0.813 (8) 0.833 (13) 0.848 (12) 0.990 (9) 0.980 (15) 0.954 (11) 0.910 (11) 0.937 (10) 0.743 (3) 0.982 (11) 0.997 (9) 0.848 (12) 0.944 (12) 0.992 (6) 0.919 (5) 0.992 (10) 1.000 (5) 0.698 (12) 1.000 (6) 0.835 (12) 0.692 (8) 0.994 (6) 0.906 (7) 0.990 (7) 0.953 (11) 0.955 (6) 0.910 (7)
0.994 (14) 1.000 (5) 0.765 (11) 0.894 (15) 0.986 (9) 1.000 (4) 1.000 (11) 0.801 (15) 0.883 (17) 0.467 (17) 0.652 (6) 0.861 (9) 0.966 (14) 0.873 (8) 0.945 (8) 0.903 (16) 1.000 (8) 0.971 (14) 0.979 (7) 0.999 (9) 0.784 (13) 0.990 (7) 0.851 (11) 0.970 (16) 0.990 (10) 0.997 (9) 0.997 (3) 0.931 (11) 0.690 (10) 0.992 (8) 0.998 (6) 0.939 (8) 0.951 (10) 0.988 (9) 0.770 (14) 0.957 (16) 0.940 (16) 0.993 (1) 1.000 (8) 0.905 (7) 0.622 (12) 0.983 (16) 0.812 (12) 0.843 (16) 0.912 (14) 0.857 (14) 0.726 (14)
0.999 (10) 0.994 (14) 0.735 (15) 0.967 (1) 0.957 (12) 0.957 (14) 0.997 (13) 0.847 (13) 0.946 (13) 0.740 (12) 0.537 (17) 0.754 (16) 1.000 (3) 0.845 (12) 0.913 (13) 0.927 (13) 1.000 (4) 0.996 (12) 0.900 (14) 0.996 (13) 0.750 (14) 0.848 (12) 0.814 (13) 0.997 (4) 0.990 (11) 0.949 (13) 0.892 (13) 0.752 (14) 0.742 (4) 0.970 (14) 0.996 (11) 0.834 (14) 0.943 (13) 0.983 (11) 0.917 (6) 0.994 (8) 0.996 (12) 0.658 (14) 1.000 (3) 0.757 (14) 0.580 (13) 0.991 (8) 0.725 (16) 0.989 (8) 0.928 (13) 0.890 (13) 0.803 (12)
0.922 (17) 0.555 (17) 0.650 (17) 0.939 (11) 0.895 (14) 0.685 (18) 0.912 (17) 0.679 (18) 0.950 (12) 0.688 (13) 0.554 (13) 0.878 (6) 0.749 (17) 0.693 (16) 0.296 (18) 0.921 (14) 0.995 (15) 0.682 (16) 0.737 (17) 0.908 (16) 0.610 (17) 0.568 (18) 0.786 (15) 0.991 (8) 0.986 (13) 0.797 (15) 0.790 (16) 0.518 (18) 0.616 (12) 0.852 (16) 0.908 (17) 0.563 (15) 0.893 (17) 0.916 (14) 0.728 (16) 0.866 (18) 0.811 (18) 0.426 (17) 0.784 (16) 0.651 (15) 0.471 (15) 0.930 (18) 0.738 (15) 0.878 (15) 0.680 (17) 0.626 (18) 0.663 (16)
0.933 (15) 0.788 (15) 0.735 (14) 0.949 (5) 0.759 (16) 0.881 (15) 0.995 (14) 0.746 (16) 0.953 (11) 0.544 (16) 0.561 (11) 0.797 (15) 0.939 (16) 0.758 (15) 0.860 (15) 0.856 (17) 0.969 (16) 0.738 (15) 0.906 (12) 0.943 (15) 0.735 (15) 0.602 (17) 0.704 (17) 0.984 (12) 0.986 (12) 0.770 (16) 0.884 (14) 0.585 (16) 0.510 (16) 0.888 (15) 0.928 (15) 0.394 (18) 0.920 (15) 0.858 (16) 0.739 (15) 0.978 (15) 0.980 (15) 0.328 (18) 0.850 (15) 0.538 (16) 0.398 (18) 0.983 (14) 0.721 (17) 0.934 (14) 0.762 (16) 0.644 (17) 0.699 (15)
0.999 (9) 1.000 (12) 0.803 (5) 0.933 (14) 0.996 (7) 1.000 (8) 1.000 (12) 0.921 (5) 0.988 (5) 0.941 (6) 0.701 (5) 0.876 (7) 1.000 (10) 0.884 (7) 0.903 (14) 0.992 (5) 1.000 (11) 0.997 (11) 0.982 (5) 0.999 (8) 0.870 (3) 0.992 (3) 0.950 (6) 0.968 (17) 0.998 (5) 0.999 (8) 0.500 (18) 0.963 (5) 0.691 (9) 0.994 (4) 0.998 (5) 0.979 (3) 0.993 (4) 0.973 (12) 0.899 (9) 0.993 (9) 0.999 (9) 0.980 (8) 1.000 (12) 0.952 (3) 0.703 (7) 0.993 (7) 0.916 (6) 0.986 (10) 0.967 (4) 0.945 (9) 0.916 (6)
0.999 (8) 1.000 (10) 0.836 (3) 0.962 (4) 0.999 (1) 1.000 (1) 1.000 (9) 0.947 (3) 0.989 (3) 0.953 (3) 0.710 (4) 0.931 (4) 1.000 (1) 0.914 (4) 0.947 (6) 0.995 (3) 1.000 (1) 0.999 (7) 0.988 (3) 1.000 (4) 0.877 (2) 0.992 (5) 0.961 (3) 0.986 (11) 0.999 (4) 0.999 (4) 1.000 (1) 0.976 (3) 0.615 (13) 0.996 (1) 0.999 (3) 0.982 (2) 0.995 (3) 0.992 (7) 0.927 (4) 0.998 (2) 1.000 (1) 0.988 (4) 1.000 (1) 0.950 (4) 0.755 (3) 0.994 (5) 0.928 (4) 0.994 (2) 0.972 (3) 0.953 (8) 0.934 (3)
1.000 (3) 1.000 (7) 0.776 (9) 0.939 (10) 0.997 (5) 1.000 (5) 1.000 (5) 0.877 (10) 0.980 (7) 0.844 (8) 0.530 (18) 0.864 (8) 1.000 (8) 0.869 (9) 0.943 (9) 0.984 (11) 1.000 (9) 0.998 (9) 0.969 (9) 1.000 (5) 0.829 (6) 0.986 (8) 0.894 (8) 0.981 (13) 0.984 (14) 0.999 (5) 0.921 (10) 0.954 (6) 0.749 (2) 0.993 (7) 0.996 (12) 0.953 (6) 0.949 (11) 0.997 (3) 0.943 (2) 0.997 (5) 1.000 (6) 0.866 (9) 1.000 (9) 0.903 (8) 0.668 (11) 0.989 (11) 0.922 (5) 0.977 (11) 0.961 (6) 0.961 (3) 0.832 (10)
0.998 (13) 1.000 (9) 0.793 (6) 0.894 (16) 0.997 (6) 1.000 (7) 1.000 (7) 0.842 (14) 0.957 (10) 0.926 (7) 0.577 (10) 0.810 (14) 0.999 (12) 0.922 (3) 0.940 (10) 0.973 (12) 0.999 (14) 0.974 (13) 0.984 (4) 0.996 (12) 0.827 (7) 0.990 (6) 0.938 (7) 0.981 (14) 0.994 (6) 1.000 (3) 0.954 (8) 0.942 (8) 0.610 (14) 0.987 (10) 0.988 (14) 0.945 (7) 0.932 (14) 0.952 (13) 0.816 (13) 0.998 (3) 1.000 (8) 0.986 (5) 1.000 (11) 0.925 (6) 0.748 (4) 0.988 (12) 0.870 (10) 0.972 (13) 0.940 (12) 0.894 (12) 0.740 (13)
1.000 (2) 1.000 (3) 0.870 (2) 0.935 (13) 0.999 (3) 1.000 (6) 1.000 (6) 0.967 (1) 0.991 (2) 0.967 (2) 0.778 (3) 0.988 (1) 1.000 (9) 0.936 (2) 0.964 (1) 0.997 (2) 1.000 (10) 1.000 (2) 0.992 (1) 1.000 (1) 0.861 (5) 0.994 (2) 0.975 (2) 0.997 (3) 1.000 (1) 1.000 (2) 0.982 (6) 0.983 (1) 0.740 (5) 0.996 (3) 0.999 (2) 0.978 (4) 0.998 (2) 0.999 (1) 0.949 (1) 0.996 (6) 1.000 (7) 0.993 (2) 1.000 (10) 0.960 (2) 0.771 (1) 0.996 (2) 0.941 (1) 0.995 (1) 0.980 (1) 0.967 (2) 0.939 (1)
1.000 (1) 1.000 (1) 0.883 (1) 0.941 (8) 0.999 (2) 1.000 (2) 1.000 (3) 0.958 (2) 0.993 (1) 0.969 (1) 0.865 (1) 0.985 (2) 1.000 (5) 0.937 (1) 0.955 (2) 0.999 (1) 1.000 (6) 1.000 (1) 0.991 (2) 1.000 (2) 0.891 (1) 0.995 (1) 0.977 (1) 0.998 (1) 1.000 (2) 1.000 (1) 1.000 (2) 0.981 (2) 0.739 (6) 0.996 (2) 0.999 (1) 0.987 (1) 1.000 (1) 0.998 (2) 0.941 (3) 0.997 (4) 1.000 (4) 0.991 (3) 1.000 (5) 0.966 (1) 0.761 (2) 0.996 (3) 0.940 (2) 0.993 (3) 0.979 (2) 0.968 (1) 0.934 (2)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.977 (6) 0.973 (6) 0.949 (6) 0.973 (6) 0.716 (7)
0.946 (13) 0.950 (13) 0.945 (12) 0.939 (14) 0.675 (13)
0.976 (9) 0.974 (5) 0.947 (10) 0.963 (12) 0.710 (9)
0.876 (16) 0.892 (16) 0.726 (17) 0.714 (18) 0.522 (17)
0.751 (18) 0.841 (18) 0.686 (18) 0.848 (15) 0.513 (18)
0.980 (4) 0.975 (3) 0.951 (5) 0.974 (5) 0.760 (4)
0.973 (10) 0.974 (4) 0.953 (3) 0.980 (2) 0.762 (3)
0.977 (7) 0.971 (10) 0.949 (7) 0.964 (11) 0.753 (5)
0.965 (12) 0.960 (12) 0.952 (4) 0.947 (13) 0.708 (10)
0.976 (8) 0.973 (7) 0.947 (9) 0.974 (4) 0.728 (6)
0.839 (17) 0.862 (17) 0.789 (15) 0.727 (17) 0.544 (16)
0.904 (15) 0.916 (15) 0.756 (16) 0.829 (16) 0.570 (15)
0.972 (11) 0.969 (11) 0.946 (11) 0.971 (8) 0.714 (8)
0.979 (5) 0.971 (9) 0.945 (13) 0.976 (3) 0.702 (11)
0.982 (3) 0.972 (8) 0.930 (14) 0.971 (7) 0.662 (14)
0.926 (14) 0.920 (14) 0.949 (8) 0.964 (10) 0.681 (12)
0.986 (1) 0.981 (2) 0.960 (1) 0.987 (1) 0.788 (1)
0.986 (2) 0.981 (1) 0.955 (2) 0.968 (9) 0.773 (2)
20news agnews amazon imdb yelp
0.885 (10) 0.965 (9) 0.893 (10) 0.910 (10) 0.953 (11)
0.822 (13) 0.916 (13) 0.817 (13) 0.822 (14) 0.898 (14)
0.866 (12) 0.970 (5) 0.892 (11) 0.917 (9) 0.964 (7)
0.573 (18) 0.654 (18) 0.519 (17) 0.480 (17) 0.572 (16)
0.708 (16) 0.743 (17) 0.538 (15) 0.476 (18) 0.534 (17)
0.925 (2) 0.973 (3) 0.930 (3) 0.943 (2) 0.976 (2)
0.886 (9) 0.968 (7) 0.930 (4) 0.937 (5) 0.975 (5)
0.926 (1) 0.971 (4) 0.923 (5) 0.939 (3) 0.975 (4)
0.875 (11) 0.955 (11) 0.911 (7) 0.919 (8) 0.957 (9)
0.917 (5) 0.969 (6) 0.914 (6) 0.928 (6) 0.970 (6)
0.634 (17) 0.771 (16) 0.512 (18) 0.572 (15) 0.510 (18)
0.748 (15) 0.863 (15) 0.536 (16) 0.532 (16) 0.582 (15)
0.893 (7) 0.961 (10) 0.897 (9) 0.909 (11) 0.953 (10)
0.887 (8) 0.965 (8) 0.911 (8) 0.919 (7) 0.957 (8)
0.923 (3) 0.925 (12) 0.811 (14) 0.868 (12) 0.918 (12)
0.818 (14) 0.903 (14) 0.851 (12) 0.854 (13) 0.903 (13)
0.921 (4) 0.977 (1) 0.940 (1) 0.949 (1) 0.978 (1)
0.915 (6) 0.975 (2) 0.931 (2) 0.939 (4) 0.976 (3)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 46: The average AUCROC performance on individual datasets under the 𝛾𝑙𝑎 = 100% setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
1.000 (10) 1.000 (10) 0.856 (6) 0.974 (2) 1.000 (8) 0.999 (11) 1.000 (9) 0.955 (11) 0.998 (3) 0.958 (6) 0.827 (3) 0.963 (4) 1.000 (15) 0.918 (6) 0.957 (4) 0.997 (6) 1.000 (16) 0.999 (10) 0.988 (6) 1.000 (8) 0.909 (5) 0.996 (2) 0.966 (7) 0.993 (9) 1.000 (5) 1.000 (4) 0.990 (6) 0.982 (5) 0.785 (7) 0.999 (5) 1.000 (4) 0.996 (4) 0.993 (7) 0.989 (12) 0.933 (8) 1.000 (9) 1.000 (14) 0.990 (6) 1.000 (13) 0.998 (5) 0.767 (6) 0.999 (6) 0.937 (4) 0.998 (3) 0.985 (5) 0.967 (3) 0.944 (4)
1.000 (8) 1.000 (8) 0.831 (8) 0.939 (12) 0.987 (10) 1.000 (2) 1.000 (10) 0.960 (10) 0.992 (8) 0.954 (7) 0.672 (6) 0.859 (14) 1.000 (3) 0.914 (7) 0.955 (6) 0.996 (9) 1.000 (3) 1.000 (4) 0.983 (9) 1.000 (6) 0.843 (9) 0.981 (9) 0.966 (6) 0.992 (10) 0.996 (7) 0.998 (8) 0.995 (3) 0.970 (8) 0.678 (14) 0.987 (11) 0.996 (13) 0.933 (10) 0.995 (6) 0.966 (14) 0.913 (11) 0.993 (13) 1.000 (2) 0.988 (8) 0.999 (14) 0.964 (8) 0.722 (8) 0.997 (9) 0.870 (13) 0.993 (9) 0.974 (10) 0.955 (10) 0.884 (10)
1.000 (7) 1.000 (3) 0.823 (10) 0.962 (6) 0.977 (11) 1.000 (4) 1.000 (12) 0.973 (2) 0.995 (6) 0.668 (15) 0.587 (11) 0.939 (7) 1.000 (9) 0.866 (9) 0.955 (5) 0.997 (5) 1.000 (10) 0.995 (14) 0.972 (10) 1.000 (7) 0.810 (13) 0.978 (10) 0.833 (12) 0.998 (4) 0.988 (11) 0.902 (13) 0.874 (15) 0.952 (11) 0.784 (8) 0.994 (9) 0.998 (10) 0.921 (11) 0.996 (5) 0.997 (6) 0.910 (12) 1.000 (4) 1.000 (7) 0.762 (11) 1.000 (7) 0.946 (10) 0.630 (13) 0.998 (8) 0.903 (9) 0.997 (6) 0.974 (9) 0.952 (11) 0.913 (9)
0.862 (17) 0.535 (18) 0.667 (18) 0.846 (18) 0.567 (18) 0.832 (16) 0.880 (17) 0.791 (16) 0.792 (18) 0.369 (18) 0.536 (17) 0.453 (18) 0.878 (18) 0.748 (15) 0.318 (18) 0.923 (17) 0.882 (18) 0.676 (17) 0.674 (18) 0.872 (17) 0.639 (18) 0.698 (16) 0.629 (18) 0.944 (18) 0.981 (15) 0.537 (18) 0.927 (10) 0.368 (18) 0.465 (18) 0.691 (18) 0.798 (18) 0.500 (16) 0.872 (18) 0.484 (18) 0.678 (17) 0.985 (16) 0.951 (18) 0.558 (16) 0.601 (18) 0.496 (18) 0.442 (17) 0.982 (15) 0.764 (15) 0.862 (18) 0.755 (17) 0.808 (15) 0.300 (18)
0.448 (18) 0.594 (17) 0.681 (17) 0.964 (5) 0.732 (17) 0.777 (18) 0.870 (18) 0.694 (17) 0.924 (17) 0.584 (17) 0.551 (14) 0.712 (17) 0.952 (17) 0.633 (17) 0.773 (16) 0.777 (18) 1.000 (7) 0.489 (18) 0.884 (16) 0.720 (18) 0.664 (17) 0.811 (15) 0.812 (15) 0.989 (14) 0.822 (18) 0.543 (17) 0.575 (18) 0.622 (16) 0.509 (17) 0.797 (17) 0.934 (15) 0.432 (18) 0.905 (17) 0.807 (17) 0.535 (18) 0.927 (18) 0.985 (17) 0.567 (15) 0.819 (17) 0.540 (17) 0.450 (16) 0.969 (17) 0.722 (18) 0.867 (17) 0.670 (18) 0.606 (18) 0.645 (16)
1.000 (13) 1.000 (12) 0.759 (15) 0.961 (7) 0.891 (15) 0.978 (14) 1.000 (2) 0.967 (5) 0.948 (16) 0.777 (12) 0.547 (15) 0.861 (13) 1.000 (4) 0.839 (12) 0.932 (13) 0.989 (13) 1.000 (4) 1.000 (8) 0.903 (14) 0.997 (13) 0.816 (12) 0.833 (14) 0.829 (13) 0.994 (6) 0.980 (16) 0.951 (12) 0.921 (11) 0.930 (13) 0.851 (1) 0.980 (13) 0.998 (11) 0.834 (13) 0.979 (12) 0.992 (10) 0.927 (10) 0.996 (11) 1.000 (3) 0.696 (13) 1.000 (2) 0.844 (13) 0.686 (12) 0.996 (12) 0.895 (10) 0.991 (14) 0.954 (13) 0.960 (6) 0.930 (6)
1.000 (11) 1.000 (9) 0.806 (11) 0.931 (14) 0.988 (9) 0.999 (10) 1.000 (3) 0.961 (9) 0.983 (11) 0.834 (9) 0.647 (9) 0.899 (9) 1.000 (6) 0.861 (10) 0.948 (10) 0.990 (12) 1.000 (6) 1.000 (6) 0.967 (11) 0.999 (10) 0.824 (10) 0.960 (11) 0.872 (9) 0.989 (13) 0.987 (13) 0.978 (9) 0.988 (7) 0.958 (10) 0.732 (12) 0.990 (10) 0.998 (9) 0.949 (9) 0.990 (11) 0.996 (8) 0.940 (5) 1.000 (2) 1.000 (4) 0.817 (10) 1.000 (4) 0.946 (11) 0.730 (7) 0.997 (11) 0.883 (11) 0.995 (8) 0.963 (11) 0.959 (9) 0.915 (8)
1.000 (12) 0.999 (13) 0.775 (12) 0.969 (3) 0.904 (13) 0.993 (13) 1.000 (5) 0.967 (4) 0.951 (15) 0.784 (10) 0.543 (16) 0.863 (12) 1.000 (8) 0.837 (13) 0.934 (11) 0.991 (11) 1.000 (9) 1.000 (7) 0.906 (13) 0.997 (12) 0.821 (11) 0.835 (13) 0.838 (11) 0.993 (7) 0.980 (17) 0.954 (11) 0.873 (16) 0.938 (12) 0.849 (2) 0.984 (12) 0.997 (12) 0.848 (12) 0.950 (13) 0.991 (11) 0.933 (9) 0.991 (15) 1.000 (6) 0.697 (12) 1.000 (6) 0.850 (12) 0.688 (11) 0.996 (13) 0.910 (7) 0.991 (12) 0.956 (12) 0.959 (8) 0.924 (7)
1.000 (4) 1.000 (4) 0.827 (9) 0.932 (13) 1.000 (3) 1.000 (5) 1.000 (13) 0.964 (7) 0.991 (9) 0.782 (11) 0.666 (7) 0.953 (5) 1.000 (10) 0.896 (8) 0.950 (8) 0.997 (7) 1.000 (11) 1.000 (3) 0.984 (8) 1.000 (4) 0.859 (7) 0.991 (8) 0.846 (10) 0.990 (12) 0.994 (8) 0.999 (7) 0.995 (4) 0.967 (9) 0.760 (11) 0.999 (6) 0.999 (6) 0.980 (7) 0.991 (10) 0.998 (5) 0.873 (14) 1.000 (5) 1.000 (8) 0.993 (1) 1.000 (8) 0.956 (9) 0.719 (9) 0.979 (16) 0.871 (12) 0.993 (10) 0.977 (8) 0.916 (13) 0.853 (13)
0.999 (14) 0.998 (14) 0.767 (13) 0.941 (11) 0.936 (12) 0.974 (15) 0.997 (14) 0.878 (14) 0.965 (13) 0.714 (13) 0.529 (18) 0.784 (16) 1.000 (5) 0.854 (11) 0.915 (14) 0.968 (14) 1.000 (5) 0.996 (13) 0.903 (15) 0.996 (14) 0.782 (14) 0.858 (12) 0.822 (14) 0.998 (5) 0.988 (10) 0.958 (10) 0.928 (9) 0.821 (14) 0.805 (5) 0.975 (14) 0.996 (14) 0.819 (14) 0.937 (15) 0.981 (13) 0.934 (6) 0.995 (12) 0.999 (15) 0.654 (14) 1.000 (3) 0.770 (14) 0.626 (14) 0.997 (10) 0.795 (14) 0.991 (13) 0.936 (14) 0.903 (14) 0.863 (12)
0.890 (16) 0.775 (16) 0.700 (16) 0.931 (16) 0.904 (14) 0.800 (17) 0.973 (16) 0.667 (18) 0.955 (14) 0.710 (14) 0.565 (12) 0.867 (11) 0.962 (16) 0.720 (16) 0.338 (17) 0.942 (16) 0.998 (17) 0.777 (16) 0.804 (17) 0.957 (16) 0.707 (16) 0.598 (18) 0.801 (16) 0.993 (8) 0.986 (14) 0.853 (16) 0.705 (17) 0.590 (17) 0.550 (15) 0.936 (15) 0.851 (17) 0.624 (15) 0.923 (16) 0.943 (15) 0.740 (16) 0.978 (17) 1.000 (9) 0.396 (17) 0.900 (16) 0.663 (15) 0.474 (15) 0.951 (18) 0.737 (17) 0.901 (16) 0.780 (16) 0.707 (16) 0.616 (17)
0.940 (15) 0.974 (15) 0.764 (14) 0.949 (10) 0.765 (16) 0.996 (12) 0.995 (15) 0.857 (15) 0.976 (12) 0.626 (16) 0.561 (13) 0.804 (15) 1.000 (1) 0.797 (14) 0.863 (15) 0.947 (15) 1.000 (1) 0.987 (15) 0.917 (12) 0.993 (15) 0.761 (15) 0.624 (17) 0.766 (17) 0.999 (3) 0.988 (12) 0.863 (15) 0.889 (14) 0.630 (15) 0.510 (16) 0.912 (16) 0.933 (16) 0.441 (17) 0.947 (14) 0.855 (16) 0.749 (15) 0.992 (14) 0.997 (16) 0.332 (18) 0.979 (15) 0.598 (16) 0.413 (18) 0.996 (14) 0.737 (16) 0.966 (15) 0.827 (15) 0.648 (17) 0.714 (15)
1.000 (9) 1.000 (11) 0.891 (4) 0.980 (1) 1.000 (7) 1.000 (9) 1.000 (11) 0.954 (12) 0.996 (5) 0.967 (4) 0.761 (5) 0.952 (6) 1.000 (14) 0.924 (5) 0.934 (12) 0.998 (4) 1.000 (15) 0.998 (12) 0.990 (4) 0.999 (11) 0.933 (3) 0.995 (5) 0.973 (4) 0.979 (17) 1.000 (4) 1.000 (6) 0.966 (8) 0.983 (4) 0.775 (10) 1.000 (4) 0.999 (5) 0.996 (5) 0.999 (4) 0.996 (7) 0.934 (7) 1.000 (8) 1.000 (13) 0.991 (3) 1.000 (12) 0.998 (6) 0.773 (5) 0.999 (7) 0.932 (5) 0.996 (7) 0.986 (4) 0.959 (7) 0.938 (5)
1.000 (6) 1.000 (1) 0.897 (3) 0.967 (4) 1.000 (1) 1.000 (1) 1.000 (1) 0.966 (6) 0.997 (4) 0.968 (3) 0.781 (4) 0.970 (3) 1.000 (2) 0.932 (3) 0.958 (3) 0.998 (3) 1.000 (2) 1.000 (9) 0.992 (3) 1.000 (9) 0.932 (4) 0.996 (4) 0.974 (3) 0.991 (11) 1.000 (1) 1.000 (3) 1.000 (1) 0.984 (3) 0.789 (6) 1.000 (1) 1.000 (2) 0.996 (3) 0.999 (3) 0.998 (4) 0.949 (4) 1.000 (1) 1.000 (1) 0.990 (7) 1.000 (1) 1.000 (3) 0.797 (1) 0.999 (4) 0.940 (3) 0.997 (5) 0.988 (3) 0.963 (5) 0.946 (2)
1.000 (5) 1.000 (5) 0.835 (7) 0.952 (8) 1.000 (4) 1.000 (6) 1.000 (6) 0.962 (8) 0.993 (7) 0.927 (8) 0.592 (10) 0.916 (8) 1.000 (11) 0.500 (18) 0.950 (9) 0.994 (10) 1.000 (12) 0.998 (11) 0.986 (7) 1.000 (5) 0.849 (8) 0.992 (7) 0.936 (8) 0.988 (15) 0.989 (9) 0.900 (14) 0.921 (12) 0.977 (6) 0.824 (4) 0.998 (8) 0.999 (8) 0.970 (8) 0.993 (8) 0.999 (3) 0.960 (2) 1.000 (6) 1.000 (10) 0.921 (9) 1.000 (9) 0.973 (7) 0.717 (10) 0.999 (3) 0.930 (6) 0.997 (4) 0.980 (7) 0.965 (4) 0.871 (11)
1.000 (3) 1.000 (7) 0.860 (5) 0.912 (17) 1.000 (6) 1.000 (8) 1.000 (8) 0.938 (13) 0.987 (10) 0.964 (5) 0.653 (8) 0.899 (10) 1.000 (13) 0.931 (4) 0.950 (7) 0.997 (8) 1.000 (14) 1.000 (5) 0.989 (5) 1.000 (3) 0.891 (6) 0.992 (6) 0.971 (5) 0.986 (16) 0.999 (6) 1.000 (5) 0.912 (13) 0.973 (7) 0.695 (13) 0.999 (7) 0.999 (7) 0.982 (6) 0.992 (9) 0.995 (9) 0.874 (13) 0.999 (10) 1.000 (12) 0.991 (5) 1.000 (11) 0.999 (4) 0.774 (4) 0.999 (5) 0.907 (8) 0.992 (11) 0.984 (6) 0.933 (12) 0.846 (14)
1.000 (2) 1.000 (6) 0.936 (2) 0.931 (15) 1.000 (5) 1.000 (7) 1.000 (7) 0.979 (1) 1.000 (1) 0.983 (2) 0.857 (2) 0.996 (1) 1.000 (12) 0.946 (1) 0.965 (1) 0.999 (2) 1.000 (13) 1.000 (2) 0.993 (1) 1.000 (1) 0.957 (2) 0.996 (3) 0.986 (2) 0.999 (2) 1.000 (3) 1.000 (2) 0.993 (5) 0.989 (2) 0.782 (9) 1.000 (3) 1.000 (3) 0.999 (1) 1.000 (2) 1.000 (1) 0.963 (1) 1.000 (7) 1.000 (11) 0.992 (2) 1.000 (10) 1.000 (2) 0.790 (3) 0.999 (2) 0.944 (2) 0.999 (1) 0.992 (1) 0.970 (2) 0.948 (1)
1.000 (1) 1.000 (2) 0.937 (1) 0.950 (9) 1.000 (2) 1.000 (3) 1.000 (4) 0.971 (3) 1.000 (2) 0.985 (1) 0.913 (1) 0.996 (2) 1.000 (7) 0.945 (2) 0.962 (2) 1.000 (1) 1.000 (8) 1.000 (1) 0.993 (2) 1.000 (2) 0.964 (1) 0.997 (1) 0.988 (1) 0.999 (1) 1.000 (2) 1.000 (1) 1.000 (2) 0.989 (1) 0.831 (3) 1.000 (2) 1.000 (1) 0.999 (2) 1.000 (1) 1.000 (2) 0.955 (3) 1.000 (3) 1.000 (5) 0.991 (4) 1.000 (5) 1.000 (1) 0.797 (2) 1.000 (1) 0.946 (1) 0.998 (2) 0.991 (2) 0.970 (1) 0.946 (3)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.985 (6) 0.977 (6) 0.954 (7) 0.996 (6) 0.751 (9)
0.977 (13) 0.973 (12) 0.958 (4) 0.998 (3) 0.743 (13)
0.985 (4) 0.982 (3) 0.949 (11) 0.996 (7) 0.748 (11)
0.880 (16) 0.892 (16) 0.726 (18) 0.798 (17) 0.522 (17)
0.759 (18) 0.844 (18) 0.741 (17) 0.852 (16) 0.518 (18)
0.984 (8) 0.977 (5) 0.951 (8) 0.983 (13) 0.785 (4)
0.979 (10) 0.979 (4) 0.955 (6) 0.993 (9) 0.797 (3)
0.978 (11) 0.972 (13) 0.950 (9) 0.977 (14) 0.773 (5)
0.983 (9) 0.975 (10) 0.960 (3) 0.996 (5) 0.757 (6)
0.978 (12) 0.974 (11) 0.947 (13) 0.988 (12) 0.745 (12)
0.826 (17) 0.878 (17) 0.794 (16) 0.775 (18) 0.548 (16)
0.937 (15) 0.939 (15) 0.857 (15) 0.975 (15) 0.572 (15)
0.984 (7) 0.976 (9) 0.950 (10) 0.995 (8) 0.748 (10)
0.985 (5) 0.977 (7) 0.949 (12) 0.996 (4) 0.752 (8)
0.988 (3) 0.976 (8) 0.929 (14) 0.989 (11) 0.682 (14)
0.969 (14) 0.959 (14) 0.962 (1) 0.992 (10) 0.754 (7)
0.991 (1) 0.986 (1) 0.961 (2) 0.999 (1) 0.823 (1)
0.989 (2) 0.985 (2) 0.957 (5) 0.999 (2) 0.802 (2)
20news agnews amazon imdb yelp
0.923 (9) 0.974 (7) 0.915 (11) 0.929 (11) 0.965 (10)
0.905 (13) 0.960 (13) 0.898 (12) 0.904 (12) 0.948 (12)
0.920 (10) 0.978 (3) 0.919 (9) 0.944 (5) 0.974 (7)
0.570 (18) 0.661 (18) 0.522 (16) 0.483 (18) 0.573 (18)
0.709 (16) 0.813 (16) 0.520 (17) 0.506 (17) 0.573 (17)
0.944 (2) 0.975 (4) 0.935 (4) 0.949 (3) 0.977 (3)
0.924 (8) 0.973 (10) 0.936 (3) 0.946 (4) 0.977 (4)
0.939 (5) 0.973 (8) 0.930 (7) 0.943 (6) 0.976 (5)
0.908 (12) 0.974 (6) 0.935 (5) 0.935 (8) 0.975 (6)
0.936 (6) 0.971 (11) 0.921 (8) 0.933 (9) 0.972 (8)
0.679 (17) 0.753 (17) 0.495 (18) 0.531 (16) 0.610 (15)
0.770 (15) 0.876 (15) 0.545 (15) 0.543 (15) 0.595 (16)
0.919 (11) 0.973 (9) 0.918 (10) 0.929 (10) 0.964 (11)
0.929 (7) 0.974 (5) 0.930 (6) 0.937 (7) 0.970 (9)
0.940 (4) 0.921 (14) 0.800 (14) 0.880 (14) 0.921 (14)
0.899 (14) 0.962 (12) 0.886 (13) 0.892 (13) 0.940 (13)
0.949 (1) 0.982 (1) 0.951 (1) 0.960 (1) 0.982 (1)
0.943 (3) 0.979 (2) 0.943 (2) 0.954 (2) 0.979 (2)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 47: The average AUCPR performance on individual datasets under the 𝑁𝑙𝑎 = 1 setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.467 (7) 0.199 (13) 0.407 (14) 0.414 (9) 0.393 (7) 0.406 (9) 0.837 (10) 0.386 (7) 0.543 (15) 0.305 (10) 0.043 (11) 0.209 (5) 0.900 (4) 0.463 (15) 0.297 (9) 0.243 (15) 0.607 (9) 0.161 (13) 0.442 (9) 0.300 (10) 0.438 (13) 0.266 (14) 0.479 (16) 0.734 (8) 0.417 (17) 0.326 (12) 0.504 (6) 0.451 (6) 0.021 (11) 0.315 (13) 0.567 (8) 0.187 (14) 0.184 (16) 0.403 (7) 0.078 (9) 0.574 (12) 0.792 (8) 0.092 (8) 0.611 (10) 0.279 (12) 0.411 (4) 0.765 (12) 0.266 (3) 0.352 (11) 0.341 (13) 0.089 (8) 0.113 (10)
0.211 (13) 0.324 (12) 0.496 (4) 0.213 (17) 0.233 (13) 0.327 (15) 0.757 (11) 0.377 (8) 0.709 (9) 0.309 (8) 0.065 (2) 0.206 (6) 0.390 (16) 0.611 (3) 0.291 (10) 0.244 (14) 0.373 (17) 0.134 (14) 0.484 (7) 0.286 (11) 0.444 (12) 0.212 (15) 0.507 (14) 0.399 (14) 0.687 (13) 0.455 (10) 0.582 (4) 0.395 (13) 0.022 (7) 0.175 (18) 0.409 (11) 0.157 (15) 0.323 (9) 0.144 (17) 0.161 (3) 0.572 (13) 0.407 (15) 0.064 (14) 0.150 (16) 0.259 (15) 0.327 (15) 0.811 (9) 0.166 (14) 0.294 (15) 0.328 (15) 0.030 (17) 0.082 (14)
0.269 (11) 0.512 (7) 0.439 (9) 0.269 (15) 0.307 (11) 0.436 (6) 0.933 (6) 0.334 (11) 0.515 (17) 0.212 (16) 0.045 (8) 0.114 (16) 0.662 (10) 0.570 (8) 0.428 (4) 0.273 (12) 0.490 (12) 0.177 (12) 0.441 (10) 0.164 (16) 0.431 (14) 0.369 (7) 0.432 (17) 0.324 (16) 0.779 (11) 0.555 (7) 0.440 (11) 0.502 (2) 0.021 (9) 0.401 (8) 0.665 (5) 0.264 (4) 0.190 (15) 0.223 (16) 0.036 (18) 0.349 (16) 0.375 (16) 0.095 (6) 0.366 (13) 0.286 (11) 0.384 (8) 0.606 (14) 0.181 (11) 0.215 (16) 0.305 (17) 0.071 (11) 0.130 (9)
0.091 (17) 0.140 (14) 0.436 (10) 0.400 (10) 0.073 (18) 0.351 (13) 0.280 (17) 0.511 (3) 0.860 (2) 0.163 (18) 0.054 (4) 0.153 (10) 0.963 (1) 0.440 (17) 0.054 (17) 0.350 (6) 0.265 (18) 0.037 (17) 0.513 (6) 0.234 (12) 0.375 (17) 0.116 (17) 0.584 (8) 0.700 (9) 0.860 (8) 0.197 (18) 0.216 (15) 0.376 (16) 0.024 (6) 0.276 (15) 0.325 (14) 0.125 (16) 0.201 (14) 0.096 (18) 0.038 (17) 0.841 (3) 0.457 (14) 0.062 (15) 0.168 (14) 0.203 (18) 0.330 (13) 0.897 (5) 0.235 (6) 0.500 (6) 0.419 (7) 0.093 (7) 0.065 (18)
0.009 (18) 0.080 (17) 0.492 (5) 0.556 (1) 0.171 (17) 0.241 (17) 0.450 (15) 0.440 (4) 0.859 (3) 0.204 (17) 0.040 (13) 0.164 (8) 0.621 (11) 0.480 (12) 0.116 (15) 0.203 (17) 0.718 (7) 0.026 (18) 0.359 (15) 0.174 (15) 0.457 (11) 0.301 (12) 0.642 (4) 0.470 (12) 0.329 (18) 0.210 (17) 0.106 (17) 0.402 (12) 0.020 (13) 0.284 (14) 0.496 (9) 0.098 (17) 0.308 (10) 0.309 (12) 0.043 (16) 0.418 (15) 0.634 (12) 0.036 (18) 0.132 (17) 0.231 (16) 0.314 (16) 0.845 (8) 0.223 (7) 0.358 (10) 0.351 (11) 0.095 (6) 0.075 (16)
0.642 (4) 0.872 (1) 0.425 (13) 0.530 (4) 0.209 (14) 0.469 (5) 1.000 (1) 0.363 (9) 0.629 (10) 0.305 (9) 0.041 (12) 0.125 (14) 0.696 (9) 0.514 (10) 0.607 (1) 0.342 (7) 0.984 (1) 0.851 (1) 0.403 (12) 0.622 (1) 0.501 (7) 0.405 (2) 0.658 (2) 0.894 (4) 0.956 (2) 0.679 (3) 0.600 (1) 0.444 (7) 0.027 (3) 0.500 (5) 0.809 (3) 0.242 (9) 0.544 (3) 0.433 (6) 0.072 (11) 0.789 (6) 0.980 (1) 0.083 (12) 0.999 (2) 0.380 (1) 0.410 (5) 0.804 (10) 0.189 (10) 0.546 (3) 0.459 (3) 0.108 (3) 0.139 (5)
0.653 (3) 0.822 (3) 0.443 (8) 0.537 (3) 0.611 (2) 0.360 (12) 1.000 (2) 0.335 (10) 0.523 (16) 0.348 (4) 0.048 (6) 0.172 (7) 0.529 (14) 0.450 (16) 0.430 (3) 0.302 (10) 0.688 (8) 0.677 (3) 0.391 (14) 0.574 (3) 0.466 (9) 0.392 (5) 0.533 (11) 0.911 (1) 0.954 (3) 0.699 (2) 0.501 (7) 0.387 (14) 0.018 (16) 0.626 (2) 0.457 (10) 0.276 (2) 0.470 (5) 0.455 (5) 0.101 (6) 0.529 (14) 0.936 (3) 0.093 (7) 0.997 (3) 0.314 (9) 0.386 (7) 0.576 (16) 0.155 (16) 0.337 (12) 0.348 (12) 0.086 (9) 0.138 (6)
0.661 (2) 0.850 (2) 0.430 (12) 0.530 (5) 0.311 (10) 0.423 (7) 1.000 (3) 0.320 (13) 0.594 (13) 0.352 (3) 0.039 (14) 0.127 (12) 0.596 (12) 0.517 (9) 0.517 (2) 0.330 (8) 0.918 (3) 0.834 (2) 0.397 (13) 0.605 (2) 0.505 (5) 0.392 (4) 0.613 (7) 0.903 (3) 0.958 (1) 0.665 (4) 0.600 (3) 0.434 (8) 0.029 (2) 0.514 (4) 0.828 (2) 0.259 (5) 0.542 (4) 0.456 (4) 0.079 (8) 0.737 (9) 0.979 (2) 0.085 (11) 1.000 (1) 0.365 (3) 0.411 (3) 0.750 (13) 0.177 (12) 0.522 (4) 0.462 (2) 0.102 (4) 0.151 (3)
0.132 (14) 0.496 (8) 0.397 (16) 0.235 (16) 0.567 (3) 0.328 (14) 0.918 (8) 0.190 (16) 0.473 (18) 0.353 (2) 0.035 (17) 0.155 (9) 0.404 (15) 0.470 (14) 0.183 (12) 0.183 (18) 0.382 (16) 0.203 (11) 0.336 (16) 0.353 (7) 0.420 (15) 0.496 (1) 0.525 (12) 0.094 (18) 0.621 (14) 0.784 (1) 0.500 (8) 0.375 (17) 0.021 (10) 0.220 (16) 0.235 (18) 0.333 (1) 0.414 (7) 0.389 (8) 0.092 (7) 0.333 (17) 0.650 (11) 0.496 (1) 0.427 (12) 0.350 (5) 0.405 (6) 0.485 (18) 0.137 (17) 0.181 (18) 0.293 (18) 0.077 (10) 0.133 (8)
0.727 (1) 0.645 (4) 0.434 (11) 0.539 (2) 0.473 (5) 0.410 (8) 0.554 (14) 0.312 (15) 0.822 (5) 0.310 (7) 0.046 (7) 0.150 (11) 0.953 (2) 0.600 (6) 0.397 (5) 0.388 (2) 0.849 (5) 0.303 (8) 0.534 (4) 0.408 (5) 0.462 (10) 0.363 (8) 0.662 (1) 0.904 (2) 0.950 (4) 0.575 (6) 0.304 (13) 0.429 (9) 0.018 (15) 0.562 (3) 0.839 (1) 0.216 (11) 0.549 (2) 0.471 (3) 0.065 (13) 0.792 (5) 0.466 (13) 0.068 (13) 0.902 (5) 0.355 (4) 0.373 (10) 0.785 (11) 0.168 (13) 0.682 (1) 0.525 (1) 0.066 (12) 0.092 (13)
0.093 (16) 0.084 (16) 0.499 (3) 0.321 (14) 0.263 (12) 0.218 (18) 0.170 (18) 0.141 (18) 0.853 (4) 0.295 (12) 0.045 (9) 0.393 (1) 0.145 (18) 0.603 (5) 0.031 (18) 0.352 (5) 0.821 (6) 0.051 (15) 0.224 (18) 0.206 (13) 0.394 (16) 0.101 (18) 0.557 (10) 0.264 (17) 0.589 (15) 0.324 (13) 0.321 (12) 0.380 (15) 0.025 (5) 0.201 (17) 0.282 (16) 0.202 (13) 0.426 (6) 0.313 (11) 0.232 (1) 0.200 (18) 0.090 (18) 0.049 (16) 0.097 (18) 0.270 (13) 0.310 (17) 0.590 (15) 0.245 (5) 0.333 (13) 0.332 (14) 0.030 (16) 0.069 (17)
0.095 (15) 0.100 (15) 0.545 (1) 0.491 (6) 0.190 (16) 0.249 (16) 0.293 (16) 0.315 (14) 0.912 (1) 0.221 (15) 0.049 (5) 0.225 (4) 0.256 (17) 0.651 (1) 0.112 (16) 0.376 (3) 0.408 (15) 0.039 (16) 0.522 (5) 0.112 (18) 0.544 (3) 0.122 (16) 0.495 (15) 0.433 (13) 0.912 (5) 0.284 (15) 0.480 (10) 0.419 (11) 0.026 (4) 0.324 (12) 0.264 (17) 0.092 (18) 0.366 (8) 0.239 (15) 0.077 (10) 0.593 (11) 0.323 (17) 0.037 (17) 0.153 (15) 0.219 (17) 0.296 (18) 0.929 (3) 0.248 (4) 0.475 (7) 0.433 (6) 0.035 (15) 0.102 (12)
0.533 (5) 0.408 (11) 0.399 (15) 0.419 (8) 0.357 (9) 0.393 (11) 0.922 (7) 0.421 (5) 0.624 (11) 0.316 (6) 0.034 (18) 0.103 (17) 0.790 (7) 0.477 (13) 0.387 (6) 0.302 (11) 0.592 (11) 0.424 (5) 0.453 (8) 0.331 (9) 0.509 (4) 0.388 (6) 0.508 (13) 0.814 (7) 0.873 (6) 0.378 (11) 0.198 (16) 0.527 (1) 0.018 (18) 0.332 (11) 0.746 (4) 0.248 (8) 0.214 (12) 0.275 (14) 0.063 (14) 0.848 (2) 0.903 (4) 0.089 (10) 0.711 (7) 0.286 (10) 0.353 (12) 0.904 (4) 0.294 (2) 0.331 (14) 0.385 (10) 0.113 (2) 0.226 (1)
0.409 (9) 0.442 (10) 0.484 (7) 0.393 (13) 0.380 (8) 0.481 (4) 0.860 (9) 0.622 (1) 0.806 (6) 0.297 (11) 0.037 (16) 0.123 (15) 0.944 (3) 0.585 (7) 0.308 (8) 0.362 (4) 0.855 (4) 0.381 (7) 0.659 (3) 0.349 (8) 0.575 (1) 0.319 (9) 0.624 (6) 0.872 (5) 0.826 (10) 0.320 (14) 0.269 (14) 0.483 (3) 0.021 (12) 0.399 (9) 0.642 (6) 0.219 (10) 0.101 (18) 0.383 (9) 0.069 (12) 0.778 (8) 0.813 (7) 0.112 (5) 0.535 (11) 0.270 (14) 0.329 (14) 0.943 (2) 0.375 (1) 0.643 (2) 0.453 (5) 0.114 (1) 0.202 (2)
0.461 (8) 0.578 (6) 0.382 (17) 0.393 (12) 0.636 (1) 0.492 (2) 1.000 (4) 0.184 (17) 0.563 (14) 0.405 (1) 0.037 (15) 0.126 (13) 0.549 (13) 0.393 (18) 0.173 (13) 0.249 (13) 0.600 (10) 0.404 (6) 0.317 (17) 0.378 (6) 0.360 (18) 0.270 (13) 0.629 (5) 0.679 (10) 0.855 (9) 0.612 (5) 0.500 (9) 0.298 (18) 0.022 (8) 0.671 (1) 0.286 (15) 0.249 (7) 0.269 (11) 0.279 (13) 0.134 (4) 0.655 (10) 0.859 (5) 0.128 (4) 0.789 (6) 0.373 (2) 0.430 (2) 0.546 (17) 0.122 (18) 0.189 (17) 0.311 (16) 0.053 (14) 0.082 (15)
0.231 (12) 0.460 (9) 0.368 (18) 0.194 (18) 0.196 (15) 0.518 (1) 0.575 (13) 0.410 (6) 0.613 (12) 0.222 (14) 0.044 (10) 0.083 (18) 0.873 (5) 0.502 (11) 0.210 (11) 0.211 (16) 0.480 (13) 0.254 (10) 0.422 (11) 0.125 (17) 0.504 (6) 0.305 (11) 0.422 (18) 0.360 (15) 0.716 (12) 0.484 (9) 0.074 (18) 0.424 (10) 0.030 (1) 0.424 (6) 0.377 (12) 0.207 (12) 0.155 (17) 0.318 (10) 0.046 (15) 0.896 (1) 0.711 (10) 0.091 (9) 0.924 (4) 0.341 (6) 0.430 (1) 0.989 (1) 0.159 (15) 0.438 (8) 0.400 (9) 0.057 (13) 0.133 (7)
0.338 (10) 0.062 (18) 0.526 (2) 0.400 (11) 0.534 (4) 0.406 (10) 0.720 (12) 0.325 (12) 0.790 (8) 0.285 (13) 0.062 (3) 0.360 (3) 0.783 (8) 0.608 (4) 0.161 (14) 0.323 (9) 0.414 (14) 0.269 (9) 0.673 (2) 0.187 (14) 0.549 (2) 0.311 (10) 0.583 (9) 0.630 (11) 0.528 (16) 0.281 (16) 0.600 (2) 0.470 (4) 0.018 (17) 0.405 (7) 0.365 (13) 0.267 (3) 0.211 (13) 0.630 (1) 0.103 (5) 0.809 (4) 0.834 (6) 0.328 (3) 0.688 (8) 0.324 (8) 0.367 (11) 0.854 (7) 0.194 (9) 0.395 (9) 0.416 (8) 0.016 (18) 0.107 (11)
0.490 (6) 0.602 (5) 0.486 (6) 0.473 (7) 0.399 (6) 0.484 (3) 0.933 (5) 0.602 (2) 0.790 (7) 0.331 (5) 0.090 (1) 0.378 (2) 0.827 (6) 0.643 (2) 0.352 (7) 0.630 (1) 0.977 (2) 0.666 (4) 0.683 (1) 0.473 (4) 0.479 (8) 0.402 (3) 0.642 (3) 0.815 (6) 0.868 (7) 0.520 (8) 0.518 (5) 0.468 (5) 0.019 (14) 0.390 (10) 0.605 (7) 0.250 (6) 0.585 (1) 0.585 (2) 0.193 (2) 0.789 (7) 0.787 (9) 0.360 (2) 0.623 (9) 0.326 (7) 0.382 (9) 0.885 (6) 0.202 (8) 0.518 (5) 0.457 (4) 0.102 (5) 0.149 (4)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.268 (12) 0.286 (12) 0.315 (8) 0.472 (9) 0.062 (11)
0.245 (13) 0.187 (16) 0.168 (16) 0.376 (18) 0.060 (15)
0.151 (17) 0.165 (18) 0.148 (18) 0.456 (10) 0.062 (12)
0.328 (7) 0.321 (9) 0.210 (14) 0.581 (3) 0.057 (17)
0.113 (18) 0.278 (13) 0.153 (17) 0.687 (1) 0.056 (18)
0.287 (10) 0.297 (11) 0.455 (5) 0.430 (12) 0.074 (1)
0.289 (9) 0.305 (10) 0.462 (4) 0.400 (15) 0.066 (4)
0.347 (6) 0.353 (7) 0.485 (3) 0.392 (17) 0.064 (8)
0.157 (16) 0.165 (17) 0.219 (13) 0.393 (16) 0.070 (2)
0.421 (3) 0.427 (4) 0.495 (1) 0.576 (4) 0.064 (6)
0.279 (11) 0.375 (5) 0.366 (7) 0.424 (13) 0.061 (13)
0.393 (5) 0.433 (3) 0.179 (15) 0.511 (7) 0.067 (3)
0.194 (15) 0.197 (15) 0.223 (12) 0.402 (14) 0.060 (14)
0.431 (2) 0.460 (2) 0.306 (9) 0.678 (2) 0.062 (10)
0.324 (8) 0.338 (8) 0.436 (6) 0.450 (11) 0.065 (5)
0.240 (14) 0.220 (14) 0.265 (10) 0.527 (6) 0.059 (16)
0.396 (4) 0.356 (6) 0.246 (11) 0.540 (5) 0.064 (7)
0.486 (1) 0.535 (1) 0.487 (2) 0.473 (8) 0.063 (9)
20news agnews amazon imdb yelp
0.160 (7) 0.133 (12) 0.071 (11) 0.069 (9) 0.086 (11)
0.099 (17) 0.094 (17) 0.055 (15) 0.056 (14) 0.058 (16)
0.134 (13) 0.094 (16) 0.075 (7) 0.061 (12) 0.089 (9)
0.058 (18) 0.088 (18) 0.052 (17) 0.048 (18) 0.046 (18)
0.113 (16) 0.147 (9) 0.050 (18) 0.051 (17) 0.050 (17)
0.198 (2) 0.174 (6) 0.101 (1) 0.090 (2) 0.142 (4)
0.160 (8) 0.166 (8) 0.075 (8) 0.076 (7) 0.120 (7)
0.209 (1) 0.197 (5) 0.096 (3) 0.090 (3) 0.163 (2)
0.157 (10) 0.109 (13) 0.096 (2) 0.085 (4) 0.096 (8)
0.192 (4) 0.205 (4) 0.089 (4) 0.084 (6) 0.147 (3)
0.152 (11) 0.227 (3) 0.074 (9) 0.066 (10) 0.070 (13)
0.148 (12) 0.295 (1) 0.054 (16) 0.055 (16) 0.060 (15)
0.129 (15) 0.095 (15) 0.072 (10) 0.066 (11) 0.088 (10)
0.134 (14) 0.143 (10) 0.060 (14) 0.057 (13) 0.074 (12)
0.182 (5) 0.167 (7) 0.080 (5) 0.094 (1) 0.180 (1)
0.171 (6) 0.101 (14) 0.066 (12) 0.073 (8) 0.124 (6)
0.158 (9) 0.137 (11) 0.062 (13) 0.056 (15) 0.061 (14)
0.194 (3) 0.239 (2) 0.080 (6) 0.085 (5) 0.140 (5)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 48: The average of AUCPR performance on individual datasets under 𝑁𝑙𝑎 = 3 setting Dataset
XGBOD
DeepSAD
REPEN
AABiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
DualMGAN
SOEL-NTL
AnoDDAE
XGB
CatB
TabMCls
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.760 (8) 0.454 (14) 0.442 (12) 0.479 (13) 0.767 (4) 0.498 (12) 0.949 (9) 0.381 (10) 0.652 (12) 0.445 (4) 0.058 (4) 0.422 (3) 0.788 (11) 0.519 (11) 0.321 (13) 0.452 (5) 0.824 (11) 0.561 (9) 0.559 (6) 0.582 (9) 0.512 (4) 0.547 (1) 0.545 (8) 0.786 (10) 0.852 (13) 0.398 (14) 0.701 (11) 0.499 (5) 0.038 (11) 0.611 (10) 0.681 (8) 0.263 (11) 0.424 (11) 0.537 (7) 0.180 (10) 0.643 (13) 0.863 (9) 0.195 (8) 0.833 (10) 0.343 (11) 0.430 (4) 0.844 (10) 0.279 (5) 0.567 (8) 0.501 (8) 0.121 (7) 0.218 (3)
0.621 (11) 0.567 (13) 0.503 (2) 0.440 (16) 0.301 (13) 0.398 (14) 0.859 (13) 0.389 (8) 0.713 (10) 0.363 (9) 0.064 (3) 0.214 (7) 0.544 (15) 0.605 (5) 0.336 (11) 0.301 (14) 0.609 (16) 0.313 (12) 0.548 (7) 0.544 (11) 0.444 (12) 0.237 (14) 0.500 (14) 0.827 (8) 0.924 (10) 0.650 (8) 0.803 (3) 0.404 (14) 0.024 (17) 0.232 (17) 0.506 (11) 0.185 (15) 0.499 (8) 0.192 (17) 0.218 (8) 0.663 (11) 0.508 (15) 0.092 (10) 0.268 (14) 0.280 (15) 0.339 (14) 0.822 (11) 0.174 (13) 0.359 (16) 0.376 (15) 0.050 (16) 0.083 (16)
0.725 (10) 0.797 (6) 0.457 (9) 0.612 (8) 0.508 (11) 0.564 (10) 0.933 (11) 0.390 (7) 0.765 (7) 0.248 (14) 0.047 (9) 0.187 (10) 0.719 (12) 0.522 (10) 0.490 (2) 0.307 (13) 0.812 (12) 0.281 (14) 0.577 (4) 0.404 (13) 0.506 (5) 0.321 (11) 0.468 (16) 0.585 (13) 0.956 (6) 0.538 (12) 0.800 (8) 0.425 (11) 0.059 (1) 0.639 (8) 0.729 (6) 0.340 (6) 0.551 (7) 0.409 (11) 0.112 (15) 0.523 (15) 0.699 (12) 0.084 (12) 0.751 (12) 0.363 (9) 0.407 (7) 0.744 (14) 0.253 (6) 0.439 (11) 0.376 (16) 0.086 (13) 0.154 (8)
0.192 (15) 0.088 (17) 0.364 (17) 0.370 (18) 0.104 (18) 0.305 (15) 0.591 (15) 0.374 (11) 0.746 (8) 0.168 (18) 0.052 (5) 0.074 (18) 0.498 (16) 0.372 (18) 0.026 (17) 0.432 (7) 0.083 (18) 0.074 (15) 0.339 (15) 0.086 (18) 0.376 (18) 0.128 (16) 0.579 (6) 0.628 (12) 0.641 (16) 0.236 (17) 0.305 (16) 0.371 (17) 0.032 (12) 0.198 (18) 0.212 (18) 0.155 (16) 0.024 (18) 0.052 (18) 0.042 (18) 0.723 (8) 0.397 (16) 0.052 (15) 0.072 (18) 0.221 (17) 0.314 (16) 0.774 (13) 0.249 (7) 0.472 (10) 0.592 (5) 0.090 (12) 0.054 (18)
0.010 (18) 0.103 (15) 0.492 (6) 0.560 (10) 0.172 (16) 0.242 (18) 0.452 (16) 0.441 (5) 0.860 (2) 0.204 (17) 0.040 (15) 0.165 (13) 0.627 (13) 0.499 (13) 0.119 (15) 0.204 (18) 0.635 (14) 0.026 (18) 0.411 (14) 0.174 (15) 0.456 (11) 0.354 (8) 0.642 (2) 0.477 (14) 0.329 (18) 0.201 (18) 0.130 (18) 0.401 (15) 0.020 (18) 0.286 (15) 0.498 (12) 0.098 (17) 0.314 (15) 0.316 (15) 0.043 (17) 0.423 (16) 0.634 (13) 0.049 (16) 0.134 (16) 0.231 (16) 0.313 (17) 0.848 (8) 0.245 (10) 0.359 (15) 0.351 (18) 0.056 (15) 0.084 (15)
0.888 (4) 0.972 (1) 0.449 (10) 0.674 (4) 0.207 (15) 0.640 (3) 1.000 (1) 0.396 (6) 0.559 (14) 0.376 (8) 0.043 (14) 0.187 (11) 0.932 (5) 0.612 (4) 0.494 (1) 0.446 (6) 1.000 (1) 0.986 (1) 0.473 (11) 0.855 (3) 0.479 (7) 0.256 (13) 0.543 (9) 0.896 (4) 0.971 (2) 0.676 (6) 0.800 (7) 0.506 (4) 0.054 (3) 0.658 (5) 0.889 (2) 0.250 (13) 0.554 (6) 0.666 (5) 0.156 (13) 0.861 (1) 1.000 (1) 0.081 (13) 1.000 (1) 0.402 (4) 0.427 (5) 0.859 (7) 0.211 (11) 0.711 (3) 0.619 (1) 0.184 (1) 0.155 (7)
0.880 (5) 0.945 (3) 0.441 (13) 0.697 (2) 0.730 (7) 0.579 (8) 1.000 (2) 0.367 (12) 0.513 (16) 0.382 (7) 0.050 (7) 0.196 (8) 0.842 (10) 0.464 (16) 0.463 (3) 0.378 (10) 0.996 (4) 0.981 (3) 0.450 (13) 0.875 (1) 0.438 (15) 0.347 (9) 0.503 (13) 0.910 (3) 0.968 (4) 0.796 (2) 0.801 (5) 0.469 (9) 0.051 (6) 0.657 (6) 0.745 (5) 0.319 (9) 0.723 (3) 0.654 (6) 0.262 (4) 0.719 (9) 0.987 (3) 0.184 (9) 1.000 (2) 0.399 (5) 0.394 (9) 0.649 (16) 0.159 (16) 0.568 (7) 0.453 (11) 0.172 (2) 0.143 (9)
0.909 (1) 0.966 (2) 0.448 (11) 0.644 (6) 0.406 (12) 0.598 (7) 1.000 (3) 0.366 (13) 0.552 (15) 0.390 (6) 0.045 (12) 0.188 (9) 0.937 (3) 0.565 (8) 0.457 (4) 0.421 (8) 1.000 (2) 0.981 (2) 0.477 (10) 0.860 (2) 0.474 (8) 0.313 (12) 0.574 (7) 0.885 (5) 0.970 (3) 0.674 (7) 0.800 (9) 0.488 (7) 0.058 (2) 0.717 (1) 0.898 (1) 0.272 (10) 0.583 (5) 0.727 (2) 0.235 (7) 0.824 (5) 1.000 (2) 0.086 (11) 1.000 (3) 0.405 (3) 0.434 (3) 0.785 (12) 0.209 (12) 0.675 (4) 0.565 (6) 0.170 (3) 0.156 (6)
0.435 (13) 0.662 (11) 0.400 (16) 0.479 (14) 0.737 (6) 0.404 (13) 0.918 (12) 0.224 (16) 0.426 (18) 0.353 (12) 0.036 (18) 0.221 (6) 0.592 (14) 0.452 (17) 0.314 (14) 0.267 (16) 0.648 (13) 0.366 (11) 0.307 (17) 0.440 (12) 0.403 (17) 0.366 (7) 0.422 (17) 0.448 (16) 0.839 (14) 0.795 (3) 0.531 (13) 0.404 (13) 0.047 (9) 0.377 (13) 0.363 (15) 0.376 (1) 0.431 (10) 0.483 (8) 0.235 (6) 0.377 (17) 0.540 (14) 0.646 (2) 0.677 (13) 0.398 (6) 0.421 (6) 0.457 (18) 0.152 (17) 0.248 (18) 0.386 (14) 0.092 (11) 0.125 (10)
0.902 (2) 0.715 (7) 0.472 (8) 0.664 (5) 0.536 (10) 0.519 (11) 0.628 (14) 0.309 (15) 0.793 (6) 0.354 (11) 0.051 (6) 0.159 (15) 0.962 (1) 0.601 (6) 0.347 (10) 0.453 (4) 0.933 (8) 0.807 (5) 0.573 (5) 0.710 (6) 0.441 (13) 0.342 (10) 0.694 (1) 0.933 (1) 0.973 (1) 0.609 (9) 0.305 (17) 0.487 (8) 0.048 (7) 0.624 (9) 0.873 (3) 0.261 (12) 0.789 (2) 0.689 (4) 0.200 (9) 0.840 (2) 0.758 (11) 0.066 (14) 0.933 (6) 0.351 (10) 0.386 (11) 0.870 (5) 0.168 (14) 0.761 (1) 0.618 (2) 0.105 (9) 0.096 (14)
0.099 (17) 0.054 (18) 0.496 (5) 0.437 (17) 0.258 (14) 0.293 (16) 0.162 (18) 0.145 (18) 0.857 (3) 0.285 (13) 0.045 (13) 0.346 (4) 0.186 (18) 0.589 (7) 0.020 (18) 0.352 (12) 0.857 (10) 0.042 (16) 0.329 (16) 0.124 (16) 0.459 (10) 0.102 (18) 0.596 (3) 0.206 (18) 0.562 (17) 0.309 (15) 0.600 (12) 0.381 (16) 0.027 (15) 0.245 (16) 0.349 (16) 0.197 (14) 0.401 (13) 0.399 (12) 0.252 (5) 0.229 (18) 0.080 (18) 0.042 (17) 0.125 (17) 0.293 (14) 0.317 (15) 0.600 (17) 0.247 (8) 0.373 (14) 0.356 (17) 0.027 (18) 0.068 (17)
0.104 (16) 0.094 (16) 0.534 (1) 0.492 (12) 0.167 (17) 0.255 (17) 0.294 (17) 0.311 (14) 0.916 (1) 0.219 (16) 0.048 (8) 0.228 (5) 0.287 (17) 0.652 (3) 0.115 (16) 0.380 (9) 0.415 (17) 0.041 (17) 0.534 (8) 0.109 (17) 0.550 (1) 0.127 (17) 0.505 (12) 0.440 (17) 0.911 (12) 0.272 (16) 0.457 (15) 0.420 (12) 0.027 (16) 0.312 (14) 0.291 (17) 0.096 (18) 0.373 (14) 0.271 (16) 0.078 (16) 0.618 (14) 0.323 (17) 0.037 (18) 0.158 (15) 0.219 (18) 0.292 (18) 0.925 (2) 0.245 (9) 0.476 (9) 0.433 (13) 0.033 (17) 0.101 (12)
0.753 (9) 0.622 (12) 0.418 (14) 0.518 (11) 0.627 (9) 0.579 (9) 0.949 (8) 0.506 (3) 0.670 (11) 0.362 (10) 0.040 (16) 0.142 (17) 0.843 (9) 0.506 (12) 0.385 (9) 0.360 (11) 0.954 (7) 0.547 (10) 0.529 (9) 0.548 (10) 0.472 (9) 0.460 (5) 0.495 (15) 0.780 (11) 0.932 (8) 0.563 (11) 0.471 (14) 0.490 (6) 0.028 (14) 0.572 (12) 0.668 (9) 0.322 (8) 0.411 (12) 0.317 (14) 0.146 (14) 0.651 (12) 0.841 (10) 0.279 (5) 0.811 (11) 0.340 (12) 0.372 (13) 0.846 (9) 0.284 (3) 0.418 (13) 0.478 (9) 0.146 (5) 0.274 (1)
0.591 (12) 0.697 (10) 0.500 (4) 0.560 (9) 0.863 (1) 0.649 (2) 0.956 (7) 0.561 (2) 0.817 (4) 0.451 (3) 0.047 (11) 0.181 (12) 0.951 (2) 0.547 (9) 0.420 (7) 0.508 (2) 0.891 (9) 0.735 (7) 0.686 (2) 0.635 (8) 0.550 (2) 0.502 (3) 0.524 (11) 0.844 (6) 0.926 (9) 0.685 (5) 0.900 (1) 0.510 (3) 0.028 (13) 0.674 (4) 0.755 (4) 0.344 (5) 0.313 (16) 0.455 (10) 0.163 (12) 0.828 (3) 0.980 (5) 0.265 (6) 0.880 (7) 0.338 (13) 0.380 (12) 0.937 (1) 0.373 (1) 0.760 (2) 0.611 (3) 0.117 (8) 0.241 (2)
0.822 (6) 0.829 (4) 0.405 (15) 0.624 (7) 0.817 (3) 0.618 (6) 1.000 (4) 0.189 (17) 0.445 (17) 0.517 (1) 0.036 (17) 0.160 (14) 0.877 (8) 0.465 (15) 0.322 (12) 0.297 (15) 0.960 (6) 0.722 (8) 0.300 (18) 0.820 (4) 0.426 (16) 0.232 (15) 0.532 (10) 0.836 (7) 0.916 (11) 0.873 (1) 0.800 (6) 0.370 (18) 0.052 (5) 0.701 (2) 0.470 (13) 0.356 (3) 0.492 (9) 0.478 (9) 0.293 (3) 0.785 (7) 0.951 (7) 0.196 (7) 0.972 (4) 0.421 (1) 0.449 (2) 0.719 (15) 0.145 (18) 0.351 (17) 0.474 (10) 0.132 (6) 0.104 (11)
0.371 (14) 0.702 (9) 0.359 (18) 0.448 (15) 0.667 (8) 0.634 (4) 0.975 (6) 0.487 (4) 0.581 (13) 0.223 (15) 0.047 (10) 0.151 (16) 0.933 (4) 0.478 (14) 0.390 (8) 0.251 (17) 0.622 (15) 0.312 (13) 0.462 (12) 0.245 (14) 0.440 (14) 0.375 (6) 0.380 (18) 0.475 (15) 0.710 (15) 0.599 (10) 0.703 (10) 0.433 (10) 0.047 (8) 0.591 (11) 0.397 (14) 0.335 (7) 0.280 (17) 0.390 (13) 0.166 (11) 0.712 (10) 0.876 (8) 0.439 (4) 0.842 (9) 0.376 (7) 0.452 (1) 0.918 (3) 0.167 (15) 0.436 (12) 0.447 (12) 0.103 (10) 0.098 (13)
0.815 (7) 0.711 (8) 0.501 (3) 0.695 (3) 0.838 (2) 0.630 (5) 0.948 (10) 0.389 (9) 0.720 (9) 0.397 (5) 0.073 (2) 0.510 (2) 0.925 (6) 0.663 (1) 0.437 (6) 0.490 (3) 0.973 (5) 0.766 (6) 0.679 (3) 0.679 (7) 0.517 (3) 0.530 (2) 0.588 (5) 0.824 (9) 0.966 (5) 0.427 (13) 0.802 (4) 0.515 (2) 0.052 (4) 0.680 (3) 0.625 (10) 0.375 (2) 0.636 (4) 0.788 (1) 0.335 (2) 0.811 (6) 0.983 (4) 0.566 (3) 0.970 (5) 0.406 (2) 0.393 (10) 0.880 (4) 0.304 (2) 0.616 (6) 0.525 (7) 0.068 (14) 0.159 (5)
0.894 (3) 0.808 (5) 0.489 (7) 0.735 (1) 0.753 (5) 0.666 (1) 0.998 (5) 0.623 (1) 0.809 (5) 0.458 (2) 0.090 (1) 0.533 (1) 0.922 (7) 0.656 (2) 0.440 (5) 0.769 (1) 0.997 (3) 0.909 (4) 0.695 (1) 0.794 (5) 0.496 (6) 0.500 (4) 0.593 (4) 0.922 (2) 0.945 (7) 0.730 (4) 0.900 (2) 0.598 (1) 0.046 (10) 0.656 (7) 0.697 (7) 0.347 (4) 0.792 (1) 0.716 (3) 0.364 (1) 0.825 (4) 0.955 (6) 0.647 (1) 0.850 (8) 0.370 (8) 0.403 (8) 0.864 (6) 0.282 (4) 0.644 (5) 0.602 (4) 0.154 (4) 0.191 (4)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.386 (11) 0.446 (9) 0.464 (7) 0.546 (7) 0.069 (9)
0.273 (13) 0.249 (17) 0.264 (13) 0.421 (17) 0.064 (13)
0.249 (14) 0.253 (16) 0.263 (14) 0.552 (6) 0.069 (11)
0.390 (10) 0.396 (12) 0.205 (16) 0.500 (12) 0.059 (17)
0.113 (18) 0.271 (14) 0.129 (18) 0.687 (2) 0.056 (18)
0.454 (6) 0.513 (4) 0.658 (3) 0.500 (13) 0.079 (2)
0.433 (8) 0.479 (8) 0.659 (2) 0.463 (15) 0.074 (5)
0.497 (4) 0.558 (3) 0.672 (1) 0.478 (14) 0.072 (6)
0.124 (17) 0.193 (18) 0.383 (12) 0.433 (16) 0.081 (1)
0.543 (2) 0.590 (2) 0.654 (4) 0.623 (4) 0.069 (7)
0.204 (16) 0.345 (13) 0.400 (10) 0.404 (18) 0.062 (16)
0.392 (9) 0.436 (10) 0.180 (17) 0.519 (10) 0.066 (12)
0.351 (12) 0.399 (11) 0.427 (9) 0.539 (8) 0.069 (8)
0.467 (5) 0.487 (6) 0.396 (11) 0.702 (1) 0.064 (15)
0.436 (7) 0.482 (7) 0.653 (5) 0.517 (11) 0.077 (3)
0.247 (15) 0.253 (15) 0.242 (15) 0.523 (9) 0.064 (14)
0.515 (3) 0.496 (5) 0.447 (8) 0.630 (3) 0.074 (4)
0.606 (1) 0.668 (1) 0.637 (6) 0.555 (5) 0.069 (10)
20news agnews amazon imdb yelp
0.219 (10) 0.172 (9) 0.085 (10) 0.089 (9) 0.166 (9)
0.126 (16) 0.125 (16) 0.060 (15) 0.063 (15) 0.073 (15)
0.185 (11) 0.130 (15) 0.085 (11) 0.080 (12) 0.097 (14)
0.084 (18) 0.088 (18) 0.051 (18) 0.049 (18) 0.053 (17)
0.114 (17) 0.153 (14) 0.053 (17) 0.050 (17) 0.050 (18)
0.325 (2) 0.325 (4) 0.159 (1) 0.171 (3) 0.394 (2)
0.267 (7) 0.258 (7) 0.103 (7) 0.121 (7) 0.290 (5)
0.319 (3) 0.309 (5) 0.149 (3) 0.172 (2) 0.344 (3)
0.222 (8) 0.171 (10) 0.117 (5) 0.148 (6) 0.226 (7)
0.298 (5) 0.326 (3) 0.127 (4) 0.154 (5) 0.317 (4)
0.156 (14) 0.170 (11) 0.077 (13) 0.090 (8) 0.103 (13)
0.150 (15) 0.294 (6) 0.055 (16) 0.055 (16) 0.061 (16)
0.219 (9) 0.162 (13) 0.086 (9) 0.087 (11) 0.150 (10)
0.184 (12) 0.169 (12) 0.070 (14) 0.067 (13) 0.108 (12)
0.288 (6) 0.330 (2) 0.155 (2) 0.193 (1) 0.402 (1)
0.161 (13) 0.121 (17) 0.093 (8) 0.065 (14) 0.124 (11)
0.301 (4) 0.197 (8) 0.083 (12) 0.087 (10) 0.169 (8)
0.327 (1) 0.361 (1) 0.116 (6) 0.161 (4) 0.260 (6)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 49: The average of AUCPR performance on individual datasets under 𝑁𝑙𝑎 = 5 setting Dataset
XGBOD
DeepSAD
REPEN
AABiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
DualMGAN
SOEL-NTL
AnoDDAE
XGB
CatB
TabMCls
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.826 (9) 0.551 (14) 0.467 (13) 0.634 (11) 0.811 (5) 0.667 (7) 0.967 (7) 0.499 (4) 0.665 (12) 0.457 (4) 0.061 (4) 0.392 (3) 0.803 (12) 0.539 (10) 0.328 (13) 0.512 (7) 0.991 (9) 0.695 (9) 0.560 (6) 0.675 (11) 0.483 (11) 0.667 (1) 0.616 (9) 0.785 (10) 0.942 (10) 0.566 (12) 0.701 (11) 0.527 (6) 0.034 (11) 0.686 (5) 0.774 (9) 0.400 (10) 0.568 (10) 0.677 (8) 0.224 (7) 0.684 (13) 0.941 (10) 0.258 (9) 0.870 (12) 0.428 (9) 0.446 (1) 0.891 (7) 0.326 (4) 0.731 (7) 0.519 (11) 0.138 (7) 0.260 (2)
0.738 (12) 0.753 (9) 0.510 (3) 0.565 (14) 0.342 (13) 0.466 (14) 0.875 (13) 0.396 (11) 0.732 (9) 0.378 (9) 0.062 (3) 0.217 (8) 0.617 (15) 0.606 (6) 0.404 (8) 0.347 (14) 0.802 (12) 0.524 (12) 0.563 (5) 0.696 (10) 0.462 (13) 0.272 (15) 0.576 (12) 0.872 (6) 0.930 (11) 0.685 (7) 0.803 (3) 0.412 (15) 0.024 (15) 0.275 (16) 0.622 (13) 0.213 (14) 0.569 (9) 0.266 (16) 0.257 (5) 0.686 (11) 0.657 (14) 0.124 (10) 0.348 (14) 0.297 (15) 0.345 (14) 0.831 (11) 0.184 (14) 0.413 (15) 0.423 (15) 0.061 (15) 0.091 (15)
0.863 (8) 0.880 (6) 0.469 (11) 0.679 (7) 0.518 (10) 0.665 (8) 0.933 (11) 0.400 (10) 0.868 (3) 0.263 (14) 0.051 (6) 0.255 (5) 0.837 (11) 0.540 (9) 0.483 (5) 0.352 (13) 0.778 (13) 0.334 (14) 0.587 (4) 0.602 (12) 0.512 (7) 0.383 (10) 0.515 (15) 0.708 (12) 0.972 (3) 0.554 (14) 0.800 (8) 0.496 (9) 0.110 (1) 0.679 (6) 0.812 (8) 0.492 (7) 0.605 (7) 0.638 (9) 0.099 (15) 0.615 (15) 0.776 (12) 0.110 (11) 0.903 (9) 0.434 (8) 0.394 (11) 0.791 (14) 0.245 (8) 0.533 (10) 0.434 (14) 0.107 (11) 0.177 (6)
0.289 (15) 0.113 (15) 0.373 (18) 0.443 (17) 0.136 (18) 0.252 (17) 0.251 (17) 0.335 (14) 0.691 (11) 0.167 (18) 0.049 (8) 0.082 (18) 0.343 (16) 0.385 (18) 0.028 (17) 0.419 (10) 0.040 (18) 0.061 (15) 0.474 (13) 0.105 (18) 0.364 (18) 0.133 (16) 0.582 (11) 0.585 (14) 0.790 (15) 0.180 (18) 0.307 (16) 0.410 (16) 0.015 (18) 0.198 (18) 0.280 (18) 0.150 (16) 0.094 (18) 0.063 (18) 0.040 (18) 0.666 (14) 0.364 (17) 0.061 (16) 0.077 (18) 0.227 (17) 0.323 (15) 0.886 (8) 0.239 (10) 0.448 (14) 0.551 (8) 0.101 (12) 0.075 (17)
0.009 (18) 0.099 (16) 0.493 (8) 0.561 (15) 0.188 (16) 0.246 (18) 0.446 (15) 0.440 (7) 0.860 (4) 0.204 (17) 0.041 (16) 0.165 (13) 0.627 (14) 0.526 (11) 0.126 (15) 0.204 (18) 0.663 (16) 0.029 (18) 0.457 (15) 0.175 (16) 0.457 (15) 0.344 (13) 0.643 (8) 0.481 (16) 0.204 (18) 0.212 (17) 0.130 (18) 0.402 (17) 0.020 (17) 0.287 (15) 0.498 (15) 0.098 (17) 0.316 (17) 0.318 (14) 0.042 (17) 0.414 (17) 0.634 (15) 0.088 (14) 0.136 (16) 0.231 (16) 0.316 (16) 0.845 (10) 0.221 (12) 0.361 (16) 0.351 (17) 0.060 (16) 0.078 (16)
0.909 (4) 0.986 (2) 0.506 (4) 0.722 (1) 0.237 (15) 0.749 (1) 1.000 (1) 0.402 (9) 0.643 (13) 0.377 (10) 0.041 (15) 0.161 (15) 0.928 (9) 0.628 (5) 0.608 (1) 0.559 (4) 1.000 (1) 0.988 (1) 0.500 (11) 0.914 (3) 0.521 (6) 0.346 (12) 0.689 (5) 0.899 (4) 0.971 (5) 0.668 (9) 0.800 (6) 0.570 (4) 0.106 (3) 0.634 (9) 0.913 (1) 0.330 (13) 0.629 (6) 0.771 (6) 0.148 (14) 0.885 (2) 1.000 (1) 0.088 (13) 1.000 (1) 0.448 (4) 0.436 (4) 0.895 (6) 0.241 (9) 0.804 (3) 0.655 (3) 0.207 (1) 0.174 (7)
0.904 (5) 0.988 (1) 0.469 (12) 0.709 (2) 0.722 (9) 0.657 (10) 1.000 (2) 0.362 (13) 0.539 (16) 0.374 (11) 0.044 (14) 0.184 (11) 0.948 (5) 0.512 (14) 0.499 (4) 0.471 (8) 1.000 (3) 0.985 (3) 0.492 (12) 0.962 (1) 0.461 (14) 0.427 (8) 0.655 (7) 0.910 (3) 0.973 (2) 0.779 (4) 0.801 (5) 0.477 (10) 0.082 (6) 0.641 (8) 0.909 (2) 0.490 (8) 0.661 (5) 0.799 (4) 0.267 (4) 0.821 (7) 1.000 (2) 0.280 (7) 1.000 (2) 0.444 (5) 0.415 (7) 0.702 (16) 0.181 (15) 0.686 (8) 0.539 (9) 0.198 (2) 0.143 (10)
0.903 (6) 0.967 (3) 0.489 (9) 0.676 (8) 0.383 (11) 0.702 (6) 1.000 (4) 0.388 (12) 0.625 (14) 0.382 (8) 0.045 (13) 0.165 (14) 0.938 (7) 0.603 (7) 0.525 (2) 0.529 (6) 1.000 (4) 0.988 (2) 0.508 (10) 0.913 (4) 0.523 (5) 0.370 (11) 0.716 (3) 0.880 (5) 0.971 (6) 0.676 (8) 0.800 (9) 0.556 (5) 0.106 (2) 0.677 (7) 0.899 (3) 0.357 (11) 0.673 (4) 0.829 (3) 0.183 (10) 0.855 (5) 1.000 (3) 0.091 (12) 1.000 (3) 0.463 (3) 0.434 (5) 0.818 (13) 0.226 (11) 0.774 (4) 0.629 (5) 0.197 (3) 0.166 (8)
0.605 (13) 0.810 (8) 0.393 (17) 0.622 (12) 0.810 (6) 0.540 (12) 0.918 (12) 0.279 (16) 0.430 (18) 0.402 (7) 0.040 (17) 0.252 (6) 0.695 (13) 0.453 (17) 0.347 (12) 0.313 (16) 0.733 (14) 0.574 (11) 0.328 (17) 0.564 (13) 0.412 (17) 0.509 (6) 0.428 (17) 0.576 (15) 0.908 (13) 0.794 (3) 0.540 (12) 0.415 (14) 0.089 (5) 0.524 (13) 0.636 (12) 0.425 (9) 0.426 (13) 0.579 (11) 0.206 (9) 0.455 (16) 0.766 (13) 0.739 (1) 0.730 (13) 0.427 (10) 0.421 (6) 0.510 (18) 0.171 (16) 0.323 (18) 0.436 (13) 0.100 (13) 0.107 (12)
0.919 (3) 0.749 (11) 0.476 (10) 0.685 (6) 0.361 (12) 0.489 (13) 0.583 (14) 0.407 (8) 0.792 (7) 0.338 (12) 0.049 (9) 0.204 (9) 0.981 (1) 0.644 (4) 0.353 (11) 0.574 (3) 1.000 (2) 0.822 (7) 0.530 (8) 0.793 (8) 0.486 (10) 0.393 (9) 0.745 (1) 0.927 (2) 0.971 (4) 0.610 (11) 0.305 (17) 0.525 (7) 0.064 (9) 0.619 (10) 0.895 (4) 0.333 (12) 0.825 (2) 0.771 (5) 0.175 (13) 0.801 (8) 0.866 (11) 0.069 (15) 0.976 (6) 0.387 (13) 0.395 (10) 0.914 (5) 0.212 (13) 0.835 (1) 0.729 (1) 0.132 (8) 0.148 (9)
0.095 (17) 0.076 (18) 0.497 (7) 0.370 (18) 0.255 (14) 0.286 (15) 0.102 (18) 0.147 (18) 0.836 (5) 0.287 (13) 0.045 (12) 0.330 (4) 0.218 (18) 0.549 (8) 0.028 (18) 0.336 (15) 0.923 (11) 0.046 (16) 0.303 (18) 0.219 (15) 0.445 (16) 0.094 (18) 0.575 (13) 0.333 (18) 0.535 (17) 0.293 (15) 0.367 (15) 0.393 (18) 0.021 (16) 0.216 (17) 0.370 (16) 0.156 (15) 0.390 (14) 0.293 (15) 0.248 (6) 0.309 (18) 0.167 (18) 0.042 (17) 0.130 (17) 0.310 (14) 0.305 (17) 0.591 (17) 0.262 (6) 0.350 (17) 0.333 (18) 0.028 (18) 0.068 (18)
0.097 (16) 0.090 (17) 0.542 (1) 0.493 (16) 0.166 (17) 0.276 (16) 0.293 (16) 0.314 (15) 0.918 (1) 0.214 (16) 0.049 (10) 0.223 (7) 0.323 (17) 0.652 (3) 0.113 (16) 0.379 (12) 0.445 (17) 0.039 (17) 0.533 (7) 0.112 (17) 0.551 (3) 0.127 (17) 0.507 (16) 0.435 (17) 0.886 (14) 0.291 (16) 0.445 (14) 0.420 (13) 0.026 (14) 0.297 (14) 0.284 (17) 0.094 (18) 0.371 (15) 0.252 (17) 0.077 (16) 0.685 (12) 0.423 (16) 0.037 (18) 0.171 (15) 0.227 (18) 0.291 (18) 0.930 (2) 0.248 (7) 0.491 (12) 0.420 (16) 0.034 (17) 0.103 (13)
0.793 (10) 0.661 (13) 0.441 (15) 0.697 (4) 0.750 (8) 0.706 (5) 0.960 (10) 0.560 (3) 0.735 (8) 0.431 (6) 0.048 (11) 0.182 (12) 0.885 (10) 0.510 (15) 0.324 (14) 0.426 (9) 1.000 (6) 0.651 (10) 0.524 (9) 0.737 (9) 0.503 (8) 0.508 (7) 0.568 (14) 0.737 (11) 0.925 (12) 0.629 (10) 0.471 (13) 0.516 (8) 0.033 (12) 0.592 (12) 0.655 (11) 0.508 (4) 0.516 (11) 0.440 (13) 0.179 (11) 0.742 (10) 0.950 (9) 0.313 (6) 0.885 (10) 0.409 (12) 0.385 (13) 0.825 (12) 0.319 (5) 0.640 (9) 0.536 (10) 0.130 (9) 0.288 (1)
0.740 (11) 0.751 (10) 0.500 (5) 0.708 (3) 0.916 (1) 0.728 (2) 0.963 (9) 0.670 (2) 0.869 (2) 0.468 (3) 0.052 (5) 0.197 (10) 0.957 (2) 0.526 (12) 0.429 (7) 0.556 (5) 0.995 (8) 0.794 (8) 0.674 (3) 0.835 (7) 0.556 (2) 0.616 (3) 0.603 (10) 0.854 (8) 0.960 (8) 0.816 (2) 0.900 (1) 0.577 (2) 0.029 (13) 0.719 (3) 0.814 (7) 0.500 (6) 0.508 (12) 0.606 (10) 0.208 (8) 0.865 (3) 0.997 (4) 0.440 (5) 0.962 (7) 0.419 (11) 0.390 (12) 0.951 (1) 0.417 (1) 0.818 (2) 0.604 (6) 0.119 (10) 0.259 (3)
0.894 (7) 0.909 (4) 0.461 (14) 0.666 (9) 0.824 (3) 0.662 (9) 1.000 (5) 0.192 (17) 0.447 (17) 0.526 (1) 0.033 (18) 0.153 (16) 0.941 (6) 0.520 (13) 0.401 (9) 0.385 (11) 0.986 (10) 0.865 (6) 0.348 (16) 0.919 (2) 0.464 (12) 0.284 (14) 0.658 (6) 0.829 (9) 0.959 (9) 0.905 (1) 0.800 (7) 0.436 (12) 0.093 (4) 0.750 (1) 0.669 (10) 0.533 (1) 0.597 (8) 0.706 (7) 0.319 (3) 0.838 (6) 0.990 (7) 0.261 (8) 1.000 (4) 0.490 (1) 0.440 (3) 0.765 (15) 0.139 (18) 0.475 (13) 0.556 (7) 0.168 (5) 0.110 (11)
0.576 (14) 0.693 (12) 0.429 (16) 0.590 (13) 0.797 (7) 0.605 (11) 1.000 (6) 0.482 (5) 0.616 (15) 0.235 (15) 0.050 (7) 0.134 (17) 0.938 (8) 0.495 (16) 0.379 (10) 0.299 (17) 0.679 (15) 0.381 (13) 0.473 (14) 0.409 (14) 0.494 (9) 0.521 (5) 0.396 (18) 0.602 (13) 0.787 (16) 0.715 (6) 0.704 (10) 0.450 (11) 0.050 (10) 0.607 (11) 0.532 (14) 0.521 (3) 0.357 (16) 0.488 (12) 0.175 (12) 0.798 (9) 0.962 (8) 0.721 (3) 0.878 (11) 0.438 (6) 0.440 (2) 0.885 (9) 0.164 (17) 0.524 (11) 0.466 (12) 0.100 (14) 0.094 (14)
0.922 (2) 0.894 (5) 0.499 (6) 0.690 (5) 0.885 (2) 0.719 (3) 0.965 (8) 0.475 (6) 0.702 (10) 0.435 (5) 0.080 (2) 0.563 (2) 0.949 (4) 0.693 (2) 0.500 (3) 0.636 (2) 0.997 (7) 0.883 (5) 0.702 (1) 0.865 (6) 0.541 (4) 0.654 (2) 0.693 (4) 0.866 (7) 0.987 (1) 0.558 (13) 0.802 (4) 0.575 (3) 0.075 (7) 0.736 (2) 0.861 (5) 0.532 (2) 0.749 (3) 0.892 (1) 0.355 (2) 0.859 (4) 0.995 (5) 0.680 (4) 0.982 (5) 0.474 (2) 0.411 (8) 0.924 (3) 0.370 (2) 0.746 (6) 0.630 (4) 0.149 (6) 0.196 (5)
0.954 (1) 0.863 (7) 0.513 (2) 0.645 (10) 0.822 (4) 0.712 (4) 1.000 (3) 0.699 (1) 0.793 (6) 0.484 (2) 0.094 (1) 0.576 (1) 0.954 (3) 0.699 (1) 0.468 (6) 0.817 (1) 1.000 (5) 0.930 (4) 0.694 (2) 0.891 (5) 0.556 (1) 0.587 (4) 0.720 (2) 0.939 (1) 0.963 (7) 0.760 (5) 0.900 (2) 0.672 (1) 0.066 (8) 0.706 (4) 0.828 (6) 0.500 (5) 0.826 (1) 0.863 (2) 0.367 (1) 0.899 (1) 0.991 (6) 0.734 (2) 0.930 (8) 0.435 (7) 0.404 (9) 0.920 (4) 0.338 (3) 0.755 (5) 0.656 (2) 0.185 (4) 0.227 (4)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.467 (9) 0.509 (9) 0.546 (8) 0.635 (5) 0.069 (12)
0.313 (13) 0.286 (15) 0.310 (14) 0.455 (17) 0.064 (15)
0.306 (14) 0.315 (14) 0.351 (13) 0.587 (9) 0.070 (10)
0.415 (11) 0.411 (12) 0.220 (16) 0.502 (15) 0.060 (17)
0.115 (18) 0.274 (17) 0.138 (18) 0.686 (2) 0.057 (18)
0.554 (5) 0.600 (4) 0.735 (2) 0.591 (8) 0.084 (2)
0.499 (7) 0.583 (6) 0.740 (1) 0.533 (13) 0.078 (6)
0.558 (4) 0.640 (2) 0.734 (3) 0.568 (11) 0.079 (5)
0.235 (16) 0.277 (16) 0.480 (11) 0.479 (16) 0.085 (1)
0.590 (2) 0.640 (3) 0.691 (6) 0.657 (4) 0.076 (7)
0.212 (17) 0.331 (13) 0.419 (12) 0.390 (18) 0.063 (16)
0.393 (12) 0.435 (11) 0.181 (17) 0.526 (14) 0.067 (14)
0.432 (10) 0.487 (10) 0.540 (9) 0.634 (6) 0.071 (9)
0.498 (8) 0.559 (8) 0.491 (10) 0.723 (1) 0.067 (13)
0.512 (6) 0.580 (7) 0.723 (4) 0.569 (10) 0.083 (3)
0.295 (15) 0.272 (18) 0.268 (15) 0.541 (12) 0.070 (11)
0.573 (3) 0.585 (5) 0.629 (7) 0.684 (3) 0.082 (4)
0.647 (1) 0.702 (1) 0.704 (5) 0.597 (7) 0.072 (8)
20news agnews amazon imdb yelp
0.257 (9) 0.240 (9) 0.087 (12) 0.115 (10) 0.171 (9)
0.149 (16) 0.144 (17) 0.066 (15) 0.070 (15) 0.082 (15)
0.209 (12) 0.171 (14) 0.087 (13) 0.081 (14) 0.124 (13)
0.083 (18) 0.085 (18) 0.052 (17) 0.047 (17) 0.050 (17)
0.114 (17) 0.152 (15) 0.049 (18) 0.047 (18) 0.049 (18)
0.365 (3) 0.409 (4) 0.186 (1) 0.187 (3) 0.380 (1)
0.299 (7) 0.348 (6) 0.137 (6) 0.158 (6) 0.277 (6)
0.349 (5) 0.401 (5) 0.179 (3) 0.190 (2) 0.350 (3)
0.241 (10) 0.222 (12) 0.131 (7) 0.144 (7) 0.216 (8)
0.353 (4) 0.412 (3) 0.153 (5) 0.169 (5) 0.309 (4)
0.176 (14) 0.177 (13) 0.092 (9) 0.111 (11) 0.107 (14)
0.150 (15) 0.294 (8) 0.055 (16) 0.054 (16) 0.061 (16)
0.258 (8) 0.234 (10) 0.090 (11) 0.126 (9) 0.167 (10)
0.219 (11) 0.224 (11) 0.091 (10) 0.089 (13) 0.143 (11)
0.329 (6) 0.458 (1) 0.182 (2) 0.211 (1) 0.377 (2)
0.182 (13) 0.145 (16) 0.084 (14) 0.095 (12) 0.125 (12)
0.377 (1) 0.305 (7) 0.116 (8) 0.131 (8) 0.226 (7)
0.371 (2) 0.439 (2) 0.154 (4) 0.183 (4) 0.309 (5)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 50: The average of AUCPR performance on individual datasets under 𝑁𝑙𝑎 = 10 setting Dataset
XGBOD
DeepSAD
REPEN
AABiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
DualMGAN
SOEL-NTL
AnoDDAE
XGB
CatB
TabMCls
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.847 (10) 0.857 (10) 0.487 (12) 0.736 (1) 0.896 (7) 0.851 (4) 0.967 (8) 0.526 (6) 0.818 (8) 0.520 (2) 0.060 (4) 0.457 (3) 0.908 (12) 0.582 (10) 0.427 (12) 0.599 (5) 0.999 (8) 0.819 (9) 0.652 (8) 0.821 (11) 0.521 (13) 0.732 (2) 0.657 (8) 0.852 (10) 0.963 (11) 0.704 (10) 0.701 (11) 0.621 (6) 0.045 (11) 0.802 (6) 0.870 (9) 0.609 (6) 0.622 (8) 0.674 (11) 0.361 (5) 0.883 (6) 0.962 (10) 0.347 (8) 0.960 (10) 0.511 (8) 0.420 (6) 0.929 (7) 0.372 (4) 0.809 (8) 0.611 (10) 0.163 (9) 0.324 (1)
0.847 (11) 0.882 (8) 0.507 (7) 0.682 (9) 0.455 (11) 0.624 (13) 0.942 (11) 0.431 (12) 0.761 (11) 0.437 (8) 0.061 (3) 0.246 (10) 0.900 (13) 0.615 (7) 0.522 (8) 0.424 (12) 0.924 (12) 0.815 (10) 0.661 (6) 0.875 (8) 0.486 (14) 0.319 (15) 0.621 (10) 0.912 (7) 0.970 (8) 0.781 (7) 0.803 (3) 0.436 (13) 0.027 (14) 0.426 (14) 0.745 (13) 0.271 (14) 0.626 (7) 0.382 (15) 0.318 (7) 0.725 (14) 0.898 (12) 0.194 (10) 0.534 (14) 0.346 (14) 0.356 (14) 0.852 (14) 0.205 (15) 0.547 (13) 0.521 (13) 0.095 (14) 0.114 (13)
0.882 (7) 0.928 (5) 0.489 (11) 0.677 (11) 0.635 (10) 0.809 (8) 0.933 (12) 0.492 (7) 0.901 (3) 0.297 (14) 0.054 (6) 0.317 (6) 0.948 (11) 0.582 (11) 0.550 (6) 0.455 (11) 0.943 (11) 0.468 (14) 0.743 (2) 0.872 (9) 0.557 (6) 0.486 (9) 0.551 (15) 0.812 (11) 0.967 (10) 0.539 (14) 0.800 (8) 0.558 (10) 0.139 (1) 0.809 (5) 0.870 (10) 0.509 (8) 0.602 (10) 0.778 (8) 0.132 (15) 0.842 (11) 0.892 (13) 0.108 (11) 0.983 (6) 0.478 (11) 0.363 (13) 0.899 (12) 0.241 (11) 0.708 (10) 0.601 (11) 0.142 (11) 0.189 (9)
0.170 (16) 0.115 (15) 0.390 (18) 0.384 (18) 0.096 (18) 0.293 (17) 0.559 (15) 0.306 (15) 0.658 (14) 0.160 (18) 0.046 (14) 0.065 (18) 0.330 (18) 0.422 (18) 0.050 (17) 0.477 (10) 0.030 (18) 0.148 (15) 0.410 (17) 0.071 (18) 0.411 (18) 0.164 (16) 0.587 (13) 0.497 (15) 0.737 (16) 0.189 (18) 0.307 (16) 0.377 (17) 0.016 (18) 0.174 (18) 0.401 (16) 0.121 (16) 0.040 (18) 0.034 (18) 0.047 (17) 0.692 (15) 0.332 (18) 0.077 (14) 0.067 (18) 0.206 (18) 0.305 (17) 0.951 (5) 0.213 (14) 0.510 (14) 0.481 (15) 0.085 (16) 0.062 (18)
0.009 (18) 0.111 (16) 0.493 (10) 0.567 (13) 0.188 (17) 0.248 (18) 0.452 (16) 0.442 (10) 0.861 (6) 0.204 (17) 0.041 (17) 0.167 (17) 0.671 (15) 0.504 (15) 0.132 (15) 0.203 (18) 0.585 (16) 0.029 (18) 0.414 (16) 0.177 (16) 0.459 (16) 0.339 (14) 0.644 (9) 0.489 (16) 0.419 (18) 0.198 (17) 0.130 (18) 0.402 (16) 0.020 (17) 0.293 (16) 0.504 (15) 0.098 (17) 0.344 (17) 0.322 (16) 0.043 (18) 0.423 (17) 0.641 (16) 0.067 (16) 0.147 (17) 0.231 (16) 0.317 (16) 0.850 (15) 0.240 (12) 0.365 (18) 0.355 (18) 0.087 (15) 0.078 (16)
0.943 (3) 0.958 (3) 0.497 (9) 0.722 (5) 0.318 (15) 0.784 (10) 1.000 (1) 0.453 (9) 0.745 (12) 0.372 (10) 0.050 (8) 0.204 (16) 0.996 (7) 0.672 (4) 0.612 (1) 0.619 (4) 1.000 (2) 0.990 (3) 0.645 (9) 0.921 (5) 0.599 (3) 0.425 (12) 0.717 (2) 0.929 (4) 0.971 (6) 0.660 (12) 0.800 (6) 0.698 (3) 0.117 (3) 0.689 (13) 0.921 (2) 0.362 (13) 0.805 (5) 0.833 (6) 0.209 (12) 0.907 (4) 1.000 (2) 0.087 (13) 1.000 (1) 0.463 (12) 0.419 (8) 0.954 (4) 0.279 (6) 0.901 (1) 0.738 (3) 0.238 (3) 0.234 (6)
0.928 (5) 0.993 (1) 0.480 (13) 0.720 (6) 0.851 (9) 0.826 (6) 1.000 (2) 0.401 (13) 0.645 (15) 0.442 (7) 0.049 (13) 0.263 (8) 1.000 (1) 0.594 (8) 0.569 (5) 0.540 (9) 1.000 (4) 0.991 (1) 0.639 (10) 0.965 (1) 0.530 (11) 0.616 (7) 0.675 (7) 0.936 (3) 0.970 (7) 0.860 (5) 0.801 (5) 0.573 (8) 0.102 (5) 0.739 (9) 0.901 (5) 0.493 (9) 0.764 (6) 0.878 (3) 0.293 (8) 0.881 (7) 1.000 (3) 0.366 (7) 1.000 (2) 0.538 (2) 0.442 (1) 0.754 (16) 0.217 (13) 0.864 (6) 0.621 (9) 0.236 (4) 0.186 (10)
0.950 (1) 0.966 (2) 0.506 (8) 0.723 (4) 0.424 (12) 0.797 (9) 1.000 (4) 0.434 (11) 0.731 (13) 0.381 (9) 0.049 (12) 0.231 (13) 1.000 (2) 0.661 (5) 0.579 (3) 0.576 (6) 1.000 (6) 0.991 (2) 0.658 (7) 0.917 (6) 0.598 (4) 0.435 (10) 0.707 (5) 0.919 (5) 0.972 (5) 0.677 (11) 0.800 (9) 0.667 (4) 0.107 (4) 0.733 (10) 0.906 (4) 0.385 (12) 0.842 (3) 0.873 (4) 0.239 (10) 0.878 (9) 1.000 (4) 0.089 (12) 1.000 (3) 0.497 (9) 0.421 (4) 0.883 (13) 0.265 (7) 0.901 (2) 0.729 (5) 0.244 (1) 0.205 (7)
0.753 (13) 0.876 (9) 0.392 (17) 0.509 (14) 0.966 (1) 0.747 (12) 0.918 (13) 0.296 (16) 0.441 (18) 0.368 (11) 0.043 (16) 0.317 (5) 0.870 (14) 0.479 (17) 0.461 (10) 0.274 (16) 0.815 (14) 0.538 (13) 0.574 (12) 0.660 (14) 0.465 (15) 0.634 (6) 0.479 (17) 0.704 (13) 0.974 (3) 0.926 (2) 0.540 (12) 0.424 (14) 0.126 (2) 0.729 (11) 0.773 (12) 0.549 (7) 0.440 (14) 0.729 (9) 0.193 (13) 0.780 (13) 0.882 (14) 0.846 (1) 0.787 (13) 0.480 (10) 0.421 (5) 0.565 (18) 0.186 (17) 0.423 (17) 0.485 (14) 0.106 (12) 0.105 (14)
0.943 (2) 0.809 (13) 0.519 (5) 0.680 (10) 0.392 (13) 0.612 (14) 0.663 (14) 0.530 (5) 0.832 (7) 0.320 (12) 0.049 (11) 0.207 (15) 0.972 (9) 0.676 (3) 0.409 (13) 0.550 (7) 1.000 (3) 0.925 (6) 0.661 (5) 0.800 (12) 0.534 (10) 0.433 (11) 0.757 (1) 0.942 (2) 0.972 (4) 0.624 (13) 0.305 (17) 0.562 (9) 0.069 (9) 0.708 (12) 0.880 (8) 0.390 (11) 0.825 (4) 0.833 (7) 0.222 (11) 0.805 (12) 0.952 (11) 0.075 (15) 0.955 (11) 0.354 (13) 0.369 (11) 0.942 (6) 0.259 (9) 0.887 (4) 0.736 (4) 0.205 (7) 0.198 (8)
0.253 (15) 0.076 (18) 0.508 (6) 0.488 (17) 0.375 (14) 0.332 (15) 0.039 (18) 0.230 (18) 0.871 (5) 0.298 (13) 0.044 (15) 0.362 (4) 0.384 (17) 0.562 (13) 0.023 (18) 0.376 (15) 0.806 (15) 0.051 (16) 0.284 (18) 0.204 (15) 0.429 (17) 0.098 (18) 0.556 (14) 0.442 (18) 0.530 (17) 0.358 (15) 0.367 (15) 0.376 (18) 0.027 (15) 0.235 (17) 0.333 (17) 0.203 (15) 0.482 (13) 0.446 (14) 0.283 (9) 0.381 (18) 0.724 (15) 0.045 (17) 0.391 (15) 0.306 (15) 0.323 (15) 0.571 (17) 0.262 (8) 0.429 (16) 0.382 (17) 0.036 (17) 0.066 (17)
0.115 (17) 0.095 (17) 0.536 (1) 0.497 (15) 0.199 (16) 0.320 (16) 0.293 (17) 0.318 (14) 0.915 (1) 0.217 (16) 0.049 (9) 0.225 (14) 0.449 (16) 0.651 (6) 0.118 (16) 0.378 (14) 0.468 (17) 0.041 (17) 0.540 (14) 0.113 (17) 0.545 (8) 0.125 (17) 0.504 (16) 0.447 (17) 0.884 (14) 0.279 (16) 0.445 (14) 0.420 (15) 0.027 (16) 0.318 (15) 0.277 (18) 0.093 (18) 0.372 (16) 0.288 (17) 0.079 (16) 0.661 (16) 0.580 (17) 0.036 (18) 0.165 (16) 0.223 (17) 0.291 (18) 0.925 (8) 0.251 (10) 0.489 (15) 0.436 (16) 0.035 (18) 0.103 (15)
0.866 (9) 0.677 (14) 0.463 (15) 0.707 (7) 0.905 (6) 0.844 (5) 0.960 (10) 0.539 (4) 0.816 (9) 0.481 (6) 0.056 (5) 0.244 (12) 0.996 (8) 0.547 (14) 0.350 (14) 0.541 (8) 0.999 (9) 0.787 (11) 0.609 (11) 0.840 (10) 0.544 (9) 0.614 (8) 0.600 (12) 0.794 (12) 0.914 (13) 0.729 (9) 0.471 (13) 0.602 (7) 0.039 (12) 0.777 (7) 0.893 (7) 0.618 (4) 0.614 (9) 0.608 (12) 0.330 (6) 0.873 (10) 0.974 (9) 0.418 (6) 0.967 (8) 0.525 (5) 0.365 (12) 0.914 (11) 0.344 (5) 0.726 (9) 0.630 (8) 0.165 (8) 0.314 (2)
0.842 (12) 0.832 (12) 0.522 (4) 0.727 (3) 0.962 (2) 0.856 (3) 0.965 (9) 0.683 (2) 0.902 (2) 0.519 (3) 0.049 (10) 0.311 (7) 0.996 (6) 0.568 (12) 0.449 (11) 0.629 (3) 1.000 (1) 0.876 (7) 0.729 (4) 0.906 (7) 0.582 (5) 0.696 (3) 0.613 (11) 0.882 (9) 0.970 (9) 0.924 (3) 0.900 (1) 0.660 (5) 0.037 (13) 0.843 (2) 0.900 (6) 0.648 (3) 0.518 (12) 0.685 (10) 0.366 (4) 0.933 (3) 1.000 (1) 0.607 (5) 0.982 (7) 0.520 (7) 0.414 (9) 0.968 (1) 0.424 (2) 0.883 (5) 0.672 (6) 0.148 (10) 0.266 (4)
0.868 (8) 0.947 (4) 0.476 (14) 0.731 (2) 0.943 (5) 0.822 (7) 1.000 (5) 0.244 (17) 0.568 (17) 0.510 (5) 0.036 (18) 0.246 (11) 1.000 (3) 0.582 (9) 0.529 (7) 0.410 (13) 0.989 (10) 0.873 (8) 0.548 (13) 0.953 (2) 0.548 (7) 0.392 (13) 0.703 (6) 0.901 (8) 0.961 (12) 0.947 (1) 0.800 (7) 0.539 (12) 0.099 (6) 0.842 (4) 0.795 (11) 0.617 (5) 0.599 (11) 0.845 (5) 0.412 (3) 0.906 (5) 1.000 (5) 0.299 (9) 0.999 (4) 0.533 (3) 0.440 (2) 0.918 (9) 0.152 (18) 0.699 (11) 0.643 (7) 0.219 (6) 0.152 (11)
0.653 (14) 0.837 (11) 0.426 (16) 0.494 (16) 0.861 (8) 0.768 (11) 0.994 (7) 0.461 (8) 0.606 (16) 0.247 (15) 0.050 (7) 0.251 (9) 0.952 (10) 0.490 (16) 0.513 (9) 0.248 (17) 0.824 (13) 0.593 (12) 0.531 (15) 0.704 (13) 0.527 (12) 0.696 (4) 0.391 (18) 0.606 (14) 0.850 (15) 0.891 (4) 0.704 (10) 0.539 (11) 0.079 (8) 0.763 (8) 0.607 (14) 0.484 (10) 0.433 (15) 0.569 (13) 0.155 (14) 0.878 (8) 0.981 (8) 0.786 (2) 0.915 (12) 0.529 (4) 0.429 (3) 0.917 (10) 0.186 (16) 0.586 (12) 0.527 (12) 0.101 (13) 0.131 (12)
0.903 (6) 0.923 (6) 0.531 (2) 0.695 (8) 0.959 (3) 0.880 (1) 1.000 (6) 0.586 (3) 0.803 (10) 0.516 (4) 0.081 (2) 0.654 (2) 0.998 (5) 0.717 (2) 0.608 (2) 0.686 (2) 1.000 (7) 0.952 (4) 0.739 (3) 0.942 (4) 0.607 (1) 0.738 (1) 0.713 (4) 0.917 (6) 0.982 (1) 0.770 (8) 0.802 (4) 0.737 (2) 0.095 (7) 0.842 (3) 0.949 (1) 0.662 (2) 0.849 (2) 0.939 (1) 0.515 (2) 0.935 (2) 1.000 (6) 0.756 (3) 0.993 (5) 0.562 (1) 0.420 (7) 0.962 (2) 0.489 (1) 0.889 (3) 0.750 (2) 0.221 (5) 0.265 (5)
0.941 (4) 0.895 (7) 0.531 (3) 0.650 (12) 0.950 (4) 0.867 (2) 1.000 (3) 0.721 (1) 0.880 (4) 0.520 (1) 0.107 (1) 0.668 (1) 0.998 (4) 0.732 (1) 0.575 (4) 0.851 (1) 1.000 (5) 0.947 (5) 0.759 (1) 0.943 (3) 0.606 (2) 0.685 (5) 0.714 (3) 0.953 (1) 0.979 (2) 0.846 (6) 0.900 (2) 0.808 (1) 0.051 (10) 0.850 (1) 0.920 (3) 0.672 (1) 0.921 (1) 0.907 (2) 0.544 (1) 0.956 (1) 0.997 (7) 0.751 (4) 0.962 (9) 0.523 (6) 0.413 (10) 0.959 (3) 0.417 (3) 0.858 (7) 0.751 (1) 0.242 (2) 0.282 (3)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.578 (9) 0.599 (9) 0.654 (8) 0.729 (3) 0.081 (11)
0.374 (15) 0.383 (15) 0.438 (14) 0.511 (16) 0.069 (14)
0.408 (11) 0.437 (12) 0.501 (12) 0.666 (10) 0.080 (12)
0.399 (12) 0.425 (14) 0.221 (16) 0.494 (17) 0.060 (17)
0.114 (18) 0.275 (18) 0.138 (18) 0.686 (8) 0.057 (18)
0.666 (4) 0.708 (4) 0.799 (2) 0.707 (6) 0.100 (2)
0.614 (7) 0.694 (6) 0.804 (1) 0.641 (12) 0.097 (5)
0.655 (5) 0.724 (3) 0.793 (3) 0.686 (7) 0.098 (4)
0.350 (16) 0.363 (16) 0.628 (11) 0.552 (14) 0.099 (3)
0.670 (2) 0.724 (2) 0.777 (4) 0.727 (4) 0.093 (7)
0.288 (17) 0.437 (11) 0.474 (13) 0.455 (18) 0.066 (15)
0.397 (13) 0.433 (13) 0.181 (17) 0.527 (15) 0.066 (16)
0.559 (10) 0.582 (10) 0.653 (9) 0.727 (5) 0.082 (9)
0.580 (8) 0.626 (8) 0.629 (10) 0.780 (1) 0.071 (13)
0.650 (6) 0.705 (5) 0.770 (6) 0.661 (11) 0.094 (6)
0.379 (14) 0.362 (17) 0.382 (15) 0.609 (13) 0.082 (10)
0.667 (3) 0.686 (7) 0.760 (7) 0.767 (2) 0.102 (1)
0.716 (1) 0.757 (1) 0.771 (5) 0.666 (9) 0.089 (8)
20news agnews amazon imdb yelp
0.294 (9) 0.338 (8) 0.141 (10) 0.161 (10) 0.242 (9)
0.188 (15) 0.202 (15) 0.093 (14) 0.080 (15) 0.114 (15)
0.257 (13) 0.259 (13) 0.110 (13) 0.108 (13) 0.207 (11)
0.083 (18) 0.086 (18) 0.049 (18) 0.046 (18) 0.053 (17)
0.115 (17) 0.151 (17) 0.051 (17) 0.047 (17) 0.049 (18)
0.426 (5) 0.523 (2) 0.224 (2) 0.334 (1) 0.462 (1)
0.348 (7) 0.459 (6) 0.203 (4) 0.260 (5) 0.359 (6)
0.437 (2) 0.508 (4) 0.224 (1) 0.311 (2) 0.438 (2)
0.296 (8) 0.315 (10) 0.158 (8) 0.212 (8) 0.284 (8)
0.432 (3) 0.492 (5) 0.200 (5) 0.278 (4) 0.410 (4)
0.203 (14) 0.203 (14) 0.084 (15) 0.095 (14) 0.149 (14)
0.153 (16) 0.296 (12) 0.054 (16) 0.055 (16) 0.060 (16)
0.293 (10) 0.334 (9) 0.145 (9) 0.169 (9) 0.240 (10)
0.273 (11) 0.301 (11) 0.110 (12) 0.135 (11) 0.179 (13)
0.430 (4) 0.552 (1) 0.195 (6) 0.284 (3) 0.413 (3)
0.264 (12) 0.173 (16) 0.132 (11) 0.125 (12) 0.180 (12)
0.447 (1) 0.443 (7) 0.182 (7) 0.229 (7) 0.343 (7)
0.399 (6) 0.520 (3) 0.211 (3) 0.260 (6) 0.394 (5)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 51: The average AUCROC performance on individual datasets under the 𝑁𝑙𝑎 = 1 setting. Dataset
XGBOD
DeepSAD
REPEN
AA-BiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
Dual-MGAN
SOEL-NTL
DDAE
XGBoost
CatBoost
TabM
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.942 (9) 0.788 (12) 0.530 (14) 0.924 (8) 0.822 (13) 0.664 (10) 0.991 (15) 0.626 (9) 0.607 (14) 0.620 (8) 0.567 (4) 0.686 (8) 0.992 (3) 0.599 (13) 0.830 (13) 0.653 (13) 0.937 (8) 0.850 (10) 0.810 (8) 0.854 (7) 0.574 (11) 0.760 (10) 0.627 (12) 0.987 (6) 0.809 (16) 0.602 (16) 0.834 (8) 0.523 (8) 0.499 (13) 0.807 (11) 0.953 (9) 0.584 (11) 0.729 (13) 0.869 (6) 0.657 (12) 0.919 (12) 0.980 (9) 0.572 (13) 0.885 (11) 0.495 (15) 0.545 (7) 0.832 (10) 0.725 (3) 0.691 (14) 0.611 (11) 0.705 (6) 0.570 (10)
0.835 (13) 0.846 (10) 0.673 (3) 0.895 (11) 0.801 (14) 0.637 (13) 0.999 (9) 0.684 (4) 0.771 (9) 0.649 (3) 0.583 (2) 0.729 (5) 0.816 (14) 0.725 (5) 0.848 (11) 0.684 (11) 0.841 (16) 0.727 (12) 0.845 (6) 0.827 (11) 0.576 (10) 0.692 (14) 0.675 (10) 0.955 (12) 0.910 (12) 0.796 (7) 0.882 (6) 0.516 (10) 0.527 (5) 0.648 (16) 0.925 (10) 0.559 (15) 0.880 (5) 0.720 (16) 0.659 (10) 0.938 (11) 0.897 (15) 0.567 (14) 0.587 (17) 0.505 (12) 0.468 (14) 0.885 (9) 0.585 (14) 0.691 (13) 0.590 (12) 0.537 (16) 0.552 (13)
0.788 (15) 0.848 (9) 0.584 (9) 0.840 (15) 0.873 (8) 0.677 (7) 1.000 (5) 0.587 (11) 0.624 (13) 0.502 (16) 0.541 (6) 0.597 (12) 0.937 (9) 0.686 (8) 0.861 (7) 0.723 (8) 0.858 (14) 0.736 (11) 0.795 (11) 0.596 (17) 0.563 (13) 0.786 (8) 0.598 (15) 0.911 (16) 0.908 (13) 0.885 (5) 0.907 (4) 0.624 (1) 0.480 (17) 0.743 (15) 0.978 (6) 0.605 (10) 0.779 (10) 0.664 (17) 0.530 (18) 0.843 (14) 0.796 (17) 0.677 (5) 0.698 (13) 0.497 (14) 0.529 (9) 0.705 (15) 0.609 (11) 0.637 (15) 0.576 (13) 0.709 (4) 0.656 (6)
0.829 (14) 0.779 (13) 0.590 (8) 0.839 (16) 0.404 (18) 0.719 (5) 0.844 (16) 0.698 (3) 0.879 (4) 0.356 (18) 0.533 (9) 0.690 (7) 0.998 (1) 0.522 (17) 0.519 (17) 0.804 (5) 0.882 (11) 0.539 (17) 0.803 (9) 0.916 (3) 0.498 (16) 0.540 (17) 0.593 (17) 0.942 (14) 0.973 (8) 0.509 (17) 0.925 (1) 0.468 (16) 0.517 (7) 0.753 (13) 0.872 (14) 0.408 (16) 0.800 (9) 0.561 (18) 0.573 (15) 0.987 (3) 0.952 (13) 0.542 (15) 0.643 (14) 0.401 (18) 0.477 (13) 0.926 (7) 0.610 (10) 0.868 (4) 0.664 (6) 0.774 (3) 0.508 (16)
0.479 (18) 0.594 (16) 0.641 (5) 0.964 (1) 0.719 (17) 0.604 (15) 0.807 (17) 0.670 (6) 0.900 (2) 0.494 (17) 0.536 (8) 0.700 (6) 0.936 (10) 0.579 (14) 0.755 (14) 0.715 (10) 0.843 (15) 0.433 (18) 0.692 (12) 0.697 (15) 0.606 (9) 0.729 (12) 0.701 (7) 0.969 (10) 0.618 (18) 0.493 (18) 0.559 (18) 0.528 (7) 0.506 (12) 0.748 (14) 0.916 (12) 0.372 (17) 0.882 (4) 0.781 (11) 0.537 (17) 0.915 (13) 0.980 (8) 0.298 (18) 0.615 (16) 0.485 (17) 0.451 (16) 0.920 (8) 0.618 (9) 0.783 (8) 0.558 (14) 0.664 (7) 0.596 (8)
0.984 (4) 0.988 (1) 0.530 (13) 0.910 (9) 0.860 (11) 0.674 (8) 1.000 (1) 0.537 (13) 0.595 (15) 0.609 (9) 0.488 (16) 0.533 (17) 0.838 (12) 0.651 (9) 0.921 (1) 0.610 (15) 0.999 (1) 0.989 (1) 0.608 (13) 0.857 (6) 0.626 (6) 0.793 (6) 0.723 (5) 0.987 (7) 0.980 (6) 0.944 (2) 0.693 (15) 0.512 (11) 0.538 (4) 0.865 (7) 0.981 (5) 0.635 (5) 0.720 (14) 0.879 (5) 0.731 (4) 0.971 (8) 0.999 (1) 0.639 (8) 1.000 (2) 0.637 (3) 0.592 (4) 0.783 (12) 0.627 (8) 0.785 (7) 0.661 (7) 0.647 (8) 0.533 (14)
0.985 (3) 0.920 (6) 0.547 (10) 0.891 (12) 0.872 (9) 0.560 (17) 1.000 (2) 0.533 (14) 0.516 (18) 0.602 (10) 0.495 (13) 0.636 (10) 0.707 (16) 0.546 (16) 0.850 (9) 0.593 (16) 0.756 (17) 0.859 (7) 0.600 (14) 0.731 (14) 0.567 (12) 0.743 (11) 0.597 (16) 0.997 (2) 0.981 (5) 0.786 (8) 0.878 (7) 0.491 (14) 0.488 (16) 0.849 (9) 0.830 (15) 0.578 (13) 0.768 (11) 0.810 (10) 0.590 (14) 0.713 (17) 0.985 (6) 0.614 (9) 1.000 (3) 0.536 (10) 0.523 (10) 0.523 (17) 0.553 (16) 0.583 (16) 0.540 (16) 0.562 (15) 0.565 (12)
0.986 (2) 0.983 (2) 0.535 (12) 0.928 (7) 0.907 (5) 0.632 (14) 1.000 (3) 0.530 (15) 0.568 (16) 0.632 (5) 0.483 (17) 0.538 (14) 0.816 (13) 0.644 (10) 0.911 (2) 0.642 (14) 0.997 (3) 0.977 (2) 0.581 (16) 0.843 (9) 0.618 (7) 0.796 (5) 0.686 (9) 0.989 (5) 0.980 (7) 0.948 (1) 0.652 (17) 0.511 (12) 0.555 (2) 0.866 (6) 0.985 (4) 0.668 (3) 0.744 (12) 0.894 (3) 0.734 (3) 0.960 (10) 0.999 (2) 0.641 (7) 1.000 (1) 0.648 (2) 0.595 (3) 0.721 (14) 0.601 (12) 0.773 (9) 0.679 (5) 0.628 (9) 0.580 (9)
0.532 (17) 0.697 (15) 0.482 (17) 0.654 (18) 0.744 (16) 0.568 (16) 1.000 (6) 0.455 (17) 0.517 (17) 0.642 (4) 0.491 (15) 0.537 (15) 0.629 (17) 0.558 (15) 0.563 (16) 0.513 (18) 0.689 (18) 0.542 (16) 0.564 (17) 0.542 (18) 0.480 (17) 0.846 (4) 0.612 (14) 0.426 (18) 0.709 (17) 0.899 (4) 0.785 (12) 0.462 (17) 0.488 (15) 0.579 (18) 0.694 (18) 0.676 (2) 0.702 (16) 0.734 (15) 0.658 (11) 0.627 (18) 0.812 (16) 0.902 (2) 0.626 (15) 0.594 (6) 0.582 (5) 0.547 (16) 0.514 (17) 0.463 (18) 0.516 (17) 0.598 (12) 0.502 (17)
0.993 (1) 0.948 (4) 0.543 (11) 0.933 (5) 0.930 (3) 0.671 (9) 0.997 (12) 0.582 (12) 0.870 (5) 0.621 (7) 0.523 (11) 0.609 (11) 0.991 (4) 0.754 (2) 0.898 (4) 0.806 (4) 0.988 (5) 0.853 (9) 0.853 (5) 0.838 (10) 0.542 (15) 0.771 (9) 0.737 (2) 0.989 (4) 0.988 (2) 0.917 (3) 0.790 (11) 0.547 (6) 0.509 (10) 0.927 (1) 0.991 (1) 0.609 (9) 0.851 (6) 0.863 (7) 0.653 (13) 0.980 (7) 0.938 (14) 0.581 (12) 0.987 (5) 0.651 (1) 0.553 (6) 0.824 (11) 0.591 (13) 0.948 (1) 0.751 (1) 0.617 (11) 0.533 (15)
0.759 (16) 0.580 (17) 0.663 (4) 0.928 (6) 0.863 (10) 0.534 (18) 0.606 (18) 0.336 (18) 0.889 (3) 0.663 (2) 0.539 (7) 0.876 (2) 0.333 (18) 0.722 (6) 0.347 (18) 0.847 (2) 0.946 (7) 0.579 (14) 0.593 (15) 0.846 (8) 0.550 (14) 0.534 (18) 0.731 (3) 0.908 (17) 0.951 (10) 0.745 (11) 0.826 (9) 0.478 (15) 0.570 (1) 0.620 (17) 0.806 (16) 0.629 (6) 0.911 (2) 0.887 (4) 0.725 (6) 0.729 (16) 0.521 (18) 0.403 (16) 0.479 (18) 0.534 (11) 0.456 (15) 0.745 (13) 0.693 (5) 0.752 (11) 0.554 (15) 0.587 (13) 0.567 (11)
0.922 (10) 0.708 (14) 0.705 (1) 0.947 (2) 0.745 (15) 0.640 (12) 0.995 (13) 0.673 (5) 0.927 (1) 0.523 (13) 0.561 (5) 0.790 (4) 0.793 (15) 0.730 (4) 0.852 (8) 0.803 (6) 0.869 (12) 0.567 (15) 0.898 (4) 0.888 (4) 0.717 (2) 0.591 (16) 0.616 (13) 0.976 (9) 0.983 (3) 0.717 (12) 0.889 (5) 0.556 (5) 0.509 (9) 0.874 (4) 0.916 (11) 0.364 (18) 0.887 (3) 0.820 (9) 0.721 (7) 0.969 (9) 0.956 (12) 0.323 (17) 0.732 (12) 0.491 (16) 0.393 (18) 0.978 (3) 0.705 (4) 0.904 (3) 0.709 (2) 0.624 (10) 0.686 (4)
0.972 (6) 0.863 (7) 0.515 (15) 0.871 (13) 0.876 (7) 0.655 (11) 1.000 (7) 0.655 (7) 0.675 (11) 0.580 (11) 0.493 (14) 0.562 (13) 0.951 (8) 0.625 (11) 0.838 (12) 0.718 (9) 0.923 (9) 0.880 (6) 0.796 (10) 0.773 (13) 0.635 (4) 0.788 (7) 0.646 (11) 0.931 (15) 0.966 (9) 0.685 (14) 0.717 (13) 0.582 (3) 0.405 (18) 0.800 (12) 0.960 (8) 0.623 (7) 0.814 (8) 0.742 (14) 0.673 (9) 0.988 (2) 0.992 (4) 0.612 (10) 0.912 (9) 0.542 (9) 0.497 (12) 0.936 (6) 0.727 (2) 0.702 (12) 0.659 (8) 0.708 (5) 0.749 (2)
0.969 (7) 0.941 (5) 0.615 (7) 0.943 (3) 0.918 (4) 0.813 (1) 0.999 (10) 0.785 (1) 0.839 (8) 0.509 (15) 0.526 (10) 0.648 (9) 0.997 (2) 0.715 (7) 0.894 (5) 0.829 (3) 0.992 (4) 0.924 (4) 0.950 (1) 0.929 (1) 0.720 (1) 0.865 (2) 0.691 (8) 0.982 (8) 0.982 (4) 0.635 (15) 0.802 (10) 0.607 (2) 0.514 (8) 0.879 (3) 0.988 (2) 0.562 (14) 0.718 (15) 0.830 (8) 0.697 (8) 0.986 (4) 0.989 (5) 0.588 (11) 0.908 (10) 0.503 (13) 0.435 (17) 0.981 (2) 0.793 (1) 0.915 (2) 0.700 (3) 0.822 (2) 0.765 (1)
0.919 (11) 0.843 (11) 0.482 (16) 0.849 (14) 0.905 (6) 0.782 (3) 1.000 (4) 0.462 (16) 0.655 (12) 0.674 (1) 0.479 (18) 0.534 (16) 0.839 (11) 0.462 (18) 0.651 (15) 0.560 (17) 0.948 (6) 0.891 (5) 0.493 (18) 0.782 (12) 0.386 (18) 0.692 (13) 0.718 (6) 0.956 (11) 0.878 (15) 0.758 (9) 0.683 (16) 0.302 (18) 0.527 (6) 0.852 (8) 0.768 (17) 0.696 (1) 0.559 (18) 0.744 (13) 0.760 (2) 0.812 (15) 0.972 (11) 0.692 (4) 0.956 (7) 0.614 (5) 0.608 (2) 0.463 (18) 0.498 (18) 0.497 (17) 0.480 (18) 0.531 (17) 0.467 (18)
0.853 (12) 0.855 (8) 0.469 (18) 0.740 (17) 0.823 (12) 0.786 (2) 0.994 (14) 0.592 (10) 0.702 (10) 0.515 (14) 0.510 (12) 0.502 (18) 0.973 (7) 0.610 (12) 0.876 (6) 0.671 (12) 0.882 (10) 0.666 (13) 0.811 (7) 0.659 (16) 0.634 (5) 0.690 (15) 0.542 (18) 0.944 (13) 0.932 (11) 0.802 (6) 0.708 (14) 0.521 (9) 0.553 (3) 0.817 (10) 0.902 (13) 0.613 (8) 0.688 (17) 0.777 (12) 0.567 (16) 0.993 (1) 0.982 (7) 0.650 (6) 0.992 (4) 0.620 (4) 0.614 (1) 0.995 (1) 0.569 (15) 0.768 (10) 0.652 (10) 0.569 (14) 0.683 (5)
0.982 (5) 0.534 (18) 0.690 (2) 0.904 (10) 0.950 (1) 0.682 (6) 0.999 (11) 0.642 (8) 0.867 (6) 0.625 (6) 0.575 (3) 0.876 (3) 0.976 (6) 0.745 (3) 0.849 (10) 0.798 (7) 0.859 (13) 0.858 (8) 0.947 (3) 0.864 (5) 0.697 (3) 0.859 (3) 0.752 (1) 0.993 (3) 0.903 (14) 0.688 (13) 0.922 (2) 0.559 (4) 0.506 (11) 0.886 (2) 0.968 (7) 0.644 (4) 0.843 (7) 0.971 (1) 0.731 (5) 0.985 (5) 0.994 (3) 0.910 (1) 0.967 (6) 0.553 (8) 0.503 (11) 0.951 (4) 0.642 (6) 0.790 (6) 0.656 (9) 0.325 (18) 0.609 (7)
0.952 (8) 0.956 (3) 0.621 (6) 0.940 (4) 0.934 (2) 0.778 (4) 1.000 (8) 0.760 (2) 0.860 (7) 0.572 (12) 0.758 (1) 0.882 (1) 0.988 (5) 0.777 (1) 0.905 (3) 0.933 (1) 0.999 (2) 0.971 (3) 0.948 (2) 0.925 (2) 0.615 (8) 0.878 (1) 0.727 (4) 0.997 (1) 0.991 (1) 0.750 (10) 0.917 (3) 0.504 (13) 0.490 (14) 0.868 (5) 0.986 (3) 0.583 (12) 0.948 (1) 0.959 (2) 0.825 (1) 0.982 (6) 0.977 (10) 0.880 (3) 0.943 (8) 0.562 (7) 0.539 (8) 0.950 (5) 0.641 (7) 0.847 (5) 0.681 (4) 0.823 (1) 0.746 (3)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.764 (9) 0.747 (10) 0.776 (7) 0.689 (7) 0.536 (8)
0.769 (8) 0.749 (9) 0.693 (14) 0.583 (17) 0.533 (11)
0.674 (16) 0.668 (15) 0.646 (17) 0.667 (9) 0.523 (15)
0.847 (5) 0.821 (7) 0.680 (15) 0.753 (4) 0.514 (17)
0.744 (12) 0.825 (6) 0.652 (16) 0.841 (1) 0.510 (18)
0.681 (15) 0.661 (16) 0.767 (8) 0.603 (14) 0.562 (2)
0.615 (17) 0.616 (17) 0.747 (10) 0.564 (18) 0.533 (12)
0.722 (13) 0.726 (12) 0.805 (3) 0.594 (15) 0.534 (9)
0.486 (18) 0.536 (18) 0.590 (18) 0.586 (16) 0.534 (10)
0.834 (6) 0.819 (8) 0.817 (2) 0.751 (5) 0.538 (7)
0.827 (7) 0.849 (5) 0.787 (6) 0.658 (10) 0.543 (4)
0.880 (4) 0.893 (3) 0.716 (11) 0.654 (11) 0.567 (1)
0.699 (14) 0.689 (14) 0.703 (13) 0.627 (12) 0.531 (14)
0.888 (2) 0.897 (2) 0.790 (5) 0.833 (2) 0.542 (5)
0.753 (10) 0.734 (11) 0.794 (4) 0.617 (13) 0.532 (13)
0.746 (11) 0.697 (13) 0.704 (12) 0.705 (6) 0.519 (16)
0.886 (3) 0.871 (4) 0.761 (9) 0.758 (3) 0.546 (3)
0.903 (1) 0.906 (1) 0.838 (1) 0.688 (8) 0.542 (6)
20news agnews amazon imdb yelp
0.661 (11) 0.681 (10) 0.583 (7) 0.564 (8) 0.623 (8)
0.588 (17) 0.618 (14) 0.517 (15) 0.524 (12) 0.530 (16)
0.617 (15) 0.594 (16) 0.559 (10) 0.518 (16) 0.606 (11)
0.479 (18) 0.665 (11) 0.508 (17) 0.456 (18) 0.481 (18)
0.702 (5) 0.740 (5) 0.492 (18) 0.465 (17) 0.513 (17)
0.655 (12) 0.659 (12) 0.649 (1) 0.621 (4) 0.679 (5)
0.643 (13) 0.649 (13) 0.569 (9) 0.584 (7) 0.638 (7)
0.695 (7) 0.700 (8) 0.643 (2) 0.627 (2) 0.739 (1)
0.615 (16) 0.574 (17) 0.633 (3) 0.589 (6) 0.594 (12)
0.731 (2) 0.757 (4) 0.628 (4) 0.622 (3) 0.724 (3)
0.670 (10) 0.772 (3) 0.542 (12) 0.562 (9) 0.575 (13)
0.726 (3) 0.849 (1) 0.530 (13) 0.523 (14) 0.565 (14)
0.620 (14) 0.602 (15) 0.570 (8) 0.559 (10) 0.609 (10)
0.698 (6) 0.736 (6) 0.549 (11) 0.519 (15) 0.621 (9)
0.691 (9) 0.695 (9) 0.613 (5) 0.632 (1) 0.734 (2)
0.695 (8) 0.567 (18) 0.510 (16) 0.555 (11) 0.677 (6)
0.705 (4) 0.734 (7) 0.518 (14) 0.524 (13) 0.563 (15)
0.743 (1) 0.802 (2) 0.608 (6) 0.621 (5) 0.704 (4)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 52: The average of AUCROC performance on individual datasets under 𝑁𝑙𝑎 = 3 setting Dataset
XGBOD
DeepSAD
REPEN
AABiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
DualMGAN
SOEL-NTL
AnoDDAE
XGB
CatB
TabMCls
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.973 (8) 0.933 (11) 0.562 (12) 0.940 (7) 0.974 (4) 0.726 (12) 0.974 (16) 0.653 (10) 0.743 (12) 0.677 (9) 0.581 (3) 0.830 (4) 0.883 (13) 0.672 (10) 0.837 (12) 0.861 (5) 0.915 (13) 0.945 (8) 0.889 (5) 0.943 (9) 0.643 (6) 0.873 (4) 0.688 (6) 0.982 (5) 0.965 (12) 0.698 (16) 0.990 (7) 0.595 (4) 0.557 (8) 0.896 (8) 0.967 (8) 0.587 (14) 0.880 (12) 0.937 (5) 0.794 (8) 0.821 (16) 0.972 (11) 0.727 (7) 0.923 (11) 0.555 (11) 0.579 (7) 0.891 (9) 0.729 (4) 0.882 (8) 0.744 (8) 0.812 (6) 0.745 (4)
0.930 (11) 0.938 (10) 0.680 (2) 0.938 (8) 0.897 (11) 0.690 (14) 0.999 (11) 0.689 (3) 0.774 (10) 0.684 (8) 0.576 (4) 0.731 (6) 0.870 (15) 0.729 (5) 0.864 (10) 0.719 (12) 0.919 (12) 0.885 (11) 0.868 (7) 0.939 (10) 0.583 (12) 0.708 (13) 0.674 (7) 0.982 (4) 0.985 (6) 0.897 (9) 0.995 (4) 0.522 (14) 0.537 (10) 0.730 (15) 0.947 (9) 0.607 (13) 0.900 (5) 0.762 (16) 0.703 (15) 0.954 (10) 0.938 (14) 0.649 (12) 0.733 (15) 0.523 (15) 0.487 (14) 0.894 (8) 0.595 (14) 0.749 (15) 0.631 (13) 0.620 (15) 0.583 (12)
0.914 (13) 0.956 (9) 0.613 (7) 0.911 (16) 0.934 (8) 0.764 (9) 1.000 (9) 0.658 (9) 0.844 (7) 0.535 (14) 0.537 (9) 0.685 (8) 0.952 (9) 0.642 (11) 0.897 (5) 0.744 (11) 0.969 (11) 0.769 (13) 0.885 (6) 0.739 (15) 0.656 (4) 0.788 (7) 0.632 (12) 0.958 (14) 0.986 (5) 0.906 (6) 0.874 (15) 0.544 (11) 0.523 (11) 0.896 (7) 0.971 (7) 0.717 (4) 0.893 (8) 0.869 (10) 0.572 (17) 0.824 (15) 0.936 (15) 0.642 (13) 0.903 (12) 0.557 (10) 0.580 (6) 0.843 (12) 0.700 (7) 0.787 (12) 0.620 (14) 0.769 (9) 0.681 (7)
0.896 (14) 0.647 (16) 0.467 (18) 0.804 (18) 0.600 (18) 0.738 (11) 0.996 (13) 0.608 (11) 0.782 (9) 0.359 (18) 0.530 (10) 0.411 (18) 0.872 (14) 0.453 (18) 0.264 (17) 0.877 (4) 0.652 (18) 0.639 (15) 0.631 (16) 0.786 (14) 0.479 (18) 0.559 (17) 0.582 (16) 0.942 (15) 0.920 (14) 0.572 (17) 0.927 (10) 0.437 (17) 0.466 (18) 0.664 (17) 0.741 (17) 0.557 (15) 0.298 (18) 0.444 (18) 0.574 (16) 0.958 (9) 0.880 (16) 0.477 (15) 0.370 (18) 0.460 (18) 0.444 (17) 0.798 (13) 0.645 (12) 0.852 (9) 0.812 (5) 0.780 (8) 0.423 (18)
0.491 (18) 0.552 (17) 0.642 (4) 0.964 (1) 0.721 (17) 0.606 (16) 0.821 (17) 0.670 (7) 0.900 (2) 0.495 (16) 0.542 (7) 0.701 (7) 0.939 (10) 0.617 (13) 0.779 (14) 0.716 (15) 0.825 (16) 0.432 (18) 0.762 (11) 0.682 (17) 0.605 (10) 0.774 (9) 0.701 (5) 0.970 (10) 0.669 (18) 0.475 (18) 0.575 (18) 0.528 (12) 0.506 (16) 0.750 (14) 0.918 (12) 0.375 (18) 0.884 (11) 0.782 (14) 0.534 (18) 0.913 (12) 0.980 (9) 0.465 (16) 0.621 (16) 0.484 (17) 0.452 (16) 0.921 (6) 0.675 (9) 0.787 (11) 0.560 (17) 0.669 (14) 0.635 (8)
0.996 (6) 0.998 (1) 0.576 (10) 0.931 (10) 0.863 (13) 0.816 (5) 1.000 (1) 0.567 (13) 0.563 (14) 0.713 (3) 0.492 (15) 0.573 (16) 0.977 (8) 0.729 (6) 0.908 (3) 0.760 (10) 1.000 (1) 0.999 (1) 0.732 (14) 0.978 (5) 0.632 (7) 0.694 (14) 0.658 (10) 0.976 (9) 0.980 (10) 0.952 (2) 0.921 (12) 0.583 (6) 0.601 (2) 0.925 (6) 0.994 (3) 0.648 (12) 0.904 (3) 0.946 (4) 0.830 (6) 0.990 (1) 1.000 (1) 0.655 (11) 1.000 (1) 0.665 (1) 0.616 (3) 0.846 (11) 0.657 (10) 0.933 (4) 0.815 (4) 0.845 (2) 0.599 (10)
0.997 (4) 0.977 (6) 0.544 (13) 0.911 (17) 0.880 (12) 0.691 (13) 1.000 (2) 0.563 (14) 0.544 (15) 0.626 (12) 0.539 (8) 0.663 (10) 0.917 (12) 0.542 (17) 0.882 (8) 0.687 (17) 1.000 (4) 0.988 (5) 0.719 (15) 0.960 (8) 0.543 (15) 0.770 (10) 0.609 (15) 0.979 (6) 0.980 (9) 0.969 (1) 0.988 (8) 0.550 (10) 0.498 (17) 0.835 (12) 0.933 (10) 0.700 (5) 0.904 (4) 0.885 (8) 0.726 (11) 0.907 (13) 1.000 (3) 0.701 (9) 1.000 (2) 0.573 (8) 0.520 (12) 0.617 (17) 0.543 (17) 0.747 (16) 0.599 (15) 0.738 (11) 0.585 (11)
0.996 (5) 0.997 (2) 0.566 (11) 0.932 (9) 0.924 (10) 0.768 (8) 1.000 (3) 0.546 (15) 0.543 (16) 0.710 (4) 0.488 (17) 0.580 (13) 0.987 (6) 0.685 (9) 0.894 (6) 0.716 (14) 1.000 (2) 0.998 (2) 0.732 (13) 0.972 (6) 0.617 (8) 0.738 (12) 0.664 (9) 0.966 (11) 0.981 (8) 0.951 (3) 0.873 (16) 0.556 (9) 0.611 (1) 0.944 (3) 0.995 (1) 0.684 (7) 0.890 (9) 0.959 (3) 0.821 (7) 0.987 (2) 1.000 (2) 0.663 (10) 1.000 (3) 0.650 (2) 0.617 (2) 0.764 (15) 0.649 (11) 0.909 (6) 0.775 (7) 0.816 (5) 0.631 (9)
0.705 (17) 0.783 (14) 0.484 (16) 0.917 (14) 0.792 (15) 0.491 (18) 1.000 (10) 0.497 (16) 0.459 (18) 0.639 (11) 0.487 (18) 0.579 (14) 0.720 (17) 0.548 (16) 0.725 (16) 0.602 (18) 0.764 (17) 0.686 (14) 0.602 (17) 0.637 (18) 0.481 (17) 0.781 (8) 0.544 (18) 0.760 (18) 0.861 (17) 0.888 (10) 0.995 (5) 0.480 (16) 0.513 (14) 0.662 (18) 0.723 (18) 0.738 (3) 0.629 (17) 0.774 (15) 0.784 (9) 0.655 (18) 0.785 (17) 0.961 (3) 0.789 (13) 0.617 (6) 0.580 (5) 0.486 (18) 0.519 (18) 0.464 (18) 0.568 (16) 0.610 (17) 0.527 (17)
0.997 (3) 0.966 (7) 0.585 (9) 0.951 (2) 0.933 (9) 0.748 (10) 0.997 (12) 0.607 (12) 0.846 (6) 0.691 (7) 0.528 (11) 0.600 (11) 0.997 (2) 0.756 (3) 0.888 (7) 0.853 (6) 0.993 (6) 0.984 (7) 0.854 (8) 0.983 (2) 0.579 (13) 0.757 (11) 0.783 (1) 0.996 (2) 0.989 (4) 0.933 (5) 0.928 (9) 0.624 (2) 0.573 (5) 0.953 (1) 0.994 (2) 0.651 (11) 0.898 (6) 0.936 (6) 0.838 (5) 0.982 (6) 0.980 (10) 0.613 (14) 0.991 (6) 0.638 (5) 0.572 (9) 0.903 (7) 0.596 (13) 0.958 (1) 0.833 (2) 0.727 (12) 0.574 (13)
0.774 (16) 0.464 (18) 0.654 (3) 0.921 (13) 0.837 (14) 0.556 (17) 0.578 (18) 0.357 (18) 0.888 (3) 0.655 (10) 0.549 (6) 0.874 (3) 0.319 (18) 0.723 (7) 0.256 (18) 0.848 (7) 0.978 (10) 0.546 (17) 0.733 (12) 0.805 (13) 0.616 (9) 0.540 (18) 0.759 (2) 0.892 (17) 0.917 (15) 0.711 (14) 0.705 (17) 0.482 (15) 0.518 (13) 0.708 (16) 0.835 (16) 0.538 (16) 0.895 (7) 0.880 (9) 0.722 (13) 0.797 (17) 0.531 (18) 0.374 (17) 0.495 (17) 0.546 (14) 0.452 (15) 0.795 (14) 0.689 (8) 0.758 (14) 0.553 (18) 0.545 (18) 0.567 (14)
0.926 (12) 0.690 (15) 0.699 (1) 0.948 (3) 0.747 (16) 0.641 (15) 0.995 (14) 0.671 (6) 0.934 (1) 0.518 (15) 0.561 (5) 0.790 (5) 0.837 (16) 0.733 (4) 0.856 (11) 0.807 (8) 0.865 (15) 0.579 (16) 0.903 (4) 0.888 (12) 0.728 (1) 0.599 (16) 0.627 (13) 0.977 (8) 0.983 (7) 0.699 (15) 0.889 (14) 0.557 (8) 0.508 (15) 0.866 (10) 0.927 (11) 0.383 (17) 0.887 (10) 0.826 (13) 0.722 (12) 0.970 (8) 0.953 (13) 0.325 (18) 0.743 (14) 0.494 (16) 0.386 (18) 0.978 (1) 0.703 (6) 0.904 (7) 0.711 (11) 0.614 (16) 0.685 (6)
0.947 (10) 0.961 (8) 0.525 (15) 0.947 (4) 0.938 (7) 0.790 (7) 0.974 (15) 0.666 (8) 0.756 (11) 0.564 (13) 0.527 (12) 0.591 (12) 0.933 (11) 0.634 (12) 0.812 (13) 0.768 (9) 0.991 (9) 0.917 (10) 0.840 (9) 0.896 (11) 0.585 (11) 0.819 (5) 0.637 (11) 0.927 (16) 0.971 (11) 0.831 (11) 0.997 (3) 0.585 (5) 0.558 (7) 0.852 (11) 0.890 (13) 0.676 (9) 0.871 (13) 0.827 (12) 0.761 (10) 0.825 (14) 0.967 (12) 0.706 (8) 0.942 (9) 0.550 (13) 0.512 (13) 0.865 (10) 0.704 (5) 0.763 (13) 0.728 (9) 0.791 (7) 0.810 (2)
0.972 (9) 0.978 (5) 0.611 (8) 0.946 (5) 0.989 (1) 0.882 (1) 1.000 (7) 0.782 (1) 0.851 (5) 0.709 (5) 0.527 (13) 0.673 (9) 0.997 (1) 0.704 (8) 0.905 (4) 0.902 (2) 0.991 (8) 0.985 (6) 0.949 (3) 0.981 (3) 0.716 (2) 0.923 (2) 0.672 (8) 0.978 (7) 0.991 (3) 0.905 (7) 1.000 (1) 0.568 (7) 0.543 (9) 0.946 (2) 0.992 (4) 0.680 (8) 0.852 (14) 0.907 (7) 0.840 (4) 0.986 (3) 0.999 (5) 0.728 (6) 0.982 (7) 0.554 (12) 0.526 (11) 0.976 (2) 0.816 (1) 0.956 (2) 0.837 (1) 0.838 (3) 0.812 (1)
0.992 (7) 0.932 (12) 0.526 (14) 0.916 (15) 0.975 (3) 0.802 (6) 1.000 (4) 0.449 (17) 0.505 (17) 0.756 (1) 0.495 (14) 0.574 (15) 0.977 (7) 0.575 (14) 0.766 (15) 0.687 (16) 0.993 (7) 0.935 (9) 0.560 (18) 0.968 (7) 0.520 (16) 0.648 (15) 0.612 (14) 0.961 (12) 0.932 (13) 0.943 (4) 0.921 (11) 0.428 (18) 0.598 (3) 0.874 (9) 0.855 (14) 0.760 (1) 0.835 (15) 0.859 (11) 0.859 (2) 0.983 (5) 0.998 (7) 0.757 (5) 0.996 (4) 0.647 (4) 0.616 (4) 0.670 (16) 0.563 (16) 0.659 (17) 0.684 (12) 0.836 (4) 0.548 (16)
0.895 (15) 0.910 (13) 0.475 (17) 0.924 (12) 0.962 (6) 0.817 (4) 1.000 (6) 0.673 (4) 0.685 (13) 0.485 (17) 0.488 (16) 0.564 (17) 0.989 (5) 0.567 (15) 0.879 (9) 0.717 (13) 0.907 (14) 0.777 (12) 0.767 (10) 0.709 (16) 0.566 (14) 0.791 (6) 0.548 (17) 0.960 (13) 0.911 (16) 0.801 (13) 0.912 (13) 0.525 (13) 0.522 (12) 0.826 (13) 0.846 (15) 0.673 (10) 0.762 (16) 0.728 (17) 0.720 (14) 0.945 (11) 0.980 (8) 0.888 (4) 0.939 (10) 0.648 (3) 0.635 (1) 0.953 (3) 0.582 (15) 0.793 (10) 0.716 (10) 0.679 (13) 0.549 (15)
0.998 (2) 0.982 (4) 0.639 (5) 0.941 (6) 0.979 (2) 0.820 (3) 1.000 (8) 0.673 (5) 0.804 (8) 0.735 (2) 0.613 (2) 0.884 (2) 0.992 (4) 0.774 (1) 0.915 (2) 0.884 (3) 0.999 (5) 0.991 (4) 0.955 (2) 0.980 (4) 0.685 (3) 0.948 (1) 0.746 (3) 0.996 (3) 0.997 (1) 0.818 (12) 0.993 (6) 0.595 (3) 0.582 (4) 0.941 (4) 0.989 (5) 0.758 (2) 0.953 (2) 0.988 (1) 0.850 (3) 0.986 (4) 1.000 (4) 0.965 (2) 0.995 (5) 0.614 (7) 0.557 (10) 0.948 (4) 0.774 (2) 0.932 (5) 0.802 (6) 0.765 (10) 0.710 (5)
0.999 (1) 0.989 (3) 0.621 (6) 0.929 (11) 0.971 (5) 0.875 (2) 1.000 (5) 0.755 (2) 0.869 (4) 0.708 (6) 0.723 (1) 0.915 (1) 0.994 (3) 0.769 (2) 0.916 (1) 0.975 (1) 1.000 (3) 0.997 (3) 0.962 (1) 0.988 (1) 0.653 (5) 0.915 (3) 0.707 (4) 0.998 (1) 0.995 (2) 0.900 (8) 1.000 (2) 0.628 (1) 0.562 (6) 0.936 (5) 0.988 (6) 0.687 (6) 0.965 (1) 0.981 (2) 0.889 (1) 0.976 (7) 0.999 (6) 0.966 (1) 0.979 (8) 0.573 (9) 0.576 (8) 0.925 (5) 0.733 (3) 0.944 (3) 0.829 (3) 0.890 (1) 0.793 (3)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.838 (8) 0.859 (7) 0.836 (8) 0.735 (6) 0.545 (11)
0.780 (12) 0.779 (14) 0.748 (13) 0.612 (16) 0.539 (14)
0.752 (14) 0.745 (16) 0.726 (14) 0.734 (7) 0.544 (12)
0.875 (6) 0.884 (6) 0.705 (17) 0.706 (10) 0.525 (16)
0.744 (16) 0.821 (12) 0.642 (18) 0.841 (2) 0.511 (18)
0.790 (11) 0.806 (13) 0.868 (5) 0.656 (13) 0.583 (1)
0.719 (17) 0.739 (17) 0.860 (6) 0.596 (18) 0.556 (6)
0.813 (9) 0.836 (10) 0.877 (3) 0.649 (14) 0.570 (2)
0.490 (18) 0.565 (18) 0.761 (12) 0.601 (17) 0.544 (13)
0.899 (4) 0.890 (5) 0.880 (2) 0.779 (4) 0.546 (8)
0.745 (15) 0.834 (11) 0.780 (11) 0.627 (15) 0.536 (15)
0.880 (5) 0.893 (4) 0.716 (16) 0.664 (12) 0.567 (3)
0.807 (10) 0.836 (9) 0.816 (10) 0.725 (8) 0.546 (10)
0.904 (3) 0.902 (2) 0.835 (9) 0.848 (1) 0.546 (9)
0.844 (7) 0.846 (8) 0.877 (4) 0.687 (11) 0.566 (4)
0.767 (13) 0.748 (15) 0.717 (15) 0.713 (9) 0.525 (17)
0.915 (2) 0.901 (3) 0.856 (7) 0.805 (3) 0.563 (5)
0.935 (1) 0.929 (1) 0.886 (1) 0.739 (5) 0.551 (7)
20news agnews amazon imdb yelp
0.736 (7) 0.740 (11) 0.617 (8) 0.626 (8) 0.730 (8)
0.614 (16) 0.663 (15) 0.530 (15) 0.528 (14) 0.562 (16)
0.665 (14) 0.637 (17) 0.573 (13) 0.589 (11) 0.618 (14)
0.555 (18) 0.663 (16) 0.495 (18) 0.467 (17) 0.553 (17)
0.698 (12) 0.750 (10) 0.506 (17) 0.470 (16) 0.517 (18)
0.760 (6) 0.781 (7) 0.733 (1) 0.734 (3) 0.856 (2)
0.708 (11) 0.698 (13) 0.621 (7) 0.623 (9) 0.740 (7)
0.764 (5) 0.789 (5) 0.709 (3) 0.750 (2) 0.819 (4)
0.670 (13) 0.681 (14) 0.665 (4) 0.697 (6) 0.788 (6)
0.783 (2) 0.824 (3) 0.665 (5) 0.709 (5) 0.836 (3)
0.662 (15) 0.750 (9) 0.571 (14) 0.578 (13) 0.649 (13)
0.727 (9) 0.848 (2) 0.527 (16) 0.524 (15) 0.570 (15)
0.716 (10) 0.721 (12) 0.616 (9) 0.629 (7) 0.717 (9)
0.735 (8) 0.773 (8) 0.593 (11) 0.579 (12) 0.711 (10)
0.771 (4) 0.811 (4) 0.713 (2) 0.764 (1) 0.861 (1)
0.598 (17) 0.598 (18) 0.601 (10) 0.454 (18) 0.680 (12)
0.772 (3) 0.785 (6) 0.593 (12) 0.623 (10) 0.693 (11)
0.793 (1) 0.867 (1) 0.657 (6) 0.734 (4) 0.817 (5)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 53: The average AUCROC performance on individual datasets under 𝑁𝑙𝑎 = 5 setting Dataset
XGBOD
DeepSAD
REPEN
AABiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
DualMGAN
SOEL-NTL
AnoDDAE
XGB
CatB
TabMCls
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.972 (10) 0.941 (12) 0.588 (12) 0.929 (12) 0.977 (4) 0.821 (7) 0.983 (14) 0.728 (4) 0.770 (11) 0.716 (6) 0.603 (3) 0.872 (3) 0.908 (14) 0.681 (9) 0.864 (11) 0.876 (5) 0.999 (9) 0.976 (9) 0.911 (4) 0.974 (10) 0.611 (11) 0.930 (4) 0.728 (7) 0.978 (8) 0.984 (6) 0.837 (14) 0.990 (7) 0.620 (8) 0.589 (11) 0.938 (6) 0.970 (8) 0.732 (12) 0.908 (8) 0.961 (7) 0.804 (8) 0.899 (15) 0.989 (10) 0.784 (8) 0.927 (12) 0.613 (10) 0.586 (5) 0.942 (7) 0.761 (4) 0.928 (6) 0.780 (8) 0.835 (7) 0.780 (4)
0.941 (12) 0.969 (8) 0.685 (2) 0.919 (13) 0.916 (9) 0.732 (12) 0.999 (11) 0.690 (6) 0.799 (9) 0.697 (9) 0.577 (4) 0.740 (6) 0.911 (13) 0.722 (6) 0.896 (9) 0.756 (14) 0.958 (13) 0.937 (11) 0.868 (7) 0.960 (11) 0.599 (13) 0.748 (14) 0.717 (9) 0.984 (4) 0.981 (9) 0.915 (8) 0.995 (5) 0.534 (12) 0.547 (13) 0.772 (14) 0.967 (10) 0.650 (14) 0.936 (3) 0.811 (15) 0.728 (12) 0.956 (11) 0.960 (14) 0.721 (11) 0.808 (14) 0.546 (15) 0.499 (14) 0.905 (10) 0.615 (14) 0.785 (15) 0.685 (13) 0.692 (13) 0.604 (11)
0.980 (8) 0.968 (9) 0.614 (9) 0.962 (3) 0.948 (8) 0.773 (10) 1.000 (9) 0.665 (11) 0.921 (2) 0.559 (14) 0.547 (8) 0.736 (7) 0.974 (10) 0.664 (11) 0.925 (4) 0.773 (12) 0.959 (12) 0.752 (14) 0.846 (8) 0.840 (15) 0.663 (7) 0.831 (9) 0.641 (14) 0.970 (11) 0.990 (4) 0.913 (9) 0.874 (15) 0.633 (6) 0.598 (10) 0.905 (8) 0.968 (9) 0.840 (3) 0.904 (9) 0.954 (9) 0.619 (16) 0.893 (16) 0.928 (15) 0.726 (10) 0.949 (10) 0.618 (9) 0.536 (13) 0.888 (12) 0.689 (9) 0.826 (12) 0.673 (15) 0.772 (11) 0.738 (6)
0.903 (15) 0.552 (18) 0.478 (17) 0.894 (16) 0.566 (18) 0.704 (14) 0.877 (16) 0.579 (12) 0.745 (12) 0.367 (18) 0.522 (12) 0.432 (18) 0.871 (15) 0.465 (18) 0.385 (17) 0.849 (6) 0.621 (18) 0.722 (15) 0.769 (12) 0.840 (16) 0.459 (18) 0.546 (17) 0.584 (16) 0.936 (15) 0.956 (14) 0.432 (18) 0.932 (10) 0.491 (17) 0.400 (18) 0.658 (18) 0.810 (18) 0.561 (15) 0.374 (18) 0.472 (18) 0.561 (17) 0.962 (10) 0.885 (16) 0.500 (16) 0.408 (18) 0.443 (18) 0.454 (16) 0.905 (9) 0.642 (11) 0.841 (10) 0.792 (7) 0.815 (9) 0.476 (18)
0.411 (18) 0.570 (17) 0.642 (5) 0.965 (2) 0.723 (17) 0.612 (17) 0.774 (17) 0.670 (9) 0.900 (3) 0.495 (16) 0.546 (9) 0.701 (9) 0.938 (12) 0.650 (12) 0.773 (16) 0.716 (17) 0.888 (14) 0.436 (18) 0.809 (11) 0.703 (18) 0.608 (12) 0.774 (12) 0.701 (12) 0.970 (12) 0.510 (18) 0.497 (17) 0.575 (18) 0.528 (14) 0.506 (16) 0.752 (15) 0.917 (13) 0.373 (18) 0.884 (12) 0.784 (17) 0.534 (18) 0.916 (14) 0.980 (12) 0.503 (15) 0.630 (16) 0.485 (17) 0.455 (15) 0.918 (8) 0.622 (13) 0.788 (14) 0.560 (17) 0.668 (15) 0.599 (12)
0.997 (6) 0.999 (2) 0.624 (7) 0.934 (11) 0.876 (13) 0.866 (4) 1.000 (1) 0.561 (13) 0.633 (14) 0.711 (8) 0.483 (17) 0.645 (14) 0.996 (7) 0.740 (4) 0.930 (1) 0.823 (8) 1.000 (1) 0.999 (1) 0.734 (15) 0.996 (3) 0.667 (6) 0.757 (13) 0.760 (5) 0.977 (9) 0.980 (10) 0.951 (4) 0.921 (12) 0.662 (4) 0.700 (1) 0.931 (7) 0.997 (2) 0.793 (8) 0.919 (6) 0.973 (4) 0.870 (6) 0.991 (1) 1.000 (1) 0.693 (13) 1.000 (1) 0.709 (2) 0.614 (3) 0.900 (11) 0.690 (8) 0.948 (5) 0.837 (4) 0.877 (3) 0.652 (9)
0.997 (4) 0.999 (1) 0.563 (15) 0.892 (17) 0.882 (12) 0.755 (11) 1.000 (2) 0.551 (15) 0.573 (16) 0.647 (13) 0.521 (13) 0.659 (10) 0.984 (9) 0.609 (15) 0.911 (7) 0.754 (15) 1.000 (3) 0.992 (5) 0.738 (14) 0.993 (6) 0.568 (15) 0.839 (8) 0.707 (11) 0.978 (7) 0.981 (7) 0.931 (6) 0.988 (8) 0.534 (11) 0.566 (12) 0.887 (10) 0.997 (1) 0.780 (10) 0.866 (15) 0.966 (5) 0.792 (10) 0.951 (12) 1.000 (2) 0.746 (9) 1.000 (2) 0.589 (12) 0.558 (10) 0.665 (17) 0.573 (16) 0.808 (13) 0.687 (12) 0.767 (12) 0.575 (13)
0.997 (5) 0.997 (3) 0.603 (11) 0.942 (6) 0.916 (10) 0.835 (6) 1.000 (4) 0.561 (14) 0.618 (15) 0.715 (7) 0.492 (16) 0.657 (11) 0.996 (5) 0.705 (7) 0.923 (5) 0.790 (11) 1.000 (4) 0.999 (2) 0.757 (13) 0.994 (5) 0.671 (5) 0.779 (11) 0.770 (4) 0.979 (6) 0.981 (8) 0.951 (3) 0.873 (16) 0.636 (5) 0.680 (3) 0.940 (5) 0.996 (5) 0.810 (6) 0.920 (5) 0.979 (3) 0.864 (7) 0.988 (5) 1.000 (3) 0.697 (12) 1.000 (3) 0.710 (1) 0.615 (2) 0.814 (14) 0.670 (10) 0.928 (7) 0.819 (6) 0.860 (4) 0.646 (10)
0.795 (17) 0.898 (13) 0.469 (18) 0.904 (15) 0.875 (14) 0.678 (15) 1.000 (10) 0.540 (16) 0.436 (18) 0.724 (5) 0.474 (18) 0.625 (16) 0.794 (17) 0.545 (17) 0.786 (15) 0.605 (18) 0.856 (17) 0.835 (12) 0.554 (18) 0.737 (17) 0.497 (17) 0.853 (6) 0.534 (18) 0.806 (18) 0.922 (15) 0.882 (11) 0.998 (3) 0.503 (16) 0.625 (6) 0.748 (16) 0.875 (17) 0.721 (13) 0.682 (17) 0.823 (14) 0.720 (14) 0.740 (18) 0.842 (17) 0.960 (4) 0.844 (13) 0.566 (13) 0.583 (7) 0.532 (18) 0.533 (18) 0.557 (18) 0.611 (16) 0.625 (16) 0.551 (16)
0.998 (3) 0.990 (6) 0.586 (13) 0.970 (1) 0.905 (11) 0.711 (13) 0.997 (12) 0.667 (10) 0.842 (7) 0.670 (11) 0.535 (11) 0.650 (12) 0.999 (1) 0.768 (3) 0.899 (8) 0.913 (3) 1.000 (2) 0.988 (7) 0.811 (10) 0.995 (4) 0.618 (10) 0.807 (10) 0.798 (1) 0.996 (3) 0.989 (5) 0.935 (5) 0.928 (11) 0.663 (3) 0.625 (7) 0.950 (3) 0.996 (4) 0.763 (11) 0.912 (7) 0.966 (6) 0.901 (2) 0.983 (7) 0.983 (11) 0.616 (14) 0.997 (6) 0.664 (5) 0.584 (6) 0.944 (6) 0.626 (12) 0.968 (1) 0.881 (1) 0.782 (10) 0.715 (7)
0.843 (16) 0.574 (16) 0.662 (3) 0.935 (10) 0.866 (15) 0.573 (18) 0.556 (18) 0.352 (18) 0.870 (5) 0.667 (12) 0.544 (10) 0.865 (4) 0.385 (18) 0.686 (8) 0.351 (18) 0.845 (7) 0.989 (11) 0.572 (16) 0.711 (16) 0.842 (14) 0.586 (14) 0.509 (18) 0.744 (6) 0.917 (16) 0.895 (17) 0.586 (16) 0.836 (17) 0.486 (18) 0.544 (14) 0.693 (17) 0.884 (16) 0.546 (16) 0.894 (10) 0.895 (11) 0.733 (11) 0.818 (17) 0.638 (18) 0.388 (17) 0.456 (17) 0.559 (14) 0.442 (17) 0.765 (15) 0.694 (7) 0.747 (16) 0.546 (18) 0.556 (18) 0.566 (15)
0.919 (13) 0.678 (15) 0.705 (1) 0.949 (5) 0.747 (16) 0.662 (16) 0.995 (13) 0.675 (7) 0.934 (1) 0.510 (15) 0.558 (5) 0.793 (5) 0.867 (16) 0.732 (5) 0.856 (12) 0.806 (10) 0.865 (16) 0.569 (17) 0.902 (5) 0.890 (12) 0.727 (1) 0.600 (16) 0.631 (15) 0.977 (10) 0.980 (11) 0.726 (15) 0.885 (14) 0.557 (10) 0.508 (15) 0.857 (11) 0.925 (11) 0.378 (17) 0.884 (13) 0.836 (13) 0.719 (15) 0.975 (8) 0.965 (13) 0.329 (18) 0.748 (15) 0.513 (16) 0.380 (18) 0.978 (2) 0.705 (6) 0.911 (8) 0.687 (11) 0.622 (17) 0.686 (8)
0.954 (11) 0.955 (11) 0.552 (16) 0.911 (14) 0.956 (7) 0.862 (5) 0.982 (15) 0.719 (5) 0.799 (10) 0.688 (10) 0.557 (6) 0.638 (15) 0.969 (11) 0.642 (13) 0.833 (13) 0.816 (9) 1.000 (6) 0.945 (10) 0.878 (6) 0.978 (9) 0.636 (8) 0.848 (7) 0.693 (13) 0.907 (17) 0.968 (13) 0.862 (13) 0.997 (4) 0.608 (9) 0.606 (9) 0.857 (12) 0.909 (15) 0.799 (7) 0.874 (14) 0.879 (12) 0.795 (9) 0.931 (13) 0.995 (9) 0.788 (7) 0.940 (11) 0.611 (11) 0.540 (12) 0.828 (13) 0.725 (5) 0.887 (9) 0.773 (9) 0.830 (8) 0.798 (3)
0.977 (9) 0.976 (7) 0.618 (8) 0.959 (4) 0.994 (1) 0.909 (1) 1.000 (8) 0.832 (2) 0.899 (4) 0.770 (2) 0.555 (7) 0.718 (8) 0.998 (2) 0.675 (10) 0.927 (3) 0.913 (4) 1.000 (8) 0.991 (6) 0.944 (3) 0.993 (7) 0.717 (2) 0.950 (2) 0.724 (8) 0.984 (5) 0.991 (3) 0.955 (2) 1.000 (1) 0.632 (7) 0.616 (8) 0.954 (1) 0.993 (7) 0.810 (5) 0.926 (4) 0.933 (10) 0.881 (5) 0.990 (4) 1.000 (4) 0.844 (5) 0.995 (7) 0.629 (8) 0.554 (11) 0.981 (1) 0.831 (1) 0.965 (2) 0.827 (5) 0.836 (6) 0.829 (1)
0.997 (7) 0.958 (10) 0.603 (10) 0.940 (8) 0.971 (5) 0.820 (8) 1.000 (5) 0.448 (17) 0.481 (17) 0.773 (1) 0.495 (15) 0.645 (13) 0.996 (6) 0.638 (14) 0.830 (14) 0.762 (13) 0.997 (10) 0.982 (8) 0.601 (17) 0.990 (8) 0.566 (16) 0.690 (15) 0.716 (10) 0.954 (13) 0.977 (12) 0.964 (1) 0.920 (13) 0.519 (15) 0.683 (2) 0.897 (9) 0.918 (12) 0.852 (2) 0.884 (11) 0.958 (8) 0.894 (4) 0.987 (6) 1.000 (7) 0.812 (6) 1.000 (4) 0.691 (3) 0.612 (4) 0.738 (16) 0.551 (17) 0.735 (17) 0.751 (10) 0.905 (1) 0.573 (14)
0.912 (14) 0.887 (14) 0.566 (14) 0.878 (18) 0.971 (6) 0.812 (9) 1.000 (6) 0.674 (8) 0.692 (13) 0.487 (17) 0.509 (14) 0.556 (17) 0.995 (8) 0.573 (16) 0.879 (10) 0.718 (16) 0.882 (15) 0.762 (13) 0.826 (9) 0.868 (13) 0.628 (9) 0.883 (5) 0.550 (17) 0.948 (14) 0.901 (16) 0.888 (10) 0.947 (9) 0.529 (13) 0.495 (17) 0.842 (13) 0.916 (14) 0.840 (4) 0.805 (16) 0.799 (16) 0.724 (13) 0.973 (9) 0.998 (8) 0.970 (3) 0.953 (9) 0.642 (6) 0.616 (1) 0.948 (5) 0.581 (15) 0.837 (11) 0.681 (14) 0.685 (14) 0.512 (17)
0.999 (2) 0.996 (4) 0.651 (4) 0.939 (9) 0.986 (2) 0.883 (2) 1.000 (7) 0.741 (3) 0.800 (8) 0.757 (4) 0.619 (2) 0.917 (2) 0.996 (4) 0.792 (2) 0.922 (6) 0.930 (2) 1.000 (7) 0.997 (4) 0.955 (2) 0.996 (2) 0.710 (3) 0.973 (1) 0.797 (2) 0.997 (2) 0.999 (1) 0.866 (12) 0.993 (6) 0.665 (2) 0.646 (5) 0.951 (2) 0.996 (3) 0.857 (1) 0.972 (2) 0.994 (1) 0.898 (3) 0.990 (3) 1.000 (5) 0.976 (1) 0.998 (5) 0.673 (4) 0.579 (8) 0.973 (3) 0.811 (2) 0.958 (4) 0.850 (3) 0.854 (5) 0.765 (5)
0.999 (1) 0.991 (5) 0.641 (6) 0.940 (7) 0.978 (3) 0.881 (3) 1.000 (3) 0.838 (1) 0.868 (6) 0.766 (3) 0.705 (1) 0.935 (1) 0.996 (3) 0.804 (1) 0.927 (2) 0.979 (1) 1.000 (5) 0.998 (3) 0.959 (1) 0.996 (1) 0.704 (4) 0.945 (3) 0.774 (3) 0.998 (1) 0.997 (2) 0.918 (7) 1.000 (2) 0.720 (1) 0.654 (4) 0.942 (4) 0.995 (6) 0.792 (9) 0.983 (1) 0.992 (2) 0.920 (1) 0.991 (2) 1.000 (6) 0.971 (2) 0.992 (8) 0.631 (7) 0.576 (9) 0.965 (4) 0.768 (3) 0.959 (3) 0.865 (2) 0.905 (2) 0.815 (2)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.884 (6) 0.880 (7) 0.861 (9) 0.792 (5) 0.562 (11)
0.801 (12) 0.797 (14) 0.769 (13) 0.637 (17) 0.543 (14)
0.785 (14) 0.787 (16) 0.775 (12) 0.760 (8) 0.552 (13)
0.879 (8) 0.879 (9) 0.717 (16) 0.701 (13) 0.527 (17)
0.747 (16) 0.818 (12) 0.651 (18) 0.841 (2) 0.514 (18)
0.856 (10) 0.861 (11) 0.902 (3) 0.728 (9) 0.599 (1)
0.761 (15) 0.812 (13) 0.903 (2) 0.646 (15) 0.577 (5)
0.849 (11) 0.880 (8) 0.902 (4) 0.706 (12) 0.587 (2)
0.650 (18) 0.690 (18) 0.846 (11) 0.645 (16) 0.570 (6)
0.901 (4) 0.899 (4) 0.881 (7) 0.794 (4) 0.564 (10)
0.682 (17) 0.792 (15) 0.750 (14) 0.603 (18) 0.541 (15)
0.879 (7) 0.894 (5) 0.715 (17) 0.668 (14) 0.568 (7)
0.871 (9) 0.872 (10) 0.858 (10) 0.785 (6) 0.564 (9)
0.917 (3) 0.917 (3) 0.863 (8) 0.858 (1) 0.561 (12)
0.887 (5) 0.894 (6) 0.899 (5) 0.726 (10) 0.582 (4)
0.798 (13) 0.740 (17) 0.745 (15) 0.726 (11) 0.536 (16)
0.931 (2) 0.919 (2) 0.892 (6) 0.834 (3) 0.586 (3)
0.945 (1) 0.939 (1) 0.904 (1) 0.766 (7) 0.566 (8)
20news agnews amazon imdb yelp
0.771 (7) 0.787 (9) 0.629 (12) 0.663 (10) 0.758 (8)
0.635 (17) 0.691 (15) 0.546 (15) 0.555 (14) 0.592 (15)
0.706 (13) 0.682 (16) 0.608 (13) 0.571 (13) 0.629 (12)
0.498 (18) 0.652 (17) 0.506 (17) 0.442 (18) 0.525 (17)
0.698 (14) 0.761 (11) 0.486 (18) 0.443 (17) 0.501 (18)
0.809 (5) 0.850 (4) 0.775 (1) 0.751 (4) 0.866 (2)
0.731 (10) 0.756 (12) 0.673 (7) 0.678 (8) 0.751 (10)
0.808 (6) 0.836 (6) 0.764 (2) 0.775 (2) 0.849 (3)
0.713 (12) 0.723 (14) 0.716 (5) 0.710 (6) 0.791 (6)
0.811 (4) 0.865 (3) 0.725 (4) 0.738 (5) 0.837 (4)
0.638 (16) 0.726 (13) 0.629 (11) 0.591 (12) 0.621 (14)
0.729 (11) 0.848 (5) 0.529 (16) 0.522 (16) 0.568 (16)
0.763 (8) 0.779 (10) 0.638 (10) 0.675 (9) 0.755 (9)
0.762 (9) 0.820 (8) 0.643 (9) 0.634 (11) 0.746 (11)
0.813 (3) 0.887 (2) 0.754 (3) 0.793 (1) 0.869 (1)
0.663 (15) 0.634 (18) 0.585 (14) 0.542 (15) 0.625 (13)
0.820 (1) 0.835 (7) 0.662 (8) 0.688 (7) 0.766 (7)
0.814 (2) 0.893 (1) 0.715 (6) 0.764 (3) 0.833 (5)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 54: The average of AUCROC performance on individual datasets under 𝑁𝑙𝑎 = 10 setting Dataset
XGBOD
DeepSAD
REPEN
AABiGAN
GANomaly
DevNet
FEAWAD
PReNet
RoSAS
DualMGAN
SOEL-NTL
AnoDDAE
XGB
CatB
TabMCls
TabR-S
TabPFN
LimiX
10_cover 11_donors 12_fault 13_fraud 14_glass 15_Hepatitis 16_http 17_InternetAds 18_Ionosphere 19_landsat 1_ALOI 20_letter 21_Lymphography 22_magic.gamma 23_mammography 24_mnist 25_musk 26_optdigits 27_PageBlocks 28_pendigits 29_Pima 2_annthyroid 30_satellite 31_satimage-2 32_shuttle 33_skin 34_smtp 35_SpamBase 36_speech 37_Stamps 38_thyroid 39_vertebral 3_backdoor 40_vowels 41_Waveform 42_WBC 43_WDBC 44_Wilt 45_wine 46_WPBC 47_yeast 4_breastw 5_campaign 6_cardio 7_Cardiotocography 8_celeba 9_census
0.997 (8) 0.988 (8) 0.650 (7) 0.962 (4) 0.995 (4) 0.929 (3) 0.983 (15) 0.749 (5) 0.877 (7) 0.786 (4) 0.642 (2) 0.890 (3) 0.977 (13) 0.726 (8) 0.875 (12) 0.921 (4) 1.000 (8) 0.986 (8) 0.942 (5) 0.986 (11) 0.646 (12) 0.953 (4) 0.778 (5) 0.983 (8) 0.994 (3) 0.894 (14) 0.990 (7) 0.712 (6) 0.663 (5) 0.969 (2) 0.996 (7) 0.845 (6) 0.953 (6) 0.964 (8) 0.879 (8) 0.947 (14) 0.996 (11) 0.852 (7) 0.974 (12) 0.675 (9) 0.574 (10) 0.971 (7) 0.792 (4) 0.964 (7) 0.828 (7) 0.863 (6) 0.845 (2)
0.968 (12) 0.991 (7) 0.688 (4) 0.945 (10) 0.947 (11) 0.831 (12) 1.000 (10) 0.711 (8) 0.824 (11) 0.739 (7) 0.568 (5) 0.768 (8) 0.987 (12) 0.739 (6) 0.907 (8) 0.805 (13) 0.983 (12) 0.980 (9) 0.916 (6) 0.992 (9) 0.625 (14) 0.776 (13) 0.750 (8) 0.989 (7) 0.990 (5) 0.957 (5) 0.995 (5) 0.559 (13) 0.568 (14) 0.874 (13) 0.981 (11) 0.716 (14) 0.955 (5) 0.871 (14) 0.784 (11) 0.978 (9) 0.994 (13) 0.822 (9) 0.916 (13) 0.590 (14) 0.515 (13) 0.919 (14) 0.634 (13) 0.872 (14) 0.759 (11) 0.751 (13) 0.646 (12)
0.992 (10) 0.982 (10) 0.647 (8) 0.927 (14) 0.968 (9) 0.889 (9) 1.000 (9) 0.771 (4) 0.937 (1) 0.567 (14) 0.540 (11) 0.834 (5) 0.994 (11) 0.710 (9) 0.929 (4) 0.831 (11) 0.996 (11) 0.856 (12) 0.946 (4) 0.956 (12) 0.724 (6) 0.885 (9) 0.671 (14) 0.994 (4) 0.989 (6) 0.909 (13) 0.874 (15) 0.682 (8) 0.600 (11) 0.961 (6) 0.993 (9) 0.840 (7) 0.929 (10) 0.977 (7) 0.735 (12) 0.967 (12) 0.995 (12) 0.726 (11) 0.998 (7) 0.630 (12) 0.507 (14) 0.948 (10) 0.704 (8) 0.917 (10) 0.795 (10) 0.822 (11) 0.778 (6)
0.847 (17) 0.609 (16) 0.487 (17) 0.735 (18) 0.658 (18) 0.731 (15) 0.996 (13) 0.613 (12) 0.731 (12) 0.350 (18) 0.506 (13) 0.377 (18) 0.824 (17) 0.517 (18) 0.360 (17) 0.877 (7) 0.432 (18) 0.584 (16) 0.615 (18) 0.818 (16) 0.516 (18) 0.632 (16) 0.593 (16) 0.911 (18) 0.914 (16) 0.434 (18) 0.932 (10) 0.447 (18) 0.414 (18) 0.610 (18) 0.883 (17) 0.491 (16) 0.516 (18) 0.405 (18) 0.548 (17) 0.961 (13) 0.812 (18) 0.581 (15) 0.313 (18) 0.388 (18) 0.434 (17) 0.973 (6) 0.585 (15) 0.867 (15) 0.738 (13) 0.770 (12) 0.448 (18)
0.431 (18) 0.534 (18) 0.642 (10) 0.964 (3) 0.723 (17) 0.617 (17) 0.887 (17) 0.670 (11) 0.901 (5) 0.496 (17) 0.553 (9) 0.703 (16) 0.942 (14) 0.618 (15) 0.781 (16) 0.715 (16) 0.874 (15) 0.433 (18) 0.748 (16) 0.710 (18) 0.609 (15) 0.769 (15) 0.702 (13) 0.972 (13) 0.764 (18) 0.474 (17) 0.575 (18) 0.528 (15) 0.507 (17) 0.758 (16) 0.919 (16) 0.374 (17) 0.886 (14) 0.787 (17) 0.536 (18) 0.923 (16) 0.982 (14) 0.437 (16) 0.650 (17) 0.486 (17) 0.457 (16) 0.919 (13) 0.653 (12) 0.792 (16) 0.565 (18) 0.730 (14) 0.599 (15)
0.999 (5) 0.999 (3) 0.631 (12) 0.956 (5) 0.891 (15) 0.921 (6) 1.000 (1) 0.608 (13) 0.729 (13) 0.713 (10) 0.493 (16) 0.745 (11) 1.000 (8) 0.772 (4) 0.929 (5) 0.870 (8) 1.000 (2) 0.998 (3) 0.862 (11) 0.996 (3) 0.746 (2) 0.808 (12) 0.785 (4) 0.994 (5) 0.979 (11) 0.950 (9) 0.921 (12) 0.776 (3) 0.715 (1) 0.954 (8) 0.998 (3) 0.807 (12) 0.961 (4) 0.979 (6) 0.905 (7) 0.992 (4) 1.000 (2) 0.693 (13) 1.000 (1) 0.720 (3) 0.622 (2) 0.969 (8) 0.721 (6) 0.983 (3) 0.892 (3) 0.912 (4) 0.735 (8)
0.999 (6) 0.999 (1) 0.585 (15) 0.928 (13) 0.961 (10) 0.870 (10) 1.000 (2) 0.593 (15) 0.658 (16) 0.731 (8) 0.559 (6) 0.736 (13) 1.000 (1) 0.668 (14) 0.931 (3) 0.782 (14) 1.000 (4) 0.997 (5) 0.832 (13) 0.994 (7) 0.627 (13) 0.936 (6) 0.729 (10) 0.981 (11) 0.979 (12) 0.981 (2) 0.988 (8) 0.595 (12) 0.577 (13) 0.939 (10) 0.997 (4) 0.826 (10) 0.963 (3) 0.980 (5) 0.828 (10) 0.923 (15) 1.000 (3) 0.789 (10) 1.000 (2) 0.682 (8) 0.574 (9) 0.697 (17) 0.606 (14) 0.942 (8) 0.747 (12) 0.843 (10) 0.638 (13)
0.999 (3) 0.999 (2) 0.630 (13) 0.948 (8) 0.918 (12) 0.910 (7) 1.000 (4) 0.600 (14) 0.717 (14) 0.740 (6) 0.499 (15) 0.765 (9) 1.000 (2) 0.745 (5) 0.922 (6) 0.837 (10) 1.000 (6) 0.998 (4) 0.873 (10) 0.996 (4) 0.745 (3) 0.814 (11) 0.774 (6) 0.994 (6) 0.981 (9) 0.953 (8) 0.873 (16) 0.744 (4) 0.645 (6) 0.956 (7) 0.997 (5) 0.827 (9) 0.951 (7) 0.984 (3) 0.907 (5) 0.990 (6) 1.000 (4) 0.694 (12) 1.000 (3) 0.740 (1) 0.627 (1) 0.883 (15) 0.701 (9) 0.982 (4) 0.878 (5) 0.905 (5) 0.692 (9)
0.900 (15) 0.937 (13) 0.470 (18) 0.946 (9) 0.994 (5) 0.799 (14) 1.000 (11) 0.549 (16) 0.423 (18) 0.662 (13) 0.483 (18) 0.717 (15) 0.874 (15) 0.570 (17) 0.857 (14) 0.574 (18) 0.862 (17) 0.749 (14) 0.786 (15) 0.771 (17) 0.559 (17) 0.912 (7) 0.586 (17) 0.924 (17) 0.987 (8) 0.970 (4) 0.998 (3) 0.493 (16) 0.623 (9) 0.868 (14) 0.937 (14) 0.781 (13) 0.749 (17) 0.910 (13) 0.705 (16) 0.877 (17) 0.965 (16) 0.987 (1) 0.891 (14) 0.614 (13) 0.589 (8) 0.569 (18) 0.559 (17) 0.625 (18) 0.629 (16) 0.657 (16) 0.516 (17)
0.999 (4) 0.993 (6) 0.645 (9) 0.954 (6) 0.912 (13) 0.812 (13) 0.997 (12) 0.735 (6) 0.869 (8) 0.667 (12) 0.536 (12) 0.727 (14) 0.998 (9) 0.775 (3) 0.907 (7) 0.907 (5) 1.000 (3) 0.994 (7) 0.900 (9) 0.993 (8) 0.690 (8) 0.840 (10) 0.813 (1) 0.997 (2) 0.988 (7) 0.939 (11) 0.928 (11) 0.695 (7) 0.688 (3) 0.963 (4) 0.995 (8) 0.822 (11) 0.928 (11) 0.963 (9) 0.907 (6) 0.973 (11) 0.997 (10) 0.649 (14) 0.995 (9) 0.656 (11) 0.558 (11) 0.974 (5) 0.676 (11) 0.986 (2) 0.886 (4) 0.845 (9) 0.773 (7)
0.872 (16) 0.578 (17) 0.662 (6) 0.923 (16) 0.899 (14) 0.596 (18) 0.695 (18) 0.434 (18) 0.893 (6) 0.675 (11) 0.542 (10) 0.878 (4) 0.561 (18) 0.677 (12) 0.238 (18) 0.863 (9) 0.948 (14) 0.589 (15) 0.646 (17) 0.859 (15) 0.591 (16) 0.510 (18) 0.725 (11) 0.943 (16) 0.870 (17) 0.725 (15) 0.836 (17) 0.461 (17) 0.555 (15) 0.697 (17) 0.859 (18) 0.573 (15) 0.883 (15) 0.914 (11) 0.720 (14) 0.852 (18) 0.926 (17) 0.417 (17) 0.704 (16) 0.574 (15) 0.462 (15) 0.749 (16) 0.695 (10) 0.791 (17) 0.608 (17) 0.612 (18) 0.550 (16)
0.929 (14) 0.696 (15) 0.700 (1) 0.951 (7) 0.743 (16) 0.682 (16) 0.995 (14) 0.677 (10) 0.931 (2) 0.514 (15) 0.558 (7) 0.793 (7) 0.870 (16) 0.733 (7) 0.859 (13) 0.806 (12) 0.869 (16) 0.581 (17) 0.904 (7) 0.891 (14) 0.721 (7) 0.597 (17) 0.626 (15) 0.978 (12) 0.980 (10) 0.711 (16) 0.885 (14) 0.558 (14) 0.508 (16) 0.863 (15) 0.921 (15) 0.371 (18) 0.888 (13) 0.848 (15) 0.712 (15) 0.975 (10) 0.978 (15) 0.318 (18) 0.752 (15) 0.500 (16) 0.383 (18) 0.978 (4) 0.708 (7) 0.913 (11) 0.711 (15) 0.624 (17) 0.688 (10)
0.991 (11) 0.934 (14) 0.595 (14) 0.968 (1) 0.988 (7) 0.923 (5) 0.982 (16) 0.721 (7) 0.858 (10) 0.721 (9) 0.576 (4) 0.754 (10) 1.000 (6) 0.676 (13) 0.833 (15) 0.892 (6) 1.000 (9) 0.975 (11) 0.904 (8) 0.992 (10) 0.679 (9) 0.901 (8) 0.724 (12) 0.963 (14) 0.967 (14) 0.915 (12) 0.997 (4) 0.673 (9) 0.625 (8) 0.946 (9) 0.992 (10) 0.859 (5) 0.949 (8) 0.911 (12) 0.874 (9) 0.978 (8) 0.999 (9) 0.866 (6) 0.989 (10) 0.682 (7) 0.532 (12) 0.938 (11) 0.776 (5) 0.925 (9) 0.826 (8) 0.855 (8) 0.838 (4)
0.995 (9) 0.985 (9) 0.684 (5) 0.965 (2) 0.998 (1) 0.944 (1) 1.000 (8) 0.820 (2) 0.927 (3) 0.807 (2) 0.554 (8) 0.797 (6) 1.000 (7) 0.707 (10) 0.899 (10) 0.930 (3) 1.000 (1) 0.995 (6) 0.955 (3) 0.996 (5) 0.735 (5) 0.972 (2) 0.745 (9) 0.983 (9) 0.991 (4) 0.981 (3) 1.000 (1) 0.725 (5) 0.606 (10) 0.967 (3) 0.997 (6) 0.885 (3) 0.938 (9) 0.957 (10) 0.916 (4) 0.995 (2) 1.000 (1) 0.924 (5) 0.998 (6) 0.673 (10) 0.591 (7) 0.988 (1) 0.855 (2) 0.981 (6) 0.871 (6) 0.860 (7) 0.853 (1)
0.998 (7) 0.969 (11) 0.633 (11) 0.927 (15) 0.992 (6) 0.909 (8) 1.000 (5) 0.535 (17) 0.606 (17) 0.778 (5) 0.484 (17) 0.739 (12) 1.000 (3) 0.692 (11) 0.879 (11) 0.760 (15) 0.998 (10) 0.977 (10) 0.793 (14) 0.995 (6) 0.665 (11) 0.772 (14) 0.762 (7) 0.981 (10) 0.978 (13) 0.988 (1) 0.919 (13) 0.653 (10) 0.693 (2) 0.939 (11) 0.979 (12) 0.885 (4) 0.909 (12) 0.984 (4) 0.923 (3) 0.990 (5) 1.000 (5) 0.838 (8) 1.000 (4) 0.704 (5) 0.619 (3) 0.935 (12) 0.552 (18) 0.880 (13) 0.808 (9) 0.929 (1) 0.681 (11)
0.961 (13) 0.952 (12) 0.548 (16) 0.838 (17) 0.987 (8) 0.857 (11) 1.000 (7) 0.690 (9) 0.667 (15) 0.509 (16) 0.505 (14) 0.653 (17) 0.995 (10) 0.573 (16) 0.902 (9) 0.632 (17) 0.949 (13) 0.838 (13) 0.861 (12) 0.953 (13) 0.666 (10) 0.945 (5) 0.555 (18) 0.953 (15) 0.916 (15) 0.957 (7) 0.947 (9) 0.620 (11) 0.584 (12) 0.923 (12) 0.947 (13) 0.828 (8) 0.793 (16) 0.841 (16) 0.722 (13) 0.988 (7) 0.999 (8) 0.978 (4) 0.975 (11) 0.711 (4) 0.617 (4) 0.960 (9) 0.582 (16) 0.890 (12) 0.737 (14) 0.714 (15) 0.601 (14)
0.999 (2) 0.997 (4) 0.694 (2) 0.929 (12) 0.996 (2) 0.943 (2) 1.000 (6) 0.797 (3) 0.862 (9) 0.806 (3) 0.627 (3) 0.954 (1) 1.000 (5) 0.822 (2) 0.936 (2) 0.954 (2) 1.000 (7) 0.998 (2) 0.966 (2) 0.998 (1) 0.752 (1) 0.983 (1) 0.809 (2) 0.997 (3) 0.999 (1) 0.947 (10) 0.993 (6) 0.829 (2) 0.641 (7) 0.973 (1) 0.998 (1) 0.905 (1) 0.980 (2) 0.997 (1) 0.938 (2) 0.995 (3) 1.000 (6) 0.986 (2) 0.999 (5) 0.729 (2) 0.594 (6) 0.986 (2) 0.880 (1) 0.986 (1) 0.906 (2) 0.913 (3) 0.819 (5)
1.000 (1) 0.994 (5) 0.689 (3) 0.943 (11) 0.995 (3) 0.928 (4) 1.000 (3) 0.854 (1) 0.916 (4) 0.816 (1) 0.711 (1) 0.951 (2) 1.000 (4) 0.836 (1) 0.938 (1) 0.984 (1) 1.000 (5) 0.998 (1) 0.971 (1) 0.998 (2) 0.744 (4) 0.970 (3) 0.804 (3) 0.998 (1) 0.999 (2) 0.957 (6) 1.000 (2) 0.848 (1) 0.673 (4) 0.961 (5) 0.998 (2) 0.897 (2) 0.993 (1) 0.995 (2) 0.938 (1) 0.997 (1) 1.000 (7) 0.983 (3) 0.996 (8) 0.698 (6) 0.596 (5) 0.982 (3) 0.820 (3) 0.982 (5) 0.910 (1) 0.921 (2) 0.842 (3)
CIFAR10 FashionMNIST MNIST-C MVTec-AD SVHN
0.925 (6) 0.917 (6) 0.896 (9) 0.847 (3) 0.584 (10)
0.829 (15) 0.837 (14) 0.819 (13) 0.669 (17) 0.553 (15)
0.833 (14) 0.840 (13) 0.839 (12) 0.805 (9) 0.583 (11)
0.880 (11) 0.891 (11) 0.721 (16) 0.680 (14) 0.525 (17)
0.745 (17) 0.821 (15) 0.650 (18) 0.841 (5) 0.515 (18)
0.921 (7) 0.910 (9) 0.926 (2) 0.815 (7) 0.623 (1)
0.845 (12) 0.883 (12) 0.929 (1) 0.723 (13) 0.605 (5)
0.908 (9) 0.916 (7) 0.922 (3) 0.791 (10) 0.616 (3)
0.766 (16) 0.771 (18) 0.891 (10) 0.673 (15) 0.605 (6)
0.941 (3) 0.934 (3) 0.919 (5) 0.841 (6) 0.610 (4)
0.642 (18) 0.781 (17) 0.729 (15) 0.607 (18) 0.536 (16)
0.882 (10) 0.894 (10) 0.716 (17) 0.670 (16) 0.567 (12)
0.917 (8) 0.910 (8) 0.897 (8) 0.845 (4) 0.588 (9)
0.928 (5) 0.922 (5) 0.881 (11) 0.886 (1) 0.558 (13)
0.941 (4) 0.932 (4) 0.913 (7) 0.789 (11) 0.602 (7)
0.840 (13) 0.798 (16) 0.794 (14) 0.769 (12) 0.556 (14)
0.951 (2) 0.939 (2) 0.918 (6) 0.879 (2) 0.621 (2)
0.957 (1) 0.950 (1) 0.921 (4) 0.811 (8) 0.600 (8)
20news agnews amazon imdb yelp
0.805 (8) 0.843 (8) 0.704 (10) 0.722 (10) 0.804 (9)
0.688 (16) 0.731 (15) 0.603 (14) 0.579 (15) 0.625 (15)
0.752 (12) 0.752 (13) 0.625 (13) 0.631 (12) 0.716 (13)
0.477 (18) 0.647 (18) 0.482 (17) 0.416 (18) 0.549 (17)
0.702 (15) 0.745 (14) 0.476 (18) 0.449 (17) 0.508 (18)
0.866 (1) 0.904 (3) 0.806 (1) 0.857 (1) 0.895 (1)
0.788 (10) 0.828 (10) 0.731 (7) 0.749 (8) 0.775 (10)
0.852 (4) 0.898 (4) 0.803 (2) 0.841 (2) 0.892 (2)
0.777 (11) 0.794 (12) 0.716 (9) 0.774 (7) 0.859 (6)
0.852 (5) 0.895 (5) 0.774 (5) 0.820 (4) 0.886 (4)
0.633 (17) 0.682 (16) 0.591 (15) 0.583 (14) 0.675 (14)
0.734 (13) 0.848 (7) 0.529 (16) 0.523 (16) 0.566 (16)
0.809 (7) 0.833 (9) 0.721 (8) 0.732 (9) 0.811 (8)
0.802 (9) 0.825 (11) 0.655 (12) 0.715 (11) 0.754 (11)
0.855 (2) 0.919 (2) 0.776 (4) 0.836 (3) 0.886 (3)
0.716 (14) 0.651 (17) 0.685 (11) 0.585 (13) 0.723 (12)
0.852 (3) 0.881 (6) 0.743 (6) 0.781 (6) 0.853 (7)
0.841 (6) 0.922 (1) 0.777 (3) 0.816 (5) 0.882 (5)
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 55: AUCPR performance on tabular datas. Performance (mean ± std, rank) under varying label ratios (𝛾𝑙𝑎 ) and label counts (𝑁𝑙𝑎 ). Label Ratio (𝛾𝑙𝑎 )
Label Count (𝑁𝑙𝑎 )
Type
Model
Key Mechanism 1%
5%
10%
1
5
10
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Aug. Data Aug. Pseudo-Labeling
0.481 ±0.257 (9) 0.379 ±0.251 (13) 0.408 ±0.259 (11) 0.308 ±0.267 (17) 0.314 ±0.261 (16) 0.555 ±0.314 (2) 0.524 ±0.296 (5) 0.547 ±0.309 (3) 0.378 ±0.249 (15) 0.518 ±0.293 (7) 0.378 ±0.258 (14)
0.623 ±0.261 (7) 0.497 ±0.276 (14) 0.538 ±0.275 (11) 0.306 ±0.263 (18) 0.315 ±0.255 (17) 0.635 ±0.304 (6) 0.646 ±0.286 (4) 0.644 ±0.295 (5) 0.498 ±0.258 (13) 0.580 ±0.297 (10) 0.385 ±0.260 (16)
0.688 ±0.257 (5) 0.568 ±0.283 (15) 0.602 ±0.276 (12) 0.310 ±0.269 (18) 0.316 ±0.260 (17) 0.666 ±0.302 (8) 0.691 ±0.281 (4) 0.676 ±0.292 (6) 0.570 ±0.272 (14) 0.603 ±0.296 (11) 0.415 ±0.272 (16)
0.385 ±0.244 (11) 0.332 ±0.223 (16) 0.353 ±0.235 (13) 0.322 ±0.274 (17) 0.310 ±0.257 (18) 0.517 ±0.315 (1) 0.469 ±0.295 (5) 0.508 ±0.309 (3) 0.346 ±0.233 (14) 0.479 ±0.290 (4) 0.344 ±0.240 (15)
0.580 ±0.258 (8) 0.474 ±0.266 (15) 0.537 ±0.274 (12) 0.282 ±0.245 (18) 0.314 ±0.258 (17) 0.610 ±0.311 (4) 0.608 ±0.300 (6) 0.608 ±0.304 (5) 0.484 ±0.250 (14) 0.571 ±0.301 (9) 0.435 ±0.273 (16)
0.657 ±0.260 (5) 0.556 ±0.283 (14) 0.607 ±0.286 (11) 0.282 ±0.249 (18) 0.319 ±0.258 (17) 0.648 ±0.304 (7) 0.662 ±0.292 (4) 0.650 ±0.300 (6) 0.554 ±0.273 (15) 0.603 ±0.299 (12) 0.491 ±0.287 (16)
Supervised
XGBoost CatBoost FTTransformer TabM TabR-S TabPFN LimiX
GBDT GBDT Deep (Sup.) Deep (Sup.) Deep (Sup.) Found. Model Found. Model
0.241 ±0.266 (18) 0.521 ±0.281 (6) 0.515 ±0.322 (8) 0.468 ±0.298 (10) 0.405 ±0.276 (12) 0.528 ±0.260 (4) 0.585 ±0.269 (1)
0.495 ±0.311 (15) 0.650 ±0.276 (3) 0.621 ±0.294 (8) 0.615 ±0.288 (9) 0.517 ±0.274 (12) 0.710 ±0.251 (2) 0.716 ±0.253 (1)
0.622 ±0.291 (10) 0.715 ±0.271 (3) 0.664 ±0.291 (9) 0.675 ±0.274 (7) 0.596 ±0.280 (13) 0.770 ±0.243 (2) 0.771 ±0.246 (1)
0.430 ±0.266 (8) 0.457 ±0.274 (6) 0.448 ±0.310 (7) 0.404 ±0.279 (10) 0.370 ±0.275 (12) 0.415 ±0.255 (9) 0.515 ±0.261 (2)
0.558 ±0.259 (11) 0.635 ±0.272 (3) 0.567 ±0.315 (10) 0.582 ±0.306 (7) 0.513 ±0.273 (13) 0.667 ±0.266 (2) 0.688 ±0.261 (1)
0.630 ±0.266 (9) 0.689 ±0.269 (3) 0.612 ±0.321 (10) 0.643 ±0.297 (8) 0.569 ±0.281 (13) 0.735 ±0.253 (2) 0.745 ±0.251 (1)
Median Unsup.
ECOD
Probabilistic
0.365 ± 0.264
Table 56: AUCPR performance on image datasets (ViT features). Performance (mean ± std, rank) under varying 𝛾𝑙𝑎 and 𝑁𝑙𝑎 . Label Ratio (𝛾𝑎 )
Label Count (𝑁𝑙𝑎 )
Type
Model
Key Mechanism 1%
5%
10%
1
5
10
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Aug. Data Aug. Pseudo-Labeling
0.388 ±0.242 (9) 0.268 ±0.168 (17) 0.288 ±0.198 (12) 0.332 ±0.246 (10) 0.277 ±0.282 (16) 0.459 ±0.296 (5) 0.449 ±0.290 (7) 0.468 ±0.300 (3) 0.285 ±0.203 (14) 0.520 ±0.307 (1) 0.292 ±0.231 (11)
0.571 ±0.301 (9) 0.395 ±0.226 (14) 0.481 ±0.264 (11) 0.318 ±0.220 (15) 0.279 ±0.282 (17) 0.629 ±0.307 (4) 0.601 ±0.304 (7) 0.624 ±0.307 (5) 0.457 ±0.271 (12) 0.636 ±0.311 (2) 0.287 ±0.235 (16)
0.633 ±0.307 (9) 0.474 ±0.249 (14) 0.592 ±0.292 (11) 0.320 ±0.227 (15) 0.279 ±0.280 (17) 0.694 ±0.312 (2) 0.666 ±0.304 (6) 0.688 ±0.309 (4) 0.548 ±0.287 (12) 0.689 ±0.314 (3) 0.292 ±0.238 (16)
0.300 ±0.198 (11) 0.217 ±0.147 (16) 0.213 ±0.175 (17) 0.314 ±0.253 (10) 0.282 ±0.284 (13) 0.333 ±0.249 (6) 0.328 ±0.242 (8) 0.349 ±0.260 (4) 0.218 ±0.173 (15) 0.421 ±0.290 (2) 0.318 ±0.262 (9)
0.471 ±0.271 (9) 0.302 ±0.180 (16) 0.350 ±0.215 (11) 0.327 ±0.231 (13) 0.278 ±0.281 (17) 0.541 ±0.299 (4) 0.515 ±0.295 (7) 0.541 ±0.304 (3) 0.341 ±0.216 (12) 0.557 ±0.308 (2) 0.305 ±0.250 (15)
0.557 ±0.297 (9) 0.376 ±0.213 (14) 0.447 ±0.243 (11) 0.324 ±0.227 (16) 0.278 ±0.282 (17) 0.625 ±0.309 (3) 0.599 ±0.303 (7) 0.619 ±0.310 (5) 0.434 ±0.257 (12) 0.627 ±0.313 (2) 0.366 ±0.258 (15)
Supervised
XGBoost CatBoost TabM TabR-S TabPFN LimiX
GBDT GBDT Deep (Sup.) Deep (Sup.) Found. Model Found. Model
0.283 ±0.241 (15) 0.448 ±0.275 (8) 0.461 ±0.293 (4) 0.287 ±0.210 (13) 0.458 ±0.266 (6) 0.518 ±0.301 (2)
0.534 ±0.284 (10) 0.579 ±0.316 (8) 0.605 ±0.306 (6) 0.411 ±0.242 (13) 0.639 ±0.305 (1) 0.633 ±0.309 (3)
0.596 ±0.294 (10) 0.635 ±0.322 (8) 0.660 ±0.313 (7) 0.488 ±0.269 (13) 0.702 ±0.307 (1) 0.679 ±0.306 (5)
0.231 ±0.151 (14) 0.403 ±0.269 (3) 0.344 ±0.256 (5) 0.284 ±0.241 (12) 0.331 ±0.234 (7) 0.422 ±0.282 (1)
0.460 ±0.269 (10) 0.491 ±0.287 (8) 0.522 ±0.291 (6) 0.308 ±0.206 (14) 0.536 ±0.287 (5) 0.565 ±0.304 (1)
0.550 ±0.293 (10) 0.566 ±0.310 (8) 0.602 ±0.304 (6) 0.385 ±0.235 (13) 0.627 ±0.306 (1) 0.622 ±0.310 (4)
Median Unsup.
IForest
Isolation-based
0.360 ± 0.270
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 57: AUCPR performance on text datasets (RoBERTa features). Performance (mean ± std, rank) under varying 𝛾𝑙𝑎 and 𝑁𝑙𝑎 . Type
Model
Label Ratio (𝛾𝑙𝑎 )
Key Mechanism
Label Count (𝑁𝑙𝑎 )
1%
5%
10%
1
5
10
Semi-supervised
XGBOD DeepSAD REPEN AA-BiGAN GANomaly DevNet FEAWAD PReNet RoSAS Dual-MGAN SOEL-NTL
Repr. Learning Repr. Learning Repr. Learning GAN-based GAN-based Score Learning Score Learning Score Learning Data Aug. Data Aug. Pseudo-Labeling
0.166 ±0.087 (8) 0.105 ±0.046 (14) 0.134 ±0.058 (12) 0.067 ±0.025 (17) 0.110 ±0.055 (13) 0.272 ±0.172 (2) 0.221 ±0.132 (6) 0.261 ±0.154 (4) 0.164 ±0.076 (9) 0.253 ±0.149 (5) 0.087 ±0.041 (16)
0.284 ±0.147 (8) 0.171 ±0.100 (13) 0.248 ±0.132 (11) 0.072 ±0.029 (17) 0.113 ±0.056 (14) 0.430 ±0.219 (1) 0.379 ±0.203 (7) 0.424 ±0.215 (2) 0.277 ±0.131 (10) 0.404 ±0.207 (5) 0.101 ±0.075 (15)
0.366 ±0.172 (8) 0.225 ±0.125 (13) 0.344 ±0.186 (11) 0.078 ±0.050 (17) 0.107 ±0.050 (14) 0.512 ±0.219 (2) 0.461 ±0.222 (7) 0.513 ±0.214 (1) 0.356 ±0.170 (10) 0.498 ±0.209 (5) 0.104 ±0.084 (15)
0.132 ±0.073 (8) 0.088 ±0.038 (16) 0.108 ±0.057 (13) 0.065 ±0.025 (17) 0.109 ±0.054 (12) 0.171 ±0.114 (4) 0.146 ±0.085 (7) 0.184 ±0.112 (2) 0.127 ±0.069 (10) 0.176 ±0.098 (3) 0.156 ±0.129 (6)
0.221 ±0.125 (8) 0.130 ±0.074 (14) 0.172 ±0.088 (12) 0.076 ±0.048 (17) 0.110 ±0.052 (16) 0.352 ±0.193 (3) 0.289 ±0.173 (7) 0.340 ±0.188 (4) 0.217 ±0.126 (10) 0.338 ±0.184 (5) 0.159 ±0.097 (13)
0.282 ±0.147 (9) 0.171 ±0.092 (14) 0.231 ±0.124 (12) 0.076 ±0.051 (17) 0.111 ±0.055 (15) 0.436 ±0.201 (2) 0.365 ±0.191 (7) 0.433 ±0.187 (3) 0.284 ±0.134 (8) 0.419 ±0.181 (4) 0.181 ±0.129 (13)
Supervised
XGBoost CatBoost FTTransformer TabM TabPFN LimiX
GBDT GBDT Deep (Sup.) Deep (Sup.) Found. Model Found. Model
0.151 ±0.079 (10) 0.143 ±0.062 (11) 0.094 ±0.077 (15) 0.278 ±0.163 (1) 0.184 ±0.088 (7) 0.266 ±0.155 (3)
0.280 ±0.145 (9) 0.224 ±0.123 (12) 0.084 ±0.123 (16) 0.407 ±0.212 (4) 0.394 ±0.188 (6) 0.420 ±0.198 (3)
0.361 ±0.179 (9) 0.317 ±0.165 (12) 0.086 ±0.140 (16) 0.478 ±0.209 (6) 0.503 ±0.191 (4) 0.506 ±0.213 (3)
0.106 ±0.052 (14) 0.120 ±0.054 (11) 0.105 ±0.076 (15) 0.163 ±0.092 (5) 0.129 ±0.069 (9) 0.187 ±0.105 (1)
0.221 ±0.126 (9) 0.195 ±0.116 (11) 0.114 ±0.141 (15) 0.352 ±0.168 (2) 0.304 ±0.159 (6) 0.356 ±0.195 (1)
0.281 ±0.149 (10) 0.251 ±0.155 (11) 0.094 ±0.140 (16) 0.437 ±0.185 (1) 0.401 ±0.173 (6) 0.411 ±0.197 (5)
Median Unsup.
VAE
Reconstruction
0.093 ± 0.036
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Xu Yao et al.
Table 58: Performance comparison of video anomaly detection models Pretrain
i3d
x3d
sf50
mvit
sf
Model
Category
AUCROC
AUCPR
TAD
UCF-Crime
XD-Violence
ShanghaiTech
Mean
TAD
UCF-Crime
XD-Violence
ShanghaiTech
Mean
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.858 ± 0.001 (5) 0.849 ± 0.042 (7) 0.808 ± 0.017 (10) 0.834 ± 0.003 (9) 0.838 ± 0.023 (8) 0.885 ± 0.004 (1) 0.808 ± 0.009 (10)
0.806 ± 0.001 (1) 0.795 ± 0.006 (4) 0.779 ± 0.014 (8) 0.806 ± 0.001 (1) 0.760 ± 0.014 (10) 0.755 ± 0.010 (11) 0.793 ± 0.002 (6)
0.871 ± 0.000 (2) 0.779 ± 0.021 (11) 0.837 ± 0.007 (6) 0.797 ± 0.003 (10) 0.827 ± 0.008 (9) 0.834 ± 0.002 (7) 0.855 ± 0.002 (4)
0.898 ± 0.003 (6) 0.688 ± 0.032 (11) 0.845 ± 0.010 (8) 0.851 ± 0.001 (7) 0.833 ± 0.016 (9) 0.906 ± 0.006 (4) 0.727 ± 0.039 (10)
0.858 (3) 0.778 (11) 0.817 (8) 0.822 (7) 0.815 (9) 0.845 (5) 0.796 (10)
0.298 ± 0.001 (4) 0.262 ± 0.060 (9) 0.268 ± 0.013 (8) 0.285 ± 0.002 (7) 0.227 ± 0.023 (10) 0.293 ± 0.007 (5) 0.216 ± 0.016 (11)
0.207 ± 0.000 (4) 0.171 ± 0.003 (10) 0.209 ± 0.024 (3) 0.216 ± 0.002 (1) 0.185 ± 0.017 (7) 0.162 ± 0.013 (11) 0.188 ± 0.003 (6)
0.658 ± 0.001 (2) 0.446 ± 0.039 (11) 0.606 ± 0.012 (5) 0.590 ± 0.007 (7) 0.600 ± 0.009 (6) 0.536 ± 0.003 (10) 0.621 ± 0.005 (3)
0.557 ± 0.002 (1) 0.154 ± 0.028 (11) 0.460 ± 0.014 (6) 0.477 ± 0.004 (4) 0.350 ± 0.066 (9) 0.384 ± 0.010 (8) 0.191 ± 0.046 (10)
0.430 (1) 0.258 (11) 0.386 (6) 0.392 (5) 0.341 (9) 0.344 (8) 0.304 (10)
IForest CatB DevNet FEAWAD DeepSAD
Isolation-based GBDT Score Learning Score Learning Repr. Learning
0.419 ± 0.010 (12) 0.875 ± 0.003 (4) 0.850 ± 0.032 (6) 0.882 ± 0.003 (2) 0.880 ± 0.003 (3)
0.504 ± 0.002 (12) 0.794 ± 0.003 (5) 0.787 ± 0.025 (7) 0.804 ± 0.004 (3) 0.776 ± 0.006 (9)
0.640 ± 0.006 (12) 0.842 ± 0.000 (5) 0.827 ± 0.027 (8) 0.856 ± 0.001 (3) 0.883 ± 0.001 (1)
0.646 ± 0.003 (12) 0.908 ± 0.001 (2) 0.904 ± 0.001 (5) 0.907 ± 0.000 (3) 0.916 ± 0.003 (1)
0.552 (12) 0.855 (4) 0.842 (6) 0.862 (2) 0.864 (1)
0.066 ± 0.002 (12) 0.302 ± 0.006 (3) 0.285 ± 0.036 (6) 0.317 ± 0.011 (1) 0.307 ± 0.003 (2)
0.085 ± 0.000 (12) 0.180 ± 0.003 (8) 0.199 ± 0.006 (5) 0.214 ± 0.006 (2) 0.174 ± 0.008 (9)
0.343 ± 0.007 (12) 0.564 ± 0.002 (9) 0.568 ± 0.071 (8) 0.616 ± 0.002 (4) 0.662 ± 0.002 (1)
0.117 ± 0.010 (12) 0.400 ± 0.001 (7) 0.520 ± 0.007 (2) 0.504 ± 0.008 (3) 0.461 ± 0.021 (5)
0.153 (12) 0.362 (7) 0.393 (4) 0.413 (2) 0.401 (3)
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.872 ± 0.002 (5) 0.894 ± 0.012 (1) 0.850 ± 0.047 (9) 0.823 ± 0.005 (11) 0.827 ± 0.019 (10) 0.885 ± 0.003 (3) 0.859 ± 0.006 (7)
0.792 ± 0.002 (4) 0.796 ± 0.003 (3) 0.778 ± 0.004 (6) 0.800 ± 0.001 (2) 0.769 ± 0.005 (9) 0.757 ± 0.009 (11) 0.769 ± 0.004 (9)
0.881 ± 0.001 (1) 0.766 ± 0.030 (11) 0.852 ± 0.015 (5) 0.787 ± 0.003 (10) 0.835 ± 0.007 (7) 0.832 ± 0.002 (8) 0.873 ± 0.003 (2)
0.910 ± 0.002 (4) 0.732 ± 0.044 (11) 0.770 ± 0.071 (10) 0.828 ± 0.004 (8) 0.854 ± 0.011 (7) 0.906 ± 0.005 (5) 0.788 ± 0.008 (9)
0.864 (2) 0.797 (11) 0.812 (9) 0.809 (10) 0.821 (8) 0.845 (5) 0.822 (7)
0.344 ± 0.002 (2) 0.331 ± 0.018 (4) 0.330 ± 0.044 (5) 0.280 ± 0.004 (10) 0.219 ± 0.016 (11) 0.293 ± 0.008 (9) 0.307 ± 0.009 (8)
0.202 ± 0.001 (4) 0.170 ± 0.002 (10) 0.212 ± 0.008 (1) 0.204 ± 0.002 (3) 0.175 ± 0.004 (8) 0.162 ± 0.012 (11) 0.183 ± 0.004 (6)
0.667 ± 0.001 (1) 0.412 ± 0.045 (11) 0.621 ± 0.020 (4) 0.577 ± 0.003 (8) 0.594 ± 0.007 (7) 0.532 ± 0.006 (9) 0.651 ± 0.011 (2)
0.521 ± 0.014 (1) 0.205 ± 0.056 (11) 0.424 ± 0.044 (7) 0.446 ± 0.012 (5) 0.369 ± 0.041 (9) 0.386 ± 0.010 (8) 0.313 ± 0.019 (10)
0.433 (1) 0.280 (11) 0.397 (4) 0.377 (6) 0.339 (10) 0.343 (9) 0.363 (8)
IForest CatB DevNet FEAWAD DeepSAD
Isolation-based GBDT Score Learning Score Learning Repr. Learning
0.433 ± 0.028 (12) 0.854 ± 0.001 (8) 0.864 ± 0.012 (6) 0.893 ± 0.005 (2) 0.881 ± 0.004 (4)
0.506 ± 0.036 (12) 0.786 ± 0.002 (5) 0.777 ± 0.002 (7) 0.801 ± 0.009 (1) 0.773 ± 0.016 (8)
0.516 ± 0.003 (12) 0.822 ± 0.002 (9) 0.846 ± 0.030 (6) 0.868 ± 0.007 (4) 0.872 ± 0.005 (3)
0.660 ± 0.016 (12) 0.918 ± 0.001 (3) 0.901 ± 0.005 (6) 0.921 ± 0.002 (2) 0.926 ± 0.002 (1)
0.529 (12) 0.845 (6) 0.847 (4) 0.871 (1) 0.863 (3)
0.067 ± 0.007 (12) 0.327 ± 0.004 (6) 0.326 ± 0.008 (7) 0.354 ± 0.007 (1) 0.334 ± 0.004 (3)
0.080 ± 0.011 (12) 0.182 ± 0.002 (7) 0.199 ± 0.002 (5) 0.208 ± 0.007 (2) 0.174 ± 0.006 (9)
0.222 ± 0.003 (12) 0.527 ± 0.003 (10) 0.600 ± 0.061 (6) 0.629 ± 0.021 (3) 0.616 ± 0.014 (5)
0.107 ± 0.018 (12) 0.425 ± 0.008 (6) 0.453 ± 0.017 (4) 0.474 ± 0.012 (3) 0.483 ± 0.031 (2)
0.119 (12) 0.365 (7) 0.395 (5) 0.417 (2) 0.402 (3)
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.869 ± 0.001 (5) 0.834 ± 0.035 (9) 0.868 ± 0.024 (6) 0.816 ± 0.006 (10) 0.867 ± 0.024 (7) 0.894 ± 0.002 (2) 0.801 ± 0.011 (11)
0.786 ± 0.001 (4) 0.748 ± 0.009 (10) 0.799 ± 0.011 (1) 0.768 ± 0.004 (7) 0.784 ± 0.009 (5) 0.745 ± 0.011 (11) 0.752 ± 0.003 (9)
0.877 ± 0.001 (2) 0.824 ± 0.058 (11) 0.846 ± 0.018 (8) 0.855 ± 0.002 (6) 0.861 ± 0.002 (5) 0.832 ± 0.002 (10) 0.865 ± 0.005 (4)
0.925 ± 0.001 (5) 0.897 ± 0.044 (7) 0.877 ± 0.023 (9) 0.884 ± 0.004 (8) 0.871 ± 0.033 (10) 0.924 ± 0.006 (6) 0.806 ± 0.032 (11)
0.864 (4) 0.826 (10) 0.848 (6) 0.831 (9) 0.846 (7) 0.849 (5) 0.806 (11)
0.313 ± 0.002 (3) 0.259 ± 0.032 (10) 0.338 ± 0.032 (1) 0.260 ± 0.005 (9) 0.263 ± 0.045 (8) 0.299 ± 0.010 (6) 0.244 ± 0.006 (11)
0.192 ± 0.002 (3) 0.147 ± 0.006 (11) 0.228 ± 0.013 (1) 0.177 ± 0.004 (6) 0.190 ± 0.006 (4) 0.166 ± 0.005 (9) 0.152 ± 0.003 (10)
0.658 ± 0.002 (1) 0.554 ± 0.127 (10) 0.611 ± 0.036 (7) 0.642 ± 0.003 (3) 0.613 ± 0.006 (6) 0.529 ± 0.010 (11) 0.628 ± 0.020 (4)
0.579 ± 0.005 (1) 0.387 ± 0.102 (10) 0.494 ± 0.064 (5) 0.531 ± 0.008 (2) 0.441 ± 0.037 (7) 0.433 ± 0.017 (8) 0.205 ± 0.053 (11)
0.435 (1) 0.337 (10) 0.418 (2) 0.403 (5) 0.377 (7) 0.357 (9) 0.307 (11)
IForest CatB DevNet FEAWAD DeepSAD
Isolation-based GBDT Score Learning Score Learning Repr. Learning
0.465 ± 0.015 (12) 0.900 ± 0.000 (1) 0.845 ± 0.047 (8) 0.889 ± 0.002 (4) 0.893 ± 0.005 (3)
0.579 ± 0.012 (12) 0.797 ± 0.005 (3) 0.762 ± 0.013 (8) 0.799 ± 0.003 (2) 0.771 ± 0.003 (6)
0.547 ± 0.015 (12) 0.846 ± 0.001 (7) 0.839 ± 0.033 (9) 0.868 ± 0.003 (3) 0.884 ± 0.007 (1)
0.649 ± 0.009 (12) 0.940 ± 0.000 (1) 0.936 ± 0.001 (3) 0.933 ± 0.001 (4) 0.937 ± 0.006 (2)
0.560 (12) 0.871 (3) 0.845 (8) 0.872 (1) 0.871 (2)
0.073 ± 0.004 (12) 0.337 ± 0.001 (2) 0.292 ± 0.051 (7) 0.309 ± 0.004 (4) 0.306 ± 0.004 (5)
0.120 ± 0.006 (12) 0.182 ± 0.002 (5) 0.167 ± 0.003 (8) 0.205 ± 0.010 (2) 0.175 ± 0.004 (7)
0.244 ± 0.008 (12) 0.554 ± 0.002 (9) 0.577 ± 0.084 (8) 0.628 ± 0.006 (5) 0.645 ± 0.024 (2)
0.117 ± 0.010 (12) 0.432 ± 0.002 (9) 0.528 ± 0.007 (4) 0.529 ± 0.008 (3) 0.493 ± 0.036 (6)
0.139 (12) 0.376 (8) 0.391 (6) 0.418 (3) 0.405 (4)
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.828 ± 0.001 (8) 0.662 ± 0.140 (12) 0.812 ± 0.009 (10) 0.818 ± 0.004 (9) 0.841 ± 0.045 (7) 0.881 ± 0.006 (6) 0.747 ± 0.017 (11)
0.793 ± 0.001 (4) 0.779 ± 0.007 (9) 0.782 ± 0.004 (8) 0.778 ± 0.001 (10) 0.748 ± 0.005 (12) 0.763 ± 0.004 (11) 0.783 ± 0.003 (7)
0.880 ± 0.000 (2) 0.808 ± 0.032 (12) 0.857 ± 0.007 (8) 0.828 ± 0.003 (11) 0.850 ± 0.006 (9) 0.835 ± 0.002 (10) 0.870 ± 0.002 (5)
0.939 ± 0.001 (5) 0.865 ± 0.048 (11) 0.730 ± 0.040 (12) 0.903 ± 0.004 (8) 0.883 ± 0.019 (10) 0.923 ± 0.006 (7) 0.894 ± 0.019 (9)
0.860 (6) 0.778 (12) 0.795 (11) 0.832 (8) 0.831 (9) 0.851 (7) 0.824 (10)
0.311 ± 0.007 (8) 0.164 ± 0.105 (12) 0.318 ± 0.007 (6) 0.318 ± 0.006 (6) 0.265 ± 0.056 (10) 0.310 ± 0.006 (9) 0.230 ± 0.020 (11)
0.176 ± 0.001 (7) 0.156 ± 0.006 (12) 0.208 ± 0.009 (1) 0.170 ± 0.001 (8) 0.165 ± 0.002 (9) 0.162 ± 0.005 (10) 0.161 ± 0.003 (11)
0.676 ± 0.001 (2) 0.511 ± 0.071 (12) 0.619 ± 0.024 (7) 0.586 ± 0.009 (10) 0.588 ± 0.012 (8) 0.529 ± 0.004 (11) 0.644 ± 0.008 (4)
0.601 ± 0.005 (1) 0.400 ± 0.084 (11) 0.433 ± 0.039 (8) 0.520 ± 0.007 (5) 0.433 ± 0.018 (8) 0.415 ± 0.003 (10) 0.367 ± 0.026 (12)
0.441 (3) 0.308 (12) 0.395 (8) 0.398 (7) 0.363 (9) 0.354 (10) 0.351 (11)
IForest CatB DevNet FEAWAD DeepSAD TabPFN
Isolation-based GBDT Score Learning Score Learning Repr. Learning Tabular Found.
0.514 ± 0.002 (13) 0.909 ± 0.001 (4) 0.882 ± 0.006 (5) 0.910 ± 0.004 (3) 0.924 ± 0.004 (1) 0.917 ± 0.002 (2)
0.441 ± 0.024 (13) 0.799 ± 0.002 (3) 0.791 ± 0.004 (6) 0.791 ± 0.003 (5) 0.803 ± 0.003 (1) 0.801 ± 0.005 (2)
0.607 ± 0.002 (13) 0.873 ± 0.001 (4) 0.862 ± 0.011 (7) 0.875 ± 0.002 (3) 0.896 ± 0.001 (1) 0.867 ± 0.003 (6)
0.647 ± 0.036 (13) 0.945 ± 0.000 (3) 0.937 ± 0.003 (6) 0.944 ± 0.003 (4) 0.953 ± 0.004 (1) 0.948 ± 0.000 (2)
0.552 (13) 0.881 (3) 0.868 (5) 0.880 (4) 0.894 (1) 0.883 (2)
0.098 ± 0.006 (13) 0.394 ± 0.006 (2) 0.359 ± 0.006 (5) 0.399 ± 0.004 (1) 0.390 ± 0.013 (3) 0.379 ± 0.006 (4)
0.069 ± 0.005 (13) 0.180 ± 0.001 (4) 0.177 ± 0.004 (6) 0.186 ± 0.009 (3) 0.193 ± 0.002 (2) 0.178 ± 0.006 (5)
0.289 ± 0.007 (13) 0.632 ± 0.002 (6) 0.640 ± 0.019 (5) 0.660 ± 0.007 (3) 0.703 ± 0.002 (1) 0.587 ± 0.007 (9)
0.116 ± 0.021 (13) 0.505 ± 0.001 (6) 0.570 ± 0.006 (2) 0.563 ± 0.021 (3) 0.526 ± 0.020 (4) 0.489 ± 0.003 (7)
0.143 (13) 0.428 (5) 0.437 (4) 0.452 (2) 0.453 (1) 0.408 (6)
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.897 ± 0.001 (6) 0.816 ± 0.091 (11) 0.880 ± 0.012 (8) 0.850 ± 0.005 (9) 0.902 ± 0.004 (3) 0.898 ± 0.007 (5) 0.841 ± 0.018 (10)
0.793 ± 0.003 (2) 0.768 ± 0.015 (7) 0.794 ± 0.010 (1) 0.778 ± 0.002 (5) 0.767 ± 0.012 (8) 0.741 ± 0.003 (11) 0.752 ± 0.009 (9)
0.884 ± 0.001 (2) 0.841 ± 0.020 (11) 0.874 ± 0.010 (4) 0.862 ± 0.003 (7) 0.863 ± 0.006 (6) 0.843 ± 0.002 (9) 0.871 ± 0.004 (5)
0.922 ± 0.001 (5) 0.867 ± 0.034 (9) 0.866 ± 0.016 (10) 0.883 ± 0.004 (8) 0.889 ± 0.004 (7) 0.913 ± 0.012 (6) 0.865 ± 0.011 (11)
0.874 (2) 0.823 (11) 0.854 (7) 0.843 (9) 0.855 (6) 0.849 (8) 0.832 (10)
0.354 ± 0.002 (3) 0.240 ± 0.066 (11) 0.364 ± 0.025 (2) 0.294 ± 0.008 (9) 0.305 ± 0.014 (8) 0.308 ± 0.015 (7) 0.294 ± 0.024 (9)
0.200 ± 0.003 (3) 0.169 ± 0.009 (8) 0.224 ± 0.010 (1) 0.191 ± 0.002 (5) 0.189 ± 0.020 (6) 0.153 ± 0.007 (11) 0.169 ± 0.007 (8)
0.669 ± 0.002 (1) 0.577 ± 0.065 (9) 0.658 ± 0.021 (3) 0.650 ± 0.003 (5) 0.600 ± 0.014 (7) 0.547 ± 0.008 (11) 0.632 ± 0.019 (6)
0.559 ± 0.005 (1) 0.361 ± 0.098 (10) 0.462 ± 0.060 (6) 0.514 ± 0.007 (3) 0.410 ± 0.015 (8) 0.393 ± 0.025 (9) 0.319 ± 0.031 (11)
0.446 (1) 0.337 (11) 0.427 (3) 0.412 (5) 0.376 (8) 0.350 (10) 0.353 (9)
IForest CatB DevNet FEAWAD DeepSAD
Isolation-based GBDT Score Learning Score Learning Repr. Learning
0.469 ± 0.008 (12) 0.911 ± 0.001 (1) 0.884 ± 0.036 (7) 0.902 ± 0.004 (4) 0.905 ± 0.005 (2)
0.590 ± 0.018 (12) 0.790 ± 0.003 (4) 0.774 ± 0.003 (6) 0.792 ± 0.007 (3) 0.746 ± 0.014 (10)
0.598 ± 0.003 (12) 0.852 ± 0.001 (8) 0.843 ± 0.039 (10) 0.875 ± 0.001 (3) 0.890 ± 0.003 (1)
0.574 ± 0.020 (12) 0.935 ± 0.001 (2) 0.926 ± 0.001 (4) 0.929 ± 0.002 (3) 0.936 ± 0.003 (1)
0.558 (12) 0.872 (3) 0.857 (5) 0.874 (1) 0.869 (4)
0.081 ± 0.008 (12) 0.366 ± 0.002 (1) 0.347 ± 0.053 (4) 0.342 ± 0.007 (6) 0.345 ± 0.008 (5)
0.126 ± 0.014 (12) 0.179 ± 0.002 (7) 0.195 ± 0.006 (4) 0.207 ± 0.007 (2) 0.159 ± 0.007 (10)
0.293 ± 0.001 (12) 0.566 ± 0.005 (10) 0.593 ± 0.074 (8) 0.650 ± 0.001 (4) 0.664 ± 0.008 (2)
0.097 ± 0.008 (12) 0.426 ± 0.003 (7) 0.501 ± 0.004 (4) 0.522 ± 0.018 (2) 0.490 ± 0.010 (5)
0.150 (12) 0.384 (7) 0.409 (6) 0.430 (2) 0.414 (4)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
KDD 2026, August 9–13, 2026, Jeju Island, Republic of Korea.
Table 59: Comparison of AUCROC results on different datasets. Pretrain
Model
32 Segments
Loss Type
200 Segments
TAD
UCF-Crime
XD-Violence
ShanghaiTech
Mean
TAD
UCF-Crime
XD-Violence
ShanghaiTech
Mean
i3d
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.858 ± 0.001 (2) 0.849 ± 0.042 (3) 0.808 ± 0.017 (7) 0.834 ± 0.003 (5) 0.838 ± 0.023 (4) 0.885 ± 0.004 (1) 0.808 ± 0.009 (6)
0.806 ± 0.001 (1) 0.795 ± 0.006 (3) 0.779 ± 0.014 (5) 0.806 ± 0.001 (1) 0.760 ± 0.014 (6) 0.755 ± 0.010 (7) 0.793 ± 0.002 (4)
0.871 ± 0.000 (1) 0.779 ± 0.021 (7) 0.837 ± 0.007 (3) 0.797 ± 0.003 (6) 0.827 ± 0.008 (5) 0.834 ± 0.002 (4) 0.855 ± 0.002 (2)
0.898 ± 0.003 (2) 0.688 ± 0.032 (7) 0.845 ± 0.010 (4) 0.851 ± 0.001 (3) 0.833 ± 0.016 (5) 0.906 ± 0.006 (1) 0.727 ± 0.039 (6)
0.858 (1) 0.778 (7) 0.817 (4) 0.822 (3) 0.815 (5) 0.845 (2) 0.796 (6)
0.858 ± 0.001 (3) 0.893 ± 0.018 (1) 0.809 ± 0.022 (7) 0.824 ± 0.007 (5) 0.848 ± 0.020 (4) 0.876 ± 0.010 (2) 0.811 ± 0.010 (6)
0.806 ± 0.001 (1) 0.795 ± 0.009 (3) 0.777 ± 0.014 (5) 0.801 ± 0.001 (2) 0.768 ± 0.006 (6) 0.758 ± 0.007 (7) 0.797 ± 0.002 (3)
0.871 ± 0.000 (1) 0.792 ± 0.034 (6) 0.843 ± 0.005 (3) 0.787 ± 0.003 (7) 0.840 ± 0.010 (4) 0.833 ± 0.004 (5) 0.867 ± 0.003 (2)
0.898 ± 0.003 (1) 0.744 ± 0.070 (6) 0.844 ± 0.021 (4) 0.829 ± 0.003 (5) 0.867 ± 0.010 (3) 0.889 ± 0.008 (2) 0.666 ± 0.046 (7)
0.858 (1) 0.806 (6) 0.818 (4) 0.810 (5) 0.831 (3) 0.839 (2) 0.785 (7)
x3d
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.872 ± 0.002 (3) 0.894 ± 0.012 (1) 0.850 ± 0.047 (5) 0.823 ± 0.005 (7) 0.827 ± 0.019 (6) 0.885 ± 0.003 (2) 0.859 ± 0.006 (4)
0.792 ± 0.002 (3) 0.796 ± 0.003 (2) 0.778 ± 0.004 (4) 0.800 ± 0.001 (1) 0.769 ± 0.005 (5) 0.757 ± 0.009 (7) 0.769 ± 0.004 (6)
0.881 ± 0.001 (1) 0.766 ± 0.030 (7) 0.852 ± 0.015 (3) 0.787 ± 0.003 (6) 0.835 ± 0.007 (4) 0.832 ± 0.002 (5) 0.873 ± 0.003 (2)
0.910 ± 0.002 (1) 0.732 ± 0.044 (7) 0.770 ± 0.071 (6) 0.828 ± 0.004 (4) 0.854 ± 0.011 (3) 0.906 ± 0.005 (2) 0.788 ± 0.008 (5)
0.864 (1) 0.797 (7) 0.812 (5) 0.810 (6) 0.821 (4) 0.845 (2) 0.822 (3)
0.872 ± 0.002 (3) 0.886 ± 0.016 (1) 0.853 ± 0.031 (6) 0.831 ± 0.008 (7) 0.856 ± 0.011 (5) 0.882 ± 0.007 (2) 0.860 ± 0.010 (4)
0.792 ± 0.002 (1) 0.781 ± 0.010 (3) 0.781 ± 0.008 (2) 0.776 ± 0.013 (5) 0.764 ± 0.017 (6) 0.738 ± 0.021 (7) 0.778 ± 0.007 (4)
0.880 ± 0.000 (2) 0.787 ± 0.064 (7) 0.854 ± 0.018 (3) 0.848 ± 0.036 (5) 0.854 ± 0.004 (4) 0.843 ± 0.004 (6) 0.888 ± 0.004 (1)
0.910 ± 0.002 (1) 0.761 ± 0.033 (7) 0.779 ± 0.056 (6) 0.847 ± 0.010 (3) 0.841 ± 0.061 (4) 0.901 ± 0.015 (2) 0.832 ± 0.016 (5)
0.863 (1) 0.804 (7) 0.817 (6) 0.826 (5) 0.829 (4) 0.841 (2) 0.839 (3)
sf50
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.869 ± 0.001 (2) 0.834 ± 0.035 (5) 0.868 ± 0.024 (3) 0.816 ± 0.006 (6) 0.867 ± 0.024 (4) 0.894 ± 0.002 (1) 0.801 ± 0.011 (7)
0.786 ± 0.001 (2) 0.748 ± 0.009 (6) 0.799 ± 0.011 (1) 0.768 ± 0.004 (4) 0.784 ± 0.009 (3) 0.745 ± 0.011 (7) 0.752 ± 0.003 (5)
0.877 ± 0.001 (1) 0.824 ± 0.058 (7) 0.846 ± 0.018 (5) 0.855 ± 0.002 (4) 0.861 ± 0.002 (3) 0.832 ± 0.002 (6) 0.865 ± 0.005 (2)
0.925 ± 0.001 (1) 0.897 ± 0.044 (3) 0.877 ± 0.023 (5) 0.884 ± 0.004 (4) 0.871 ± 0.033 (6) 0.924 ± 0.006 (2) 0.806 ± 0.032 (7)
0.864 (1) 0.826 (6) 0.848 (3) 0.831 (5) 0.846 (4) 0.849 (2) 0.806 (7)
0.871 ± 0.001 (3) 0.827 ± 0.044 (5) 0.856 ± 0.037 (4) 0.813 ± 0.005 (7) 0.881 ± 0.021 (2) 0.889 ± 0.003 (1) 0.821 ± 0.015 (6)
0.786 ± 0.001 (2) 0.750 ± 0.011 (6) 0.793 ± 0.010 (1) 0.767 ± 0.004 (4) 0.786 ± 0.008 (3) 0.734 ± 0.009 (7) 0.758 ± 0.004 (5)
0.877 ± 0.000 (1) 0.796 ± 0.058 (7) 0.831 ± 0.016 (6) 0.855 ± 0.002 (4) 0.863 ± 0.002 (3) 0.833 ± 0.004 (5) 0.875 ± 0.004 (2)
0.925 ± 0.001 (1) 0.865 ± 0.049 (7) 0.891 ± 0.012 (3) 0.886 ± 0.002 (5) 0.885 ± 0.024 (6) 0.922 ± 0.010 (2) 0.888 ± 0.009 (4)
0.865 (1) 0.809 (7) 0.843 (4) 0.830 (6) 0.854 (2) 0.844 (3) 0.836 (5)
mvit
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.828 ± 0.001 (3) 0.662 ± 0.140 (7) 0.812 ± 0.009 (5) 0.818 ± 0.004 (4) 0.841 ± 0.045 (2) 0.881 ± 0.006 (1) 0.747 ± 0.017 (6)
0.793 ± 0.001 (1) 0.779 ± 0.007 (4) 0.782 ± 0.004 (3) 0.778 ± 0.001 (5) 0.748 ± 0.005 (7) 0.763 ± 0.004 (6) 0.783 ± 0.003 (2)
0.880 ± 0.000 (1) 0.808 ± 0.032 (7) 0.857 ± 0.007 (3) 0.828 ± 0.003 (6) 0.850 ± 0.006 (4) 0.835 ± 0.002 (5) 0.870 ± 0.002 (2)
0.939 ± 0.001 (1) 0.865 ± 0.048 (6) 0.730 ± 0.040 (7) 0.903 ± 0.004 (3) 0.883 ± 0.019 (5) 0.923 ± 0.006 (2) 0.894 ± 0.019 (4)
0.860 (1) 0.779 (7) 0.795 (6) 0.832 (3) 0.831 (4) 0.851 (2) 0.824 (5)
0.829 ± 0.001 (3) 0.727 ± 0.142 (7) 0.810 ± 0.011 (5) 0.815 ± 0.003 (4) 0.854 ± 0.036 (2) 0.878 ± 0.014 (1) 0.772 ± 0.027 (6)
0.793 ± 0.001 (1) 0.781 ± 0.010 (4) 0.786 ± 0.010 (3) 0.777 ± 0.001 (5) 0.751 ± 0.005 (7) 0.758 ± 0.010 (6) 0.789 ± 0.004 (2)
0.880 ± 0.000 (1) 0.815 ± 0.036 (7) 0.860 ± 0.003 (3) 0.827 ± 0.003 (6) 0.850 ± 0.007 (4) 0.840 ± 0.003 (5) 0.875 ± 0.006 (2)
0.940 ± 0.000 (1) 0.884 ± 0.030 (6) 0.738 ± 0.026 (7) 0.906 ± 0.005 (3) 0.899 ± 0.017 (5) 0.918 ± 0.011 (2) 0.899 ± 0.012 (4)
0.860 (1) 0.802 (6) 0.798 (7) 0.831 (5) 0.838 (3) 0.849 (2) 0.834 (4)
sf
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.897 ± 0.001 (3) 0.816 ± 0.091 (7) 0.880 ± 0.012 (4) 0.850 ± 0.005 (5) 0.902 ± 0.004 (1) 0.898 ± 0.007 (2) 0.841 ± 0.018 (6)
0.793 ± 0.003 (2) 0.768 ± 0.015 (4) 0.794 ± 0.010 (1) 0.778 ± 0.002 (3) 0.767 ± 0.012 (5) 0.741 ± 0.003 (7) 0.752 ± 0.009 (6)
0.884 ± 0.001 (1) 0.841 ± 0.020 (7) 0.874 ± 0.010 (2) 0.862 ± 0.003 (5) 0.863 ± 0.006 (4) 0.843 ± 0.002 (6) 0.871 ± 0.004 (3)
0.922 ± 0.001 (1) 0.867 ± 0.034 (5) 0.866 ± 0.016 (6) 0.883 ± 0.004 (4) 0.889 ± 0.004 (3) 0.913 ± 0.012 (2) 0.865 ± 0.011 (7)
0.874 (1) 0.823 (7) 0.853 (3) 0.843 (5) 0.855 (2) 0.849 (4) 0.832 (6)
0.898 ± 0.001 (2) 0.853 ± 0.044 (7) 0.881 ± 0.013 (4) 0.856 ± 0.004 (6) 0.899 ± 0.008 (1) 0.891 ± 0.003 (3) 0.860 ± 0.013 (5)
0.794 ± 0.003 (1) 0.775 ± 0.005 (4) 0.783 ± 0.010 (2) 0.777 ± 0.002 (3) 0.771 ± 0.007 (5) 0.742 ± 0.005 (7) 0.748 ± 0.017 (6)
0.882 ± 0.001 (2) 0.833 ± 0.047 (7) 0.865 ± 0.015 (4) 0.861 ± 0.002 (5) 0.868 ± 0.008 (3) 0.839 ± 0.003 (6) 0.882 ± 0.003 (1)
0.922 ± 0.001 (1) 0.893 ± 0.010 (4) 0.879 ± 0.009 (6) 0.884 ± 0.004 (5) 0.878 ± 0.024 (7) 0.909 ± 0.012 (2) 0.894 ± 0.005 (3)
0.874 (1) 0.839 (7) 0.852 (3) 0.845 (6) 0.854 (2) 0.845 (5) 0.846 (4)
Table 60: Comparison of AUCPR results on different datasets. Pretrain
Model
32 Segments
Loss Type
200 Segments
TAD
UCF-Crime
XD-Violence
ShanghaiTech
Mean
TAD
UCF-Crime
XD-Violence
ShanghaiTech
Mean
i3d
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.298 ± 0.001 (1) 0.262 ± 0.060 (5) 0.268 ± 0.013 (4) 0.285 ± 0.002 (3) 0.227 ± 0.023 (6) 0.293 ± 0.007 (2) 0.216 ± 0.016 (7)
0.207 ± 0.000 (3) 0.171 ± 0.003 (6) 0.209 ± 0.024 (2) 0.216 ± 0.002 (1) 0.185 ± 0.017 (5) 0.162 ± 0.013 (7) 0.188 ± 0.003 (4)
0.658 ± 0.001 (1) 0.446 ± 0.039 (7) 0.606 ± 0.012 (3) 0.590 ± 0.007 (5) 0.600 ± 0.009 (4) 0.536 ± 0.003 (6) 0.621 ± 0.005 (2)
0.557 ± 0.002 (1) 0.154 ± 0.028 (7) 0.460 ± 0.014 (3) 0.477 ± 0.004 (2) 0.350 ± 0.066 (5) 0.384 ± 0.010 (4) 0.191 ± 0.046 (6)
0.430 (1) 0.258 (7) 0.386 (3) 0.392 (2) 0.340 (5) 0.344 (4) 0.304 (6)
0.298 ± 0.001 (2) 0.329 ± 0.023 (1) 0.257 ± 0.007 (5) 0.277 ± 0.008 (4) 0.246 ± 0.035 (6) 0.287 ± 0.017 (3) 0.229 ± 0.012 (7)
0.208 ± 0.000 (1) 0.170 ± 0.006 (6) 0.204 ± 0.019 (3) 0.205 ± 0.003 (2) 0.175 ± 0.005 (5) 0.168 ± 0.012 (7) 0.193 ± 0.006 (4)
0.658 ± 0.001 (1) 0.471 ± 0.077 (7) 0.596 ± 0.023 (4) 0.576 ± 0.004 (5) 0.599 ± 0.016 (3) 0.544 ± 0.005 (6) 0.641 ± 0.009 (2)
0.554 ± 0.002 (1) 0.219 ± 0.074 (6) 0.460 ± 0.020 (2) 0.449 ± 0.012 (3) 0.390 ± 0.043 (5) 0.392 ± 0.018 (4) 0.150 ± 0.034 (7)
0.429 (1) 0.297 (7) 0.380 (2) 0.377 (3) 0.353 (4) 0.348 (5) 0.303 (6)
x3d
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.344 ± 0.002 (1) 0.331 ± 0.018 (2) 0.330 ± 0.044 (3) 0.280 ± 0.004 (6) 0.219 ± 0.016 (7) 0.293 ± 0.008 (5) 0.307 ± 0.009 (4)
0.202 ± 0.001 (3) 0.170 ± 0.002 (6) 0.212 ± 0.008 (1) 0.204 ± 0.002 (2) 0.175 ± 0.004 (5) 0.162 ± 0.012 (7) 0.183 ± 0.004 (4)
0.667 ± 0.001 (1) 0.412 ± 0.045 (7) 0.621 ± 0.020 (3) 0.577 ± 0.003 (5) 0.594 ± 0.007 (4) 0.532 ± 0.006 (6) 0.651 ± 0.011 (2)
0.521 ± 0.014 (1) 0.205 ± 0.056 (7) 0.424 ± 0.044 (3) 0.446 ± 0.012 (2) 0.369 ± 0.041 (5) 0.386 ± 0.010 (4) 0.313 ± 0.019 (6)
0.434 (1) 0.279 (7) 0.397 (2) 0.377 (3) 0.339 (6) 0.343 (5) 0.364 (4)
0.344 ± 0.002 (1) 0.337 ± 0.011 (2) 0.316 ± 0.035 (4) 0.296 ± 0.015 (6) 0.273 ± 0.022 (7) 0.323 ± 0.010 (3) 0.307 ± 0.009 (5)
0.203 ± 0.001 (2) 0.171 ± 0.005 (6) 0.216 ± 0.010 (1) 0.199 ± 0.003 (3) 0.175 ± 0.011 (5) 0.163 ± 0.018 (7) 0.192 ± 0.001 (4)
0.664 ± 0.001 (2) 0.459 ± 0.126 (7) 0.610 ± 0.035 (4) 0.640 ± 0.039 (3) 0.599 ± 0.016 (5) 0.548 ± 0.005 (6) 0.695 ± 0.015 (1)
0.522 ± 0.011 (1) 0.198 ± 0.057 (7) 0.374 ± 0.098 (3) 0.454 ± 0.014 (2) 0.354 ± 0.092 (5) 0.369 ± 0.027 (4) 0.330 ± 0.030 (6)
0.433 (1) 0.291 (7) 0.379 (4) 0.398 (2) 0.350 (6) 0.351 (5) 0.381 (3)
sf50
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.313 ± 0.002 (2) 0.259 ± 0.032 (6) 0.338 ± 0.032 (1) 0.260 ± 0.005 (5) 0.263 ± 0.045 (4) 0.299 ± 0.010 (3) 0.244 ± 0.006 (7)
0.192 ± 0.002 (2) 0.147 ± 0.006 (7) 0.228 ± 0.013 (1) 0.177 ± 0.004 (4) 0.190 ± 0.006 (3) 0.166 ± 0.005 (5) 0.152 ± 0.003 (6)
0.658 ± 0.002 (1) 0.554 ± 0.127 (6) 0.611 ± 0.036 (5) 0.642 ± 0.003 (2) 0.613 ± 0.006 (4) 0.529 ± 0.010 (7) 0.628 ± 0.020 (3)
0.579 ± 0.005 (1) 0.387 ± 0.102 (6) 0.494 ± 0.064 (3) 0.531 ± 0.008 (2) 0.441 ± 0.037 (4) 0.433 ± 0.017 (5) 0.205 ± 0.053 (7)
0.436 (1) 0.337 (6) 0.418 (2) 0.403 (3) 0.377 (4) 0.357 (5) 0.307 (7)
0.314 ± 0.003 (2) 0.252 ± 0.036 (6) 0.328 ± 0.034 (1) 0.259 ± 0.003 (5) 0.285 ± 0.051 (4) 0.305 ± 0.010 (3) 0.251 ± 0.011 (7)
0.193 ± 0.002 (2) 0.150 ± 0.008 (7) 0.210 ± 0.015 (1) 0.177 ± 0.002 (4) 0.190 ± 0.004 (3) 0.160 ± 0.009 (5) 0.155 ± 0.005 (6)
0.656 ± 0.001 (1) 0.481 ± 0.133 (7) 0.568 ± 0.041 (5) 0.640 ± 0.002 (3) 0.615 ± 0.008 (4) 0.540 ± 0.006 (6) 0.645 ± 0.019 (2)
0.583 ± 0.003 (1) 0.350 ± 0.123 (6) 0.525 ± 0.043 (2) 0.525 ± 0.029 (3) 0.458 ± 0.017 (4) 0.429 ± 0.036 (5) 0.344 ± 0.043 (7)
0.436 (1) 0.308 (7) 0.408 (2) 0.400 (3) 0.387 (4) 0.359 (5) 0.349 (6)
mvit
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.311 ± 0.007 (3) 0.164 ± 0.105 (7) 0.318 ± 0.007 (1) 0.318 ± 0.006 (2) 0.265 ± 0.056 (5) 0.310 ± 0.006 (4) 0.230 ± 0.020 (6)
0.176 ± 0.001 (2) 0.156 ± 0.006 (7) 0.208 ± 0.009 (1) 0.170 ± 0.001 (3) 0.165 ± 0.002 (4) 0.162 ± 0.005 (5) 0.161 ± 0.003 (6)
0.676 ± 0.001 (1) 0.511 ± 0.071 (7) 0.619 ± 0.024 (3) 0.586 ± 0.009 (5) 0.588 ± 0.012 (4) 0.529 ± 0.004 (6) 0.644 ± 0.008 (2)
0.601 ± 0.005 (1) 0.400 ± 0.084 (6) 0.433 ± 0.039 (4) 0.520 ± 0.007 (2) 0.433 ± 0.018 (3) 0.415 ± 0.003 (5) 0.367 ± 0.026 (7)
0.441 (1) 0.308 (7) 0.394 (3) 0.399 (2) 0.363 (4) 0.354 (5) 0.351 (6)
0.310 ± 0.006 (1) 0.205 ± 0.111 (7) 0.309 ± 0.006 (2) 0.306 ± 0.004 (4) 0.277 ± 0.048 (5) 0.309 ± 0.013 (3) 0.259 ± 0.021 (6)
0.176 ± 0.001 (2) 0.160 ± 0.007 (7) 0.205 ± 0.008 (1) 0.169 ± 0.000 (3) 0.166 ± 0.002 (5) 0.163 ± 0.010 (6) 0.167 ± 0.005 (4)
0.676 ± 0.001 (1) 0.527 ± 0.075 (7) 0.611 ± 0.006 (3) 0.583 ± 0.008 (5) 0.584 ± 0.012 (4) 0.542 ± 0.004 (6) 0.650 ± 0.019 (2)
0.605 ± 0.005 (1) 0.442 ± 0.070 (5) 0.432 ± 0.025 (6) 0.528 ± 0.008 (2) 0.457 ± 0.014 (4) 0.472 ± 0.026 (3) 0.425 ± 0.027 (7)
0.442 (1) 0.333 (7) 0.389 (3) 0.397 (2) 0.371 (6) 0.372 (5) 0.375 (4)
sf
AR-Net MGFN RTFM Sultani UR-DMU VadCLIP GCN-Anomaly
Dynamic MIL Magnitude MIL Magnitude MIL Vanilla MIL Uncertainty-Aware MIL Language-Guided MIL Label Denoising
0.354 ± 0.002 (2) 0.240 ± 0.066 (7) 0.364 ± 0.025 (1) 0.294 ± 0.008 (5) 0.305 ± 0.014 (4) 0.308 ± 0.015 (3) 0.294 ± 0.024 (6)
0.200 ± 0.003 (2) 0.169 ± 0.009 (6) 0.224 ± 0.010 (1) 0.191 ± 0.002 (3) 0.189 ± 0.020 (4) 0.153 ± 0.007 (7) 0.169 ± 0.007 (5)
0.669 ± 0.002 (1) 0.577 ± 0.065 (6) 0.658 ± 0.021 (2) 0.650 ± 0.003 (3) 0.600 ± 0.014 (5) 0.547 ± 0.008 (7) 0.632 ± 0.019 (4)
0.559 ± 0.005 (1) 0.361 ± 0.098 (6) 0.462 ± 0.060 (3) 0.514 ± 0.007 (2) 0.410 ± 0.015 (4) 0.393 ± 0.025 (5) 0.319 ± 0.031 (7)
0.445 (1) 0.337 (7) 0.427 (2) 0.412 (3) 0.376 (4) 0.350 (6) 0.354 (5)
0.355 ± 0.002 (2) 0.270 ± 0.048 (7) 0.360 ± 0.023 (1) 0.297 ± 0.008 (6) 0.297 ± 0.016 (5) 0.307 ± 0.015 (3) 0.307 ± 0.018 (4)
0.200 ± 0.003 (2) 0.173 ± 0.006 (5) 0.213 ± 0.013 (1) 0.191 ± 0.003 (3) 0.186 ± 0.014 (4) 0.152 ± 0.006 (7) 0.164 ± 0.011 (6)
0.666 ± 0.002 (1) 0.565 ± 0.111 (6) 0.643 ± 0.026 (4) 0.648 ± 0.002 (3) 0.609 ± 0.019 (5) 0.544 ± 0.009 (7) 0.663 ± 0.017 (2)
0.562 ± 0.005 (1) 0.375 ± 0.055 (6) 0.476 ± 0.016 (3) 0.513 ± 0.007 (2) 0.402 ± 0.016 (4) 0.395 ± 0.024 (5) 0.360 ± 0.025 (7)
0.445 (1) 0.346 (7) 0.423 (2) 0.412 (3) 0.374 (4) 0.350 (6) 0.373 (5)