arXiv:2609.24980v1 [cs.CR] 21 Sep 2026
Residual Community Prototypes Under-Reject Held-Out Malware Families in FCG-MFD 1st Junru Zhu
2nd Yixin Yang
3rd Xiaoqing Ding
4th Ruoyu Qi
Independent Researcher Seattle, WA, USA [email protected]
Independent Researcher New York, NY, USA [email protected]
University of Chicago Chicago, IL, USA [email protected]
Independent Researcher Charlotte, NC, USA [email protected]
Abstract—Open-set malware-family recognition must classify known families while rejecting families absent from training. We test whether Louvain-community summaries add rejection information beyond a graph neural network embedding and dimension-matched generic topology. The study uses a deduplicated, conflict-audited FCG-MFD corpus, five held-out families, and three optimization seeds. Community features are residualized against generic topology using known-family training data before nearest-prototype scoring. Residual community does not produce stable held-out-family rejection. Ranking effects reverse across families, the false-positive rate at 95 percent unknown recall worsens for every held-out family, and a validationfitted threshold rejects only 4.48 percent of unknown samples. Accepted-known macro F1 improves in every family, but with five independent family units the exact two-sided sign-flip pvalue is 0.0625, the smallest attainable value. The score remains associated with graph scale, while simple classifier uncertainty performs better on ranking, high-recall rejection, and OSCR. In this GIN/FCG-MFD setting, community-enriched prototypes change known-class geometry without creating a stable unknown margin. Graph open-set evaluations should pair structural features with matched topology controls, operational thresholds, and held-out-family analysis. Index Terms—malware family recognition, open-set recognition, graph neural networks, function-call graphs, community structure
I. Introduction Malware-family classification supports the reuse of analysis across related samples, but deployed systems also encounter families absent from training. Function-call graphs (FCGs) represent program structure, and graph neural networks (GNNs) can learn family representations from these graphs [1], [2]. Closed-set accuracy alone does not measure whether such a representation can recognize when a sample falls outside its known-family support. Open-set recognition (OSR) adds this rejection requirement. Activation-space models and metric-learning methods assign unknown scores from the relation between a test representation and known-class support [3], [4]. Malware studies have applied FCG transformations, learned distances, and contrastive representations to unknown or shifted samples [5]–[7]. These methods make the representation central to rejection: an added graph descriptor is useful only when it separates unseen families beyond information already present in the embedding and generic graph topology.
Community organization is a candidate descriptor because modularity, community-size distributions, and crosscommunity connectivity summarize mesoscopic structure. The same summaries, however, vary with graph size, density, and degree. A prototype can therefore respond to structural scale rather than family novelty. This confounding is especially important under held-out-family evaluation, where the size distribution itself may shift. A controlled test must compare community information with generic topology and examine operating thresholds, scale dependence, and family-specific effects rather than relying on a pooled ranking metric. We conduct this test on a cleaned FCG-MFD corpus [8]. Exact graph conflicts are removed as groups, cross-label Weisfeiler–Lehman signatures are excluded, and same-label signatures are collapsed before splitting. Five malware families are held out in turn, with three optimization seeds per family. A common multiclass GIN provides the graph embedding. The primary comparison augments this embedding with either an eight-dimensional generic-topology block or an eightdimensional community block residualized against generic topology on known-family training data. In this GIN/FCG-MFD setting, the embedding plus residualcommunity prototype does not establish reliable unknownfamily rejection. Its mean AUROC and AUPRC changes are small relative to the dimension-matched generic-topology control. FPR@95TPR is worse across the held-out-family means, mean OSCR does not improve, and the validation-fitted threshold accepts 95.52% of held-out-family samples. The mean known-family macro F1 is 0.0082 higher, but family-clustered bootstrap uncertainty includes zero. Scores remain associated with graph scale, and effects reverse between winwebsec and WannaCry_Worm. Together, these observations locate the failure in the tested representation and operating regime rather than supporting a general benefit from community features. Our contributions are: • a leakage-audited FCG-MFD protocol with deterministic graph-group splitting, five held-out families, three optimization seeds, and retained sample-level scores; • a dimension-matched comparison that tests whether residual community structure adds unknown-family information beyond a learned embedding and generic topology; • a threshold-, scale-, and family-aware negative result showing that the tested structural prototype under-rejects
held-out families and trails classifier-uncertainty baselines in this GIN/FCG-MFD setting. II. Related Work A. Graph Representations for Malware Families Graph-based malware learning represents programs through call and control flow rather than only flat feature vectors [1]. Function call graph embeddings have been used for Android malware detection and family categorization [9], while hierarchical GNNs combine inter-function call graphs with intrafunction control-flow graphs [2]. Contrastive objectives have also been applied to malware familial classification [7]. FCGMFD extends this line with a function-call-graph benchmark spanning malware families [8]. These studies establish graph encoders for closed-set prediction. Our study instead holds a common encoder fixed and tests whether an added community block supports rejection of a family absent from training. B. Open-Set and Shift-Aware Malware Recognition General OSR methods derive unknown scores from class activation tails or distance in a learned representation space [3], [4]. The closest graph-specific malware study learns representations from transformed FCGs and applies distance-based rejection to unseen families [5]. CADE similarly learns a contrastive distance for detecting and explaining individual drift samples [6]. CNS-Net synthesizes conservative novelty examples for broader open-set malware recognition [10]. These methods establish rejection strategies; our novelty claim is limited to the matched structural control and diagnostic failure analysis, not a new rejection rule. Malware reliability research has also examined predictive uncertainty and selective classification under dataset, adversarial, and temporal shift [11]–[13]. TESSERACT further shows that spatial and temporal split choices can bias malware-classification evaluation [14]. These lines motivate classifier-uncertainty baselines, held-out-family reporting, and operating-point metrics. Our comparison asks a narrower structural question: after an embedding and generic topology are available, does residual community information improve family rejection? C. Community Structure and Scale Controls Louvain detects communities by optimizing modularity through local moves and graph coarsening [15]. Modularity optimization has a resolution limit: the smallest recoverable modules depend on total network size and their interconnection pattern [16]. Community count, modularity, and sizedistribution summaries are therefore not independent of graph scale. Under held-out-family shift, these quantities can change prototype distance without encoding family-specific novelty. We isolate this ambiguity by residualizing an eightdimensional community block against generic topology on known-family training data and comparing it with an eightdimensional generic-topology control. Score–size correlations and size-stratified metrics then test whether rejection behavior remains associated with structural scale.
III. Method A. Study Design and Leakage Controls We use the FCG-MFD corpus [8] to evaluate whether community structure adds open-set information beyond an FCG embedding and generic topology. The source archive contains 46,714 graph rows. Cleaning precedes all splitting: empty and edgeless graphs are removed, exact canonical graphs with conflicting family labels are dropped as groups, cross-label Weisfeiler–Lehman signatures are excluded, and same-label signatures are collapsed to one representative. Benign graphs are excluded because the task is malware-family recognition. The resulting model pool contains 20,870 structural signatures across 32 labels. Twenty-three malware families meet the 100-signature modeling threshold. Five families selected before model outcomes are held out in turn: BazaLoader (859 graphs), Mirai Botnet (1,212), Zeus Trojan (1,045), winwebsec (2,996), and WannaCry_Worm (2,714). Each fold assigns the entire held-out family to unknown testing. Within every remaining family, a deterministic graph-group split assigns 70% to training, 15% to validation, and 15% to known testing. The partitions are fixed across seeds. The unknown family is absent from encoder training, checkpoint selection, feature scaling, residualization, prototype fitting, and threshold selection. The encoder and structural score models use known-family training data; checkpoint selection and rejection thresholds use known-family validation data. The unknown-family split is accessed only for final evaluation. B. Graph and Structural Representations Directed call edges define five node features: log in-degree, log out-degree, log total degree, and source and sink indicators. Message passing uses the symmetrized edge set. A three-layer GIN [17] with 64 hidden units and a 64dimensional graph embedding applies global mean and max pooling. Training uses class-weighted cross entropy, dropout 0.2, Adam with learning rate 10−3 and weight decay 10−4 , batches of 64 graphs, and early stopping with patience 12 within an 80-epoch cap. The selected checkpoint maximizes known-validation macro F1, with negative log-likelihood as the tie-breaker. Optimization seeds are 7, 23, and 47. Structural covariates are computed on the undirected simple graph induced by the FCG. The eight-dimensional generictopology block contains log node and directed-edge counts, undirected density, mean and standard deviation of degree, maximum-degree ratio, component ratio, and transitivity. Louvain community detection uses resolution 1 and seed 0 [15]. Its eight-dimensional block contains modularity, log community count, community-count ratio, normalized community-size entropy, maximum-community ratio, intracommunity edge ratio, bridge-node ratio, and size-weighted within-community density. C. Dimension-Matched Prototype Scores Let z(x) ∈ R64 denote the GIN embedding, u(x) ∈ R8 the generic block, and q(x) ∈ R8 the community block. Trainingset standardizers Su and Sq put the structural coordinates on
comparable scales. A multi-output ridge model g with penalty 1 predicts standardized community structure from standardized generic topology. The residual is r(x) = Sq (q(x)) − g(Su (u(x))) .
(1)
Thus, residualization removes the linear component predictable from generic topology in known-family training data; it does not impose independence under held-out-family shift. For a feature space with blocks hb (x) of dimensions db , each √ block is standardized on training data and scaled by 1/ db before concatenation. The prototype and unknown score are µc =
1 X h̄(xi ), |Tc | xi ∈Tc
s(x) = min h̄(x) − µc 2 , (2) c
TABLE I Mean performance over five held-out families and three seeds. Classifier uncertainty is stronger than the concatenated prototypes on unknown ranking and OSCR. Higher is better except FPR@95TPR. Method
AUROC AUPRC FPR95 OSCR
MSP Entropy Energy Embed.+generic Embed.+resid.
0.6013 0.6149 0.6072 0.5196 0.5319
D. Metrics and Family-Clustered Analysis Unknown AUROC and AUPRC measure ranking. FPR@95TPR reports the known-sample false-positive rate when unknown recall reaches 95%; lower is better. OSCR combines correct known classification with unknown rejection [19]. We also report known-family macro F1, known acceptance, and unknown rejection at the validation-fitted threshold. Each metric is computed within a held-out-family and optimization-seed cell. Seeds are repeated optimization runs, whereas the five held-out families are the units of generalization. For the primary comparison, we compute paired method differences within each cell, average the three seeds within each family, and obtain 95% intervals by resampling the five family means 10,000 times with replacement [20]. Oriented effects are positive when residual community is better; FPR@95TPR differences are sign-reversed. Absolute means over the 15 cells give each family equal weight because every family has three seeds. Because only five family units are available, percentile intervals are descriptive. We also enumerate all 32 sign assignments to the five family means for an exact two-sided sign-flip test; the smallest attainable p-value is 0.0625.
0.4191 0.7152 0.2839 0.2479 0.4262 0.7101 0.2822 0.2479 0.4322 0.7742 0.2558 0.2479 0.3719 0.8941 0.1814 0.2061 0.3820 0.9160 0.1725 0.2143
TABLE II Residual-community versus generic-topology paired effects. Descriptive intervals resample the five seed-averaged held-out-family effects 10,000 times. Positive values favor residual community; FPR@95TPR is sign-reversed. Metric
where Tc is the training set for known family c. The scaling prevents the 64-dimensional embedding from dominating an eight-dimensional structural block solely through dimension. The primary comparison uses h̄ = [z, u] for the generictopology control and h̄ = [z, r] for the residual-community score. We additionally evaluate MSP using 1 − maxc p(c | x), normalized predictive entropy, energy [18], and prototypes built from the embedding, generic block, or raw community block alone. Larger scores always indicate stronger evidence for rejection. For each method, the rejection threshold is the empirical 95th percentile of its known-validation scores, targeting 95% known-family acceptance without using unknown samples.
F1
Mean
95% interval Fam. +
Unknown AUROC +0.0123 [−0.0408, +0.0686] Unknown AUPRC +0.0101 [−0.0241, +0.0448] FPR@95TPR −0.0219 [−0.0371, −0.0069] OSCR −0.0090 [−0.0302, +0.0162] Known F1 +0.0082 [−0.0025, +0.0177] Accepted F1 +0.0138 [+0.0088, +0.0196]
3/5 4/5 0/5 1/5 4/5 5/5
IV. Experiments and Analysis A. Classifier Uncertainty Is the Stronger Baseline Table I reports means over the 15 family–seed cells. The three classifier-uncertainty scores outperform both concatenated prototype scores on unknown AUROC and OSCR. Entropy attains the highest AUROC (0.6149) and lowest FPR@95TPR (0.7101), while MSP attains the highest OSCR (0.2839). The residual-community prototype reaches 0.5319 AUROC, 0.9160 FPR@95TPR, and 0.1725 OSCR. Its ranking is therefore weaker than simple uncertainty, and its high FPR shows that reaching 95% unknown recall requires rejecting most known samples. B. Residual Community Does Not Improve Family-Level Rejection Table II gives the primary paired comparison at the heldout-family level. Residual community changes unknown AUROC by +0.0123 and AUPRC by +0.0101, but both familyclustered intervals include zero. OSCR changes by −0.0090 and improves for only one of five family means. The clearest adverse result is FPR@95TPR: its oriented effect is −0.0219, the interval remains below zero, and all five held-out families favor the generic-topology control. Its exact family-level signflip p-value is 0.0625, so the unanimous direction is suggestive rather than conventionally significant. Figure 1 shows why the mean ranking changes are inconclusive. The AUROC effect ranges from −0.0767 for WannaCry_Worm to +0.1132 for winwebsec. AUPRC exhibits the same reversal, and winwebsec is the only family with a positive OSCR effect. Averaging these opposing cases produces a small positive AUROC mean without a family-general improvement.
Residual community better
Generic topology better
(a) Unknown AUROC BazaLoader
-0.001
Mirai Botnet
+0.002
WannaCry Worm
(b) Unknown AUPRC +0.006 +0.013
-0.077
-0.048 +0.024
Zeus Trojan
+0.012 +0.113
winwebsec
+0.068
(c) FPR@95TPR
(d) OSCR
-0.022
BazaLoader
-0.016
-0.004
Mirai Botnet
-0.010
-0.050
WannaCry Worm
-0.042 -0.002
Zeus Trojan
-0.014
-0.031
winwebsec −0.08
−0.04 0.00 0.04 0.08 Oriented paired effect (positive is better)
+0.037 0.12
−0.08
−0.04 0.00 0.04 0.08 Oriented paired effect (positive is better)
0.12
Fig. 1. Seed-averaged paired effects by held-out family for residual community versus generic topology. FPR@95TPR is sign-reversed, so rightward values always favor residual community. Ranking effects reverse across families, while FPR@95TPR is worse for every held-out family.
Classification and rejection also separate. Known-family macro F1 changes from 0.2061 to 0.2143, but its familyclustered interval includes zero. Among accepted known samples, macro F1 increases by 0.0138 with a positive descriptive interval and improves for all five families. Its exact two-sided sign-flip p-value is also 0.0625. This conditional gain concerns classification after acceptance; it does not establish that the score separates unknown families. C. Validation Thresholds Accept Nearly All Unknowns At the known-validation threshold, residual community rejects 4.48% of held-out-family samples, compared with 3.92% for generic topology. The paired rejection increase is 0.56 percentage points [0.01, 1.09] and is positive for four families, but the absolute operating point remains poor: 95.52% of unknown samples are accepted. Known acceptance is 94.93%, close to the intended 95%, so the failure is not a missed known-acceptance target. Instead, the known- and unknownscore distributions overlap at the validation-selected boundary. The misses are not concentrated near that boundary: accepted unknowns lie 1.093 score units below the residual threshold on average versus 0.839 for generic topology, and the mean cell-wise maximum margin is 1.608 versus 1.270. Residual community moves only a small subset across the boundary while leaving most misses deep inside the acceptance region. The shared classifier is also weak: MSP-based known macro F1 is 0.2479, and the residual-community prototype obtains
TABLE III Unknown-sample decision transitions at the known-validation threshold. The two structural scores agree on 95.86% of pooled family–seed predictions, mostly by accepting the unknown sample.
Decision pair
Count
Fraction
Both accept Both reject Residual rejects only Generic rejects only
24,808 573 613 484
93.69% 2.16% 2.32% 1.83%
0.2143. This limits operational interpretation because OSCR requires both correct known classification and unknown rejection. It does not explain away the controlled comparison, however: all methods use the same encoder, and classifier uncertainty still achieves higher AUROC, lower FPR@95TPR, and higher OSCR. D. Decision Transitions Localize the Difference The decision-transition audit in Table III pools 26,478 unknown-sample predictions from the 15 family–seed cells. The counts represent model decisions rather than independent graph draws because each test graph is evaluated by three seed-specific models. We therefore use this audit descriptively; family-level sign patterns and exact tests bound the smallsample interpretation.
TABLE IV Known-sample transitions for predictions that are both accepted and assigned to the correct family. Counts pool 41,133 family–seed predictions. Residual community improves macro F1 without increasing the total number of accepted-and-correct decisions.
Decision pair
Count Fraction
Both correct 9,488 Neither correct 27,366 Residual correct only 1,925 Generic correct only 2,354
23.07% 66.53% 4.68% 5.72%
Size-stratified FPR@95TPR sharpens this pattern. Residual community improves over generic topology by 0.0397 in the smallest node-count quartile, then worsens by 0.0616, 0.0215, and 0.0235 in the next three quartiles. These observations do not identify graph size as the sole cause of acceptance, but they show that the added coordinates are not scale-neutral in the tested score. Together with the winwebsec–WannaCry reversal, they bound the result: the representation has family-specific effects rather than a stable unknown-family signature. V. Discussion A. Classification and Rejection Use Different Geometry
Residual-only rejections exceed generic-only rejections by 129 predictions, or 0.49 percentage points in this sampleweighted pool. The direction is not uniform across families. Residual-only rejection is more frequent in four families, while WannaCry_Worm has 86 residual-only and 115 generic-only rejections. Winwebsec contributes 290 of the 613 residualonly cases. The small aggregate increase thus comes from limited, family-concentrated movements across the threshold rather than broad separation of unknown samples from known support. Known-sample transitions provide a complementary view of the accepted-known macro F1 result. The two scores agree on whether a sample is accepted and correct in 89.60% of pooled decisions. Generic-only correct decisions exceed residual-only decisions by 429, or 1.04 percentage points. This does not conflict with the positive macro F1 effect: macro F1 gives equal weight to class-specific precision and recall, whereas the transition total weights every sample equally. The residual block therefore redistributes known-class errors without increasing the total number of accepted-and-correct predictions. The redistribution is also family-dependent. Residual-only correct decisions exceed generic-only decisions for winwebsec (406 versus 320), but the generic control has more in BazaLoader (509 versus 361), Mirai Botnet (476 versus 425), WannaCry_Worm (459 versus 369), and Zeus Trojan (590 versus 364). Together, the known and unknown transition audits show that the feature block changes a limited set of decisions and that those changes remain concentrated by heldout family.
Nearest-prototype classification depends on which class centroid is closest, whereas rejection depends on the absolute distance to that closest centroid. An added block can therefore improve accepted-known assignments without moving unknowns outside known support. Residual community improves accepted-known macro F1 in every family, yet AUROC, AUPRC, and OSCR are not stable across families and FPR@95TPR deteriorates in all five. Both unanimous directions have exact two-sided p-value 0.0625, the finest resolution available here. The isotropic Euclidean score has no objective that reserves an unknown margin; its F1 gain is evidence about conditional classification, not open-set separation. B. Two Operating Points Reveal Score Overlap FPR@95TPR is a diagnostic operating point selected with test labels: the residual-community score must falsely reject 91.60% of known samples to recover 95% of unknown samples. The deployable threshold is selected only from known validation data. It retains 94.93% known acceptance but rejects only 4.48% of unknown samples. The score therefore provides no useful trade-off near either high unknown recall or high known acceptance. AUROC averages ranking over every threshold and can improve when a subset of unknown samples moves upward, even if the operational boundary remains inside heavy known– unknown overlap. OSCR also requires correct known classification and degrades on average [19]. The winwebsec– WannaCry reversal further shows why a pooled mean is insufficient: a deployment encounters one unseen family, while the five-family analysis remains descriptive.
E. Scale Dependence and Family Reversals
C. Residualization Does Not Imply Scale Invariance
The residual-community score remains associated with structural scale after training-set residualization. On known samples, its mean Spearman correlations with node count, directed-edge count, and community count are −0.405, −0.381, and −0.406; on unknown samples they are −0.410, −0.391, and −0.402. The generic-topology control has substantially smaller absolute correlations, ranging from 0.021 to 0.060 on known samples and 0.127 to 0.148 on unknown samples. Residualization removes a linear training-set prediction, but the resulting prototype distance still tracks scale under held-out-family shift.
The ridge residual removes the linear component of standardized community features predicted by generic topology in known-family training data. It does not remove nonlinear dependence or held-out-family shift, and the minimumdistance operation can restore scale association. This matters for Louvain summaries because modularity resolution depends on network size and interconnection [16]. Correlations near −0.4 and deterioration in the three larger node-count quartiles indicate remaining scale sensitivity, not a causal size mechanism. The dimension-matched generic control therefore remains essential.
D. Designing the Next Test A stronger test should separate representation quality, calibration, and incremental structural information. It should first establish a stronger known-family representation, using rejection-aware metric learning or contrastive pretraining [7]. A nested protocol can use development-only unknown families for calibration while preserving distinct final families. Community effects should also be tested in node-count and density-matched strata, with family-level endpoints retained. VI. Limitations The study covers structural FCGs, five held-out families, and one GIN encoder; richer program semantics or other representations may yield different score geometry. Five family units limit exact two-sided inference to p-values no smaller than 0.0625. Cleaning ambiguous and duplicate graphs changes the source distribution, two cells reach the 80-epoch cap, and score–size correlations are diagnostic rather than causal. VII. Conclusion We tested whether residual Louvain-community summaries add open-set information beyond a GIN embedding and dimension-matched generic topology in FCG-MFD. They do not provide stable held-out-family rejection in this tested setting. AUROC and AUPRC effects remain uncertain across held-out families, FPR@95TPR worsens for all five family means, OSCR does not improve, and the validation-fitted threshold accepts 95.52% of unknown samples. The consistent accepted-known F1 gain shows that the added coordinates can refine classification among accepted examples without creating an unknown margin. The score also remains associated with structural scale, and its effect reverses between winwebsec and WannaCry_Worm. In this GIN/FCG-MFD study, communityenriched prototypes change known-class geometry without yielding stable open-set separation. Graph OSR evaluations should pair structural features with dimension-matched topology controls, operational thresholds, and held-out-family analysis. References [1] T. Bilot, N. El Madhoun, K. Al Agha, and A. Zouaoui, “A survey on malware detection with graph representation learning,” ACM Computing Surveys, vol. 56, no. 11, pp. 1–36, 2024. [2] X. Ling, L. Wu, W. Deng, Z. Qu, J. Zhang, S. Zhang, T. Ma, B. Wang, C. Wu, and S. Ji, “MalGraph: Hierarchical graph neural networks for robust windows malware detection,” in IEEE INFOCOM 2022, 2022, pp. 1998–2007. [3] A. Bendale and T. E. Boult, “Towards open set deep networks,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 1563–1572. [4] M. Hassen and P. K. Chan, “Learning a neural-network-based representation for open set recognition,” in Proceedings of the 2020 SIAM International Conference on Data Mining, 2020, pp. 154–162. [5] J. Jia and P. K. Chan, “Representation learning with function call graph transformations for malware open set recognition,” in 2022 International Joint Conference on Neural Networks, 2022, pp. 1–8. [6] L. Yang, W. Guo, Q. Hao, A. Ciptadi, A. Ahmadzadeh, X. Xing, and G. Wang, “CADE: Detecting and explaining concept drift samples for security applications,” in 30th USENIX Security Symposium, 2021.
[7] Y. Wu, S. Dou, D. Zou, W. Yang, W. Qiang, and H. Jin, “Contrastive learning for robust android malware familial classification,” IEEE Transactions on Dependable and Secure Computing, 2024. [8] H. J. Hadi, Y. Cao, S. Li, N. Ahmad, and M. A. Alshara, “FCGMFD: Benchmark function call graph-based dataset for malware family detection,” Journal of Network and Computer Applications, vol. 233, p. 104050, 2025. [9] P. Xu, C. Eckert, and A. Zarras, “Detecting and categorizing android malware with graph neural networks,” in Proceedings of the 36th Annual ACM Symposium on Applied Computing, 2021, pp. 409–412. [10] J. Guo, S. Guo, S. Ma, Y. Sun, and Y. Xu, “CNS-Net: Conservative novelty synthesizing network for malware recognition in an open-set scenario,” arXiv preprint arXiv:2305.01236, 2023. [11] D. Li, T. Qiu, S. Chen, Q. Li, and S. Xu, “Can we leverage predictive uncertainty to detect dataset shift and adversarial examples in android malware detection?” in Annual Computer Security Applications Conference, 2021, pp. 596–608. [12] H. Li, G. Xu, L. Wang, X. Xiao, X. Luo, G. Xu, and H. Wang, “MalCertain: Enhancing deep neural network based android malware detection by tackling prediction uncertainty,” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering, 2024, pp. 1–13. [13] A. Herzog, A. Eusebi, and L. Cavallaro, “On the reliability and stability of selective methods in malware classification tasks,” arXiv preprint arXiv:2505.22843, 2025. [14] F. Pendlebury, F. Pierazzi, R. Jordaney, J. Kinder, and L. Cavallaro, “TESSERACT: Eliminating experimental bias in malware classification across space and time,” in 28th USENIX Security Symposium, 2019, pp. 729–746. [15] V. D. Blondel, J.-L. Guillaume, R. Lambiotte, and E. Lefebvre, “Fast unfolding of communities in large networks,” Journal of Statistical Mechanics: Theory and Experiment, vol. 2008, no. 10, p. P10008, 2008. [16] S. Fortunato and M. Barthélemy, “Resolution limit in community detection,” Proceedings of the National Academy of Sciences, vol. 104, no. 1, pp. 36–41, 2007. [17] K. Xu, W. Hu, J. Leskovec, and S. Jegelka, “How powerful are graph neural networks?” in International Conference on Learning Representations, 2019. [18] W. Liu, X. Wang, J. D. Owens, and Y. Li, “Energy-based out-ofdistribution detection,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 21 464–21 475. [19] A. R. Dhamija, M. Günther, J. Ventura, and T. E. Boult, “Reducing network agnostophobia,” in Advances in Neural Information Processing Systems, vol. 31, 2018. [20] B. Efron, “Bootstrap methods: Another look at the jackknife,” The Annals of Statistics, vol. 7, no. 1, pp. 1–26, 1979.