arXiv:2609.24389v1 [cs.CR] 21 Sep 2026
Name2Pkg: Lightweight One-Class Android Malware Screening via Name-Package Correspondence Modeling Changyeop Sung
Yeonjae Kang
Jaeho Shin
Huy Kang Kim
School of Cybersecurity Korea University Seoul, Republic of Korea [email protected]
School of Cybersecurity Korea University Seoul, Republic of Korea [email protected]
IT Planning Department Hana Bank Seoul, Republic of Korea [email protected]
School of Cybersecurity Korea University Seoul, Republic of Korea [email protected]
Abstract—Deep learning-based malware detection has been widely adopted in security-critical services. Most detection methods rely on internal features extracted from APK files or runtime behavior. However, extracting these features is computationally expensive. This limits their use in large-scale, early-stage screening. Malicious apps may exhibit weak correspondence between their user-facing app names and package names, providing a low-cost screening signal. We present Name2Pkg, a lightweight one-class classification method. It leverages only the app name and the package name. We formulate malware screening as a sequence anomaly detection problem. A character-level sequenceto-sequence model estimates the conditional likelihood of a package name given the app name. The length-normalized negative log-likelihood serves as the anomaly score. We train the model and calibrate the threshold using only benign data. Using a dataset of 67,129 real-world applications, Name2Pkg achieves an area under the receiver operating characteristic curve (ROCAUC) of 0.982 and malware recall of 0.885 at an achieved falsepositive rate of 0.044 on held-out test data. It has a 3.57 MiB checkpoint and a CPU inference latency of 28.20 ms per sample. Name2Pkg provides an efficient and effective pre-filtering signal for large-scale security systems. Index Terms—Android malware screening, One-class anomaly detection, Name-package correspondence
I. I NTRODUCTION Android malware detection commonly relies on static inspection of artifacts within Android application packages (APKs), dynamic analysis, or hybrid approaches [1]–[3]. Although these methods can achieve strong detection performance, they often require extracting permissions and API-call features, analyzing bytecode or graphs, or collecting runtime behavior. These requirements make them less suitable for large-scale triage and early-stage screening, where a low-cost signal is needed before deeper analysis. Prior work has therefore explored lightweight models, compact representations, and reduced feature sets [4]–[6]. App-market metadata and identifier strings provide another source of low-cost security signals [7]–[9]. SeqDroid, for example, learns representations from package names and other metadata strings together with permissions and intent actions [10]. Studies of fake apps and app squatting also show
that app names and package names can be manipulated or diverge in identity-abuse scenarios [11], [12]. This suggests that the relationship between the two identifiers may itself provide a screening signal. However, existing methods generally combine identifiers with additional features, treat them as independent attributes, or compare apps against known references. The correspondence between the app name and package name within a single app remains underexplored as a malware-screening signal. In this paper, we focus on two low-cost textual identifiers: the app name and the package name. A package name is a namespace-style identifier commonly exposed through Android APIs and app metadata [13], whereas the app name is the primary user-facing label. Because both identifiers can be obtained from app-market records or lightweight manifestlevel metadata without bytecode analysis or runtime execution, they are natural inputs for low-cost early-stage screening. A key challenge is that package names alone can already provide a useful signal. Some malicious apps use unusual, random-looking, or weakly meaningful package strings, and such patterns may be detected even without considering the app name. However, package-name morphology alone does not fully capture the relationship between what an app claims to be and how it is identified. A package name may look plausible in isolation but still be weakly related to the corresponding app name. Therefore, an identifier-level detector should model not only whether a package name looks normal, but also whether it is plausible given the app name. This restricted input also fits a one-class setting, in which benign regularities are learned without requiring complete or stable malware labels [14]–[17]. We propose Name2Pkg, a lightweight one-class screening method based on name-conditioned package scoring. Name2Pkg uses a character-level sequence-to-sequence model trained only on benign app-name and package-name pairs [18]–[20]. Given an app name, the model estimates the conditional likelihood of the corresponding package name. The average negative log-likelihood of the package-name sequence is used as the anomaly score, with higher values indicating
greater deviation from benign name-package regularities. The decision threshold is selected exclusively from a separate benign calibration set, and malware samples are used only for final evaluation. Name2Pkg is intended as a pre-analysis filter that prioritizes suspicious apps for more expensive static, dynamic, or hybrid analysis. We evaluate Name2Pkg on a dataset collected by a commercial bank in South Korea, consisting of 62,730 benign apps and 4,399 malware samples. It achieves an area under the receiver operating characteristic curve (ROC-AUC) of 0.982 and, at a target false-positive rate (FPR) of 0.05, an achieved FPR of 0.044 on the benign test set and malware recall of 0.885. A package-only control reaches a ROC-AUC of 0.930 and recall of 0.692, showing that package-name morphology is useful but does not account for Name2Pkg’s full performance. In a shuffled-pair control, breaking valid benign name-package correspondence increases anomaly scores. Name2Pkg has a 3.57 MiB checkpoint size and 28.20 ms CPU latency per sample. The main contributions of this paper are as follows: • We formulate within-app name-package correspondence as a lightweight anomaly signal for Android malware screening. • We develop Name2Pkg, a benign-only character-level sequence model that estimates conditional package-name likelihood and calibrates thresholds using benign data only. • We evaluate Name2Pkg against identifier-level baselines, isolate the contribution of name-package correspondence through package-only and shuffled-pair controls, and report checkpoint size and CPU inference latency. The remainder of this paper is organized as follows. Section II reviews related work. Section III details the Name2Pkg methodology. Section IV presents the experimental setup, evaluation results, and ablation studies. Finally, Section V concludes the paper. II. R ELATED W ORK A. Static, Dynamic, and Hybrid Android Malware Detection Android malware detection has been extensively studied through static, dynamic, and hybrid analyses of APK artifacts and runtime behavior [1]–[3]. Static methods inspect manifests, permissions, API calls, bytecode, or graph representations without executing the app. DREBIN uses features extracted from Android applications for static malware detection [1], while GSEDroid represents apps using API-call graphs enriched with permission and opcode-semantic information [21]. Dynamic methods execute apps in controlled environments and observe behaviors such as system calls, file and network operations, and sensitive API usage. AppsPlayground and CopperDroid are representative systems for automated runtime analysis and behavior reconstruction [2], [22]. Hybrid methods combine static and dynamic evidence. DeepAMD analyzes features from both layers using a deep neural network [23],
whereas MPDroid integrates static and dynamic representations within a multimodal framework [3]. Although these approaches provide rich structural and behavioral evidence, they require APK-internal feature extraction, graph construction, controlled execution, or a combination of these operations. Name2Pkg is intended as a lower-cost pre-analysis signal that can be applied before such methods. B. Lightweight Android Malware Detection To reduce the cost of conventional analysis, prior work has explored compact models and restricted feature sets. Krzyszton et al. developed an on-device detector using features obtained through the Koodous platform [24]. RevealDroid uses a lightweight, obfuscation-resilient representation for malware detection and family identification [4]. Ma et al. proposed a lightweight two-layer detection framework [6], while PacDroid uses selected permissions, Intent Actions, and Intent Categories extracted from Android manifest files [5]. These methods reduce computational cost by limiting model complexity or narrowing the feature space. However, they generally continue to depend on APK-derived information, including manifest entries, permissions, intents, API usage, native-code indicators, or bytecode-level features. Name2Pkg targets a more restricted setting in which screening requires only the app name and package name. Its lightweight design therefore applies both to model size and to the amount of application information required at inference time. C. Metadata, Identifier-Level Signals, and App Identity Abuse App-market metadata can provide useful signals before detailed APK analysis. Teufl et al. used descriptions, permissions, ratings, and developer information for pre-installation malware detection [7]. Munoz et al. identified predictive Google Play metadata, including developer, certificate, and intrinsic app features [9], and Martin et al. further demonstrated that market metadata can support early-stage malware detection [8]. At the identifier-string level, SeqDroid shows that package names and certificate owner names can provide useful textual signals for obfuscated Android malware detection [10]. Research on camouflaged apps, fake apps, and app squatting further shows that externally visible identifiers can be manipulated. Kywe et al. examined camouflaged applications that imitate visible identity cues such as app names and icons [25]. Hu et al. showed that both app names and package names can be systematically manipulated in mobile app squatting [11]. Tang et al. found that fake apps frequently imitate official app names but rarely reuse official package names [12]. These studies establish the security relevance of metadata and identifier strings, but they typically combine several fields, incorporate APK-derived features, or compare an app with a known reference. Name2Pkg instead models the relationship between the app name and package name within the same app. It therefore requires neither an official reference app nor comparison against a set of suspected clones.
Name2Pkg: Name-conditioned Package Scoring Dataset App Name GRU
GRU
Package Name
GRU
Attention context
Package Name
GRU
GRU
GRU
avgNLL Anomaly Score
threshold
App Name
Not flagged
Flagged
Fig. 1. Overview of Name2Pkg. A character-level sequence-to-sequence model scores the observed package name conditioned on the app name. The lengthnormalized negative log-likelihood (avgNLL) serves as the anomaly score. Samples with scores at or above a benign-calibrated threshold are flagged for further analysis.
D. One-Class and Inconsistency-Based Android Malware Detection Benign-only, weak-label, and anomaly-based formulations have been studied to reduce dependence on complete and reliable malware labels. Wang et al. used a one-class support vector machine (SVM) trained on benign apps as part of a cloud-based hybrid detection framework [15]. DeLoach et al. applied positive–unlabeled learning to malware detection under weak ground truth [14]. Wang and Zheng evaluated oneclass feature-selection and classification methods for zero-day detection using benign samples [16], while MalGAE learns benign attributed function-call graphs using a stacked graph autoencoder [17]. A related line of work detects inconsistencies between an app’s stated purpose and its implementation. CHABADA groups apps according to description topics and identifies APIusage outliers using one-class SVMs [26]. BERTDetect revisits this formulation using BERTopic to model app descriptions and detect anomalous API usage [27]. These methods detect deviations between two information sources. However, they require full app descriptions and APIusage features. Name2Pkg restricts its input to two short identifier strings. It applies benign-only anomaly detection directly to within-app name-package correspondence. III. M ETHODOLOGY Name2Pkg treats weak name-package correspondence as an identifier-level anomaly signal. The method learns regularities from benign app-name and package-name pairs and then measures how unlikely a given package name is under benign correspondence patterns learned from benign data. Name2Pkg uses a character-level sequence-to-sequence (Seq2Seq) model trained only on benign pairs. Given a normalized app name, the model estimates the conditional likelihood of the corresponding normalized package name. The average negative log-likelihood of the package-name sequence is used as the anomaly score. Higher scores indicate weaker name-package correspondence under benign regularities. Figure 1 provides an overview of the screening process. A. Input Preprocessing The app name and package name are each normalized to reduce superficial formatting variation. For an app name a,
spaces are inserted at CamelCase boundaries, delimiters such as ., _, and - are replaced with spaces, consecutive spaces are collapsed, and the result is lowercased. For example, SampleTask_App is normalized to sample task app. For a package name p, underscores, hyphens, and spaces are replaced with periods, repeated periods are collapsed, and the result is lowercased. For example, com.example.sample_task is normalized to com. example.sample.task. The normalized app name â is used as the source sequence, and the normalized package name p̂ is used as the target sequence. B. Name-conditioned Package Scoring Name2Pkg models name-package correspondence using a character-level encoder-decoder architecture following prior sequence-to-sequence models [18], [19]. The encoder is a bidirectional gated recurrent unit (GRU) over app-name characters, and the decoder is a GRU that predicts package-name characters using additive attention [20]. Given â, the decoder predicts each package-name character conditioned on the previous package-name characters and the encoded app-name representation. Name2Pkg is trained by minimizing the average negative log-likelihood over benign training pairs. At inference time, the model scores the observed package-name sequence using its length-normalized negative log-likelihood (avgNLL): |p̂|
S(â, p̂) = −
1 X log Pθ (p̂t | p̂<t , â). |p̂| t=1
(1)
Length normalization prevents longer package names from being penalized solely because they contain more characters. Lower scores indicate plausible benign name-package correspondence, whereas higher scores indicate greater deviation from learned benign regularities. C. Benign-only Calibration and Decision Rule Name2Pkg follows a benign-only one-class calibration protocol. After training on benign training samples, the anomaly scores of a separate benign calibration set are used to select B the decision threshold. Let Scal denote the set of calibration scores computed from benign calibration pairs. For a target
TABLE I SUMMARY OF DATASET. THE MODEL IS TRAINED AND CALIBRATED EXCLUSIVELY ON BENIGN SAMPLES. Dataset
Samples
Benign (train) Benign (calibration) Benign (test) Malware (test)
50,184 6,273 6,273 4,399
Total
67,129
observed on the unseen benign test set may differ from the target value. We therefore report the achieved false-positive rate, denoted by aFPR. We also report ROC-AUC, precision, recall, and F1-score. ROC-AUC is used as a threshold-independent summary. Precision, recall, and F1-score are computed on the combined evaluation set consisting of the benign test set and the malware test set, whereas aFPR is computed only on the benign test set. C. Experimental Setup
false-positive rate γ, the threshold and decision rule are defined by B τγ = Q1−γ Scal , (2) ŷ = 1 {S(â, p̂) ≥ τγ } . Here, Q1−γ denotes the (1 − γ)-quantile, while ŷ = 1 denotes an anomalous sample and ŷ = 0 denotes a nonanomalous sample. The threshold is determined entirely from benign calibration scores; malware samples are not used for training, model selection, or threshold calibration. IV. E XPERIMENTS Under the benign-only one-class protocol, we evaluate Name2Pkg’s calibrated detection performance, comparisons against identifier-level baselines, computational efficiency, and the contribution of app-name conditioning. A. Dataset The evaluation uses an Android application dataset collected by a commercial bank in South Korea. The benign samples were collected between March 2023 and September 2023, whereas the malware samples were collected between October 2023 and June 2025. Each sample includes two textual identifiers used as input to Name2Pkg: the user-facing app name and the package name. Before filtering, the collected app names were written in a variety of scripts, including Latin, Hangul, CJK Unified Ideographs, Cyrillic, Katakana, and Hiragana. To ensure consistent character-level preprocessing across Name2Pkg and the identifier-level baselines, we retained apps whose names were primarily composed of ASCII Latin characters. Overall, preprocessing and filtering retained 62,730 of 115,388 benign samples and 4,399 of 12,992 malware samples, excluding 45.6% and 66.1%, respectively. The benign data were split into training, calibration, and test sets at an 80/10/10 ratio. The calibration split was used for threshold selection, and the benign test split and all malware samples were held out for final evaluation. Table I summarizes the dataset split used for training, calibration, and evaluation. B. Evaluation Metrics Name2Pkg is evaluated at benign-calibrated operating points with target false-positive rates γ ∈ {0.01, 0.05, 0.10}. For each method and each target FPR, the threshold is selected from benign calibration scores using Eq. (2). Because the threshold is selected on calibration data, the false-positive rate
All experiments follow the benign-only protocol. All trainable models are trained only on the benign training split, thresholds are selected only from benign calibration scores, and malware samples are used only for final evaluation. Unless otherwise stated, the 80/10/10 benign split was generated using random seed 42. 1) Shared Preprocessing and Vocabulary: All methods use the normalization procedure described in Section III. App names and package names are truncated or padded to a maximum length of 64 characters. The character vocabulary is built only from benign training app names and package names to avoid leakage from calibration, benign test, or malware test samples. Name2Pkg uses <PAD>, <UNK>, <BOS>, and <EOS> as special tokens. Package-only additionally uses <NULL_APP> as a fixed source token. Conditional GRU-LM uses <APP>, </APP>, <PKG>, and </PKG> to mark the app-name and package-name regions in the concatenated character sequence. 2) Name2Pkg: Name2Pkg is implemented as a characterlevel encoder-decoder Seq2Seq model. The encoder is a onelayer bidirectional GRU over app-name characters, and the decoder is a one-layer GRU over package-name characters with additive attention over encoder hidden states. The character embedding dimension is 96, the encoder hidden dimension is 192 per direction, and the decoder hidden dimension is 192. Dropout is set to 0.1. The model is trained with Adam using a learning rate of 10−3 , batch size 256, and 8 epochs. Early stopping is not used. During training, the decoder uses teacher forcing with the shifted package sequence. The anomaly score is computed using Eq. (1). 3) Identifier-Level Baselines: SBERT Top-1 is an embedding-based baseline using Sentence-BERT [28], designed to measure semantic similarity between an app name and the most similar package-name segment. We use the pretrained all-MiniLM-L6-v2 checkpoint [29], based on the compact MiniLM architecture [30], without task-specific fine-tuning. The normalized package name is split into dot-delimited segments, and common namespace tokens such as com, org, net, kr, jp, io, co, de, and air are removed. The anomaly score is one minus the maximum cosine similarity between the app-name embedding and the embeddings of the remaining package segments. Thresholds are selected from benign calibration scores using the same protocol as Name2Pkg.
TABLE II P ERFORMANCE COMPARISON FOR EACH TARGET FPR. A FPR DENOTES THE ACHIEVED FPR ON THE BENIGN TEST SET. Target FPR = 0.01 Method SBERT Top-1 Package-only Conditional GRU-LM Name2Pkg (ours)
ROC-AUC aFPR Prec. Recall 0.903 0.930 0.948 0.982
0.011 0.010 0.008 0.010
0.820 0.977 0.981 0.979
0.070 0.607 0.622 0.656
Target FPR = 0.05
Target FPR = 0.10
F1
aFPR Prec. Recall
F1
aFPR Prec. Recall
F1
0.128 0.749 0.761 0.786
0.055 0.051 0.049 0.044
0.519 0.784 0.809 0.909
0.103 0.098 0.098 0.097
0.690 0.804 0.840 0.916
Package-only is a Seq2Seq control with a fixed null appname input. It uses the same Seq2Seq architecture and hyperparameters as Name2Pkg, but replaces the app-name input with the fixed token <NULL_APP> during both training and inference. Its anomaly score is the same average negative loglikelihood used in Eq. (1), with the fixed null source input. This control tests whether package-name morphology alone can explain the observed detection performance. Conditional GRU-LM is a conditional GRU language model used as a simpler recurrent baseline. Each input is represented as a single character sequence: app-name characters followed by package-name characters, with explicit delimiter tokens marking the regions. The model is a one-layer character-level GRU language model with embedding dimension 96, hidden dimension 192, and dropout 0.1. It is trained with Adam using a learning rate of 10−3 , batch size 256, and 8 epochs. The loss and anomaly score are computed only over package-name character positions after the <PKG> delimiter. This baseline tests whether a single recurrent conditional language model suffices without the explicit encoder–decoder structure used by Name2Pkg. 4) Efficiency Measurement: We measure single-sample CPU inference latency with a batch size of 1 on an Intel Core i7-9700K CPU using 8 threads. Latency is averaged over five runs on the same 500 apps, with 32 warm-up samples processed before each run. Timing includes input encoding and scoring, but excludes shared string normalization, loading, and disk I/O. Both Seq2Seq models compute encoder attention projections once per input and reuse them across decoding steps. D. Main Detection Results Table II compares Name2Pkg with identifier-level baselines and the package-only control under benign-calibrated operating points. Name2Pkg achieves the strongest overall performance, with an ROC-AUC of 0.982. At a target FPR of 0.05, it achieves an aFPR of 0.044 on the benign test set, with malware recall of 0.885. At this operating point, the calibrated threshold is 2.408, yielding 3,893 true positives, 506 false negatives, 273 false positives, and 6,000 true negatives. Precision and F1-score on this evaluation set are 0.934 and 0.909, respectively. The comparison with Package-only shows that packagename morphology is informative but insufficient. The packageonly control reaches an ROC-AUC of 0.930 and recall of 0.692 at the target FPR of 0.05, indicating that anomalous
0.828 0.906 0.912 0.934
0.378 0.692 0.727 0.885
0.805 0.846 0.856 0.875
0.605 0.766 0.826 0.962
package names alone provide a meaningful signal. However, Name2Pkg improves recall from 0.692 to 0.885 at the same target FPR. This gap supports our central hypothesis that conditioning on the app name adds discriminative information beyond standalone package plausibility. The remaining baselines clarify how this correspondence should be modeled. Conditional GRU-LM also uses appname context and outperforms the package-only control, but it remains below Name2Pkg across the calibrated operating points. This suggests that an attention-based encoder–decoder formulation is better suited for modeling name-conditioned package regularities than a single recurrent language model over a concatenated sequence. In contrast, SBERT Top-1 performs substantially worse, especially at low false-positive budgets. This indicates that off-the-shelf semantic similarity between an app name and package segments does not adequately capture character-level identifier correspondence, where package names often contain abbreviations, namespaces, developer tokens, or short fragments. Overall, the results show a consistent pattern: packagename morphology provides a useful baseline signal, app-name conditioning strengthens that signal, and the attention-based encoder-decoder formulation gives the strongest calibrated identifier-level screening performance. E. Efficiency Analysis Table III reports the computational profiles of the identifierlevel methods under single-sample CPU inference. The benchmark includes input encoding and anomaly-score computation, but excludes shared string normalization, loading, and disk I/O. Name2Pkg requires a 3.57 MiB checkpoint and 28.20 ms per sample on CPU. Its footprint is nearly identical to the package-only control because both use the same Seq2Seq architecture; the main difference is whether the source input is the actual app name or a fixed null token. The small latency gap between the two models indicates that app-name conditioning adds limited inference overhead relative to the package-only control. Conditional GRU-LM is substantially smaller and faster, requiring 0.74 MiB and 6.57 ms per sample. However, this efficiency comes with lower calibrated detection performance: at the target FPR of 0.05, its recall is 0.727 compared with 0.885 for Name2Pkg. Name2Pkg has higher single-sample latency than SBERT Top-1 despite having fewer parameters. Sequential decoding of package-name characters contributes
TABLE III M ODEL SIZE AND SINGLE - SAMPLE CPU INFERENCE LATENCY. L ATENCY IS AVERAGED OVER FIVE RUNS . L ATENCY INCLUDES INPUT ENCODING AND SCORING , BUT EXCLUDES SHARED STRING NORMALIZATION , LOADING , AND DISK I/O. Method SBERT Top-1 Package-only Conditional GRU-LM Name2Pkg (ours)
Parameters 22,713,216 935,445 191,894 934,676
Checkpoint size (MiB)
CPU latency (ms/sample)
86.66 3.58 0.74 3.57
9.63 27.95 6.57 28.20
to this latency. SBERT Top-1 has a much larger pretrained model footprint and substantially weaker recall under the same benign-calibrated protocol. These results indicate that the key trade-off is not latency alone, but whether the model captures the intended name-package correspondence signal. Name2Pkg is not the smallest or fastest method, but it provides the best calibrated detection performance while remaining compact enough for pre-analysis triage. Its compact checkpoint and measured CPU latency support its intended role as a pre-analysis signal before more expensive static, dynamic, or hybrid analysis. F. Ablation and Control Analysis We further examine whether Name2Pkg exploits pairwise name-package correspondence rather than relying only on package-name morphology. Because the package-only control is already included in the main comparison, this section focuses on score-distribution analysis and a name–package shuffle control. 1) Distributional Effect of App-name Conditioning: As described above, the package-only control isolates the effect of app-name conditioning by replacing the app-name input with a fixed null token. Given the performance gap between Name2Pkg and the package-only control, we examine the score distributions to identify where app-name conditioning improves separation. Figure 2 compares the anomaly-score distributions of the package-only control and Name2Pkg. Malware scores are bimodal in both models. In the package-only control, the highscore malware mode is already largely separated from benign samples, which is consistent with package-name morphology providing a useful anomaly signal. However, the low-score malware mode remains close to the benign distribution. In the name-conditioned model, this low-score mode is more clearly separated from benign samples. The distributional comparison suggests that app-name conditioning adds a useful signal where package-name morphology alone provides weaker separation. Together with the package-only comparison, this supports the contribution of name-conditioned package plausibility. We interpret this as distributional evidence, not as evidence of fixed semantic categories among malware samples. 2) Name-Package Shuffle Control: We perform a namepackage shuffle control on the benign test set. The Name2Pkg
TABLE IV NAME - PACKAGE SHUFFLE CONTROL ON THE BENIGN TEST SET. A PP NAMES ARE PERMUTED WHILE PACKAGE NAMES ARE FIXED ; THE MODEL IS NOT RETRAINED OR RECALIBRATED . P95 AND P99 DENOTE THE 95 TH AND 99 TH PERCENTILES , RESPECTIVELY. Set
Mean
Std.
Median
P95
P99
Original benign Shuffled benign Malware
1.326 2.077 3.632
0.593 0.473 0.978
1.254 2.077 3.785
2.360 2.851 5.049
2.909 3.336 5.447
model is trained on the original benign training pairs, and the thresholds are selected from the original benign calibration set. We then randomly permute app names within the benign test set while keeping package names fixed. The model is not retrained or recalibrated. This control preserves the marginal distributions of benign app names and package names, but breaks their valid pairwise correspondence. If Name2Pkg modeled only packagename morphology, this shuffle would have little effect on the anomaly-score distribution. In contrast, a score increase on shuffled pairs suggests that the model is sensitive to pairwise name-package correspondence. Table IV summarizes the shuffle control. The mean anomaly score increases from 1.326 for the original benign test pairs to 2.077 for the name-shuffled pairs. The median score also increases from 1.254 to 2.077. This score increase suggests that breaking valid name–package correspondence makes otherwise benign package names less plausible to the nameconditioned model. At the same time, the mean and median scores for the nameshuffled benign pairs remain below the corresponding malware values of 3.632 and 3.785, respectively. This is expected because the shuffled samples still contain benign package names, whereas malware samples may exhibit both unusual packagename morphology and weak name-package correspondence. Therefore, the shuffle control should be interpreted as a test of sensitivity to pairwise correspondence, not as a transformation that makes benign samples equivalent to malware. Overall, the ablation and control results support a layered interpretation of the proposed signal. Package names alone provide a meaningful baseline, but conditioning on the app name improves detection by evaluating whether a package name is plausible given the app name. The shuffle control further suggests that Name2Pkg is sensitive to the relationship between the two identifiers rather than merely learning the standalone plausibility of package strings. G. Discussion and Limitations The results suggest that Name2Pkg captures two complementary identifier-level signals: standalone package-name regularity and name-conditioned package plausibility. The package-only control confirms that package-name morphology alone carries useful information, while Name2Pkg’s improvement over this control shows that app-name conditioning adds further discriminative value. The comparison with Conditional
(a) Package-only
(b) Name2Pkg (ours)
1.0
Benign (Test) Malware
Density
0.8
0.6
0.4
0.2
0.0
1
2
3
4
5
Anomaly Score (avgNLL)
1
2
3
4
5
Anomaly Score (avgNLL)
Fig. 2. Anomaly-score distributions for the package-only control and Name2Pkg. Name2Pkg (right) reduces the overlap between benign and malware score distributions.
GRU-LM also indicates that the modeling formulation matters, suggesting that the attention-based encoder-decoder architecture better captures the proposed correspondence signal. Finally, the score-distribution and shuffle controls support the interpretation that Name2Pkg is sensitive to pairwise namepackage correspondence rather than only standalone package plausibility. In deployment, Name2Pkg should be used as a triage filter: flagged apps can be prioritized for deeper inspection, while unflagged apps receive lower priority. Its decision relies only on identifier-level evidence, not on permissions, API calls, bytecode, network behavior, native code, or runtime traces. This limited input surface is central to its intended use: Name2Pkg provides a low-cost screening signal before more expensive static, dynamic, or hybrid analysis is applied. The evaluation set contains 41.2% malware. Since precision and F1-score depend on malware prevalence, these metrics may differ in deployment environments with different class proportions. As with other lightweight screening signals, Name2Pkg has a clear limitation under fully adaptive evasion. An attacker aware of the detection principle may choose an app name and package name that are mutually plausible under benign naming conventions. For example, a malicious app may use a banking-related app name together with a banking-related package name. Such samples may receive low anomaly scores despite being malicious. This does not invalidate the intended use of Name2Pkg: its value lies in detecting identifier-level irregularities at low cost and prioritizing suspicious apps for deeper inspection, rather than serving as a standalone defense against fully adaptive attackers. False positives may also occur for legitimate apps with weakly aligned or opaque identifiers. Some benign apps may use package names that reflect internal project names, legacy namespaces, vendor identifiers, abbreviations, or opaque organizational structures rather than the user-facing app name. Such identifiers may be legitimate engineering or organiza-
tional artifacts, but they can still appear anomalous to a nameconditioned model. In deployment, the decision threshold should therefore be selected according to the acceptable falsepositive budget, as this setting controls how many legitimate but weakly aligned apps are escalated for further analysis. Because benign and malware samples were collected during different periods, differences in collection periods may contribute to the observed separation between the two classes. This study also restricts the input to app names written in Latin script. This restriction provides a controlled setting for character-level preprocessing and comparison with the English-centric SBERT baseline, while leaving the multilingual Android ecosystem for future work. App names written in non-Latin scripts, such as Hangul and CJK Unified Ideographs, may exhibit different naming patterns and different relationships with their package names. Even within Latinscript app names, language-specific naming conventions may affect name-package correspondence. Extending Name2Pkg to support multilingual app names, Unicode-aware preprocessing, and language-specific naming conventions is an important direction for future work. V. C ONCLUSION This paper presented Name2Pkg, a lightweight anomaly detection method for Android malware screening that uses only the app name and package name. Name2Pkg employs a character-level sequence-to-sequence model and computes length-normalized negative log-likelihood scores. The experimental results show that name-package correspondence is an effective identifier-level screening signal. Name2Pkg achieves an ROC-AUC of 0.982 and, at a target FPR of 0.05, an aFPR of 0.044 on the benign test set with malware recall of 0.885. Compared with the package-only control, Name2Pkg improves the ROC-AUC from 0.930 to 0.982 and recall from 0.692 to 0.885 at the same target FPR, showing that app-name conditioning adds discriminative information beyond package-name morphology.
The ablation and control analyses further support this conclusion. The distributional comparison shows that app-name conditioning improves separation between benign apps and low-score malware samples, while the name-package shuffle control shows that breaking valid benign pairings increases anomaly scores. These findings suggest that Name2Pkg exploits pairwise name-package correspondence rather than relying only on standalone package-name plausibility. Name2Pkg also maintains a compact computational profile. It contains 934,676 parameters, has a checkpoint size of 3.57 MiB, and takes 28.20 ms per sample for CPU inference. Although it is not the smallest or fastest method among the compared approaches, it provides the strongest calibrated detection performance among the evaluated identifier-level methods while remaining feasible for pre-analysis screening. Overall, Name2Pkg demonstrates that lightweight identifierlevel modeling can provide a useful early signal for Android malware triage. Rather than replacing full static, dynamic, or hybrid malware analysis, Name2Pkg is designed to operate as a low-cost pre-analysis screening filter before such methods are applied. Apps whose name-package pairings are unlikely under benign naming regularities can then be prioritized for deeper inspection. R EFERENCES [1] D. Arp, M. Spreitzenbarth, M. Hübner, H. Gascon, and K. Rieck, “DREBIN: Effective and explainable detection of Android malware in your pocket,” in Proceedings of the Network and Distributed System Security Symposium (NDSS). Internet Society, 2014. [2] K. Tam, S. J. Khan, A. Fattori, and L. Cavallaro, “CopperDroid: Automatic reconstruction of Android malware behaviors,” in Proceedings of the Network and Distributed System Security Symposium (NDSS). Internet Society, 2015, pp. 1–15. [3] S. Zhang, H. Su, H. Liu, and W. Yang, “MPDroid: A multimodal pretraining Android malware detection method with static and dynamic features,” Computers & Security, vol. 150, p. 104262, 2025. [4] J. Garcia, M. Hammad, and S. Malek, “Lightweight, obfuscationresilient detection and family identification of Android malware,” ACM Transactions on Software Engineering and Methodology, vol. 26, no. 3, pp. 1–29, 2018. [5] A. Kadir and S. K. Peddoju, “PacDroid: Lightweight Android malware detection using permissions and intent features,” Multimedia Tools and Applications, vol. 84, no. 27, pp. 32 351–32 379, 2025. [6] R. Ma, S. Yin, X. Feng, H. Zhu, and V. S. Sheng, “A lightweight deep learning-based Android malware detection framework,” Expert Systems with Applications, vol. 255, p. 124633, 2024. [7] P. Teufl, M. Ferk, A. Fitzek, D. Hein, S. Kraxberger, and C. Orthacker, “Malware detection by applying knowledge discovery processes to application metadata on the Android Market (Google Play),” Security and Communication Networks, vol. 9, no. 5, pp. 389–419, 2016. [8] I. Martı́n, J. A. Hernández, A. Muñoz, and A. Guzmán, “Android malware characterization using metadata and machine learning techniques,” Security and Communication Networks, vol. 2018, p. 5749481, 2018. [9] A. Muñoz, I. Martı́n, A. Guzmán, and J. A. Hernández, “Android malware detection from Google Play meta-data: Selection of important features,” in 2015 IEEE Conference on Communications and Network Security (CNS). IEEE, 2015, pp. 701–702. [10] W. Y. Lee, J. Saxe, and R. Harang, “SeqDroid: Obfuscated Android malware detection using stacked convolutional and recurrent neural networks,” in Deep Learning Applications for Cyber Security, M. Alazab and M. Tang, Eds. Cham: Springer, 2019, pp. 197–210. [11] Y. Hu, H. Wang, R. He, L. Li, G. Tyson, I. Castro, Y. Guo, L. Wu, and G. Xu, “Mobile app squatting,” in Proceedings of The Web Conference 2020. ACM, 2020, pp. 1727–1738.
[12] C. Tang, S. Chen, L. Fan, L. Xu, Y. Liu, Z. Tang, and L. Dou, “A largescale empirical study on industrial fake apps,” in 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2019, pp. 183–192. [13] Google, “Configure the app module,” https://developer.android.com/ build/configure-app-module, 2026, last updated: 2026-02-26; accessed: 2026-04-20. [14] J. DeLoach, D. Caragea, and X. Ou, “Android malware detection with weak ground truth data,” in 2016 IEEE International Conference on Big Data (Big Data). IEEE, 2016, pp. 3457–3464. [15] X. Wang, Y. Yang, and Y. Zeng, “Accurate mobile malware detection and classification in the cloud,” SpringerPlus, vol. 4, p. 583, 2015. [16] Y. Wang and J. Zheng, “An evaluation of one-class feature selection and classification for zero-day Android malware detection,” in 17th International Conference on Information Technology–New Generations (ITNG 2020), ser. Advances in Intelligent Systems and Computing. Springer, 2020, pp. 105–111. [17] F. Deldar, M. Abadi, and M. Ebrahimifard, “Android malware detection using one-class graph neural networks,” The ISC International Journal of Information Security, vol. 14, no. 3, pp. 51–60, 2022. [18] K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder–decoder for statistical machine translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Doha, Qatar: Association for Computational Linguistics, 2014, pp. 1724–1734. [19] I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in Neural Information Processing Systems, vol. 27. Curran Associates, Inc., 2014, pp. 3104–3112. [20] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in 3rd International Conference on Learning Representations (ICLR 2015), 2015. [21] J. Gu, H. Zhu, Z. Han, X. Li, and J. Zhao, “GSEDroid: GNNbased Android malware detection framework using lightweight semantic embedding,” Computers & Security, vol. 140, p. 103807, 2024. [22] V. Rastogi, Y. Chen, and W. Enck, “AppsPlayground: Automatic security analysis of smartphone applications,” in Proceedings of the Third ACM Conference on Data and Application Security and Privacy. ACM, 2013, pp. 209–220. [23] S. I. Imtiaz, S. ur Rehman, A. R. Javed, Z. Jalil, X. Liu, and W. S. Alnumay, “DeepAMD: Detection and identification of Android malware using high-efficient deep artificial neural network,” Future Generation Computer Systems, vol. 115, pp. 844–856, 2021. [24] M. Krzysztoń, B. Bok, M. Lew, and A. Sikora, “Lightweight on-device detection of Android malware based on the Koodous platform and machine learning,” Sensors, vol. 22, no. 17, p. 6562, 2022. [25] S. M. Kywe, Y. Li, R. H. Deng, and J. I. Hong, “Detecting camouflaged applications on mobile application markets,” in Information Security and Cryptology – ICISC 2014, ser. Lecture Notes in Computer Science. Cham: Springer, 2015, vol. 8949, pp. 241–254. [26] A. Gorla, I. Tavecchia, F. Gross, and A. Zeller, “Checking app behavior against app descriptions,” in Proceedings of the 36th International Conference on Software Engineering. ACM, 2014, pp. 1025–1035. [27] N. Ranaweera, J. Xu, S. Seneviratne, and A. Seneviratne, “BERTDetect: A neural topic modelling approach for Android malware detection,” in Companion Proceedings of the ACM on Web Conference 2025. ACM, 2025, pp. 1802–1810. [28] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using siamese BERT-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Hong Kong, China: Association for Computational Linguistics, Nov. 2019, pp. 3982–3992. [29] Sentence-Transformers, “all-MiniLM-L6-v2,” https://huggingface.co/ sentence-transformers/all-MiniLM-L6-v2, 2025, model card; accessed: 2026-04-20. [30] W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, “MiniLM: Deep self-attention distillation for task-agnostic compression of pretrained transformers,” in Advances in Neural Information Processing Systems, vol. 33. Curran Associates, Inc., 2020, pp. 5776–5788.