Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection Syed Ali Ahmed1 , Malaika Raza2 , Muhammad Shoaib Siddiqui3 (Senior Member, IEEE), and Muhammad Rafi4 (Member, IEEE) 1
National University of Computer and Emerging Sciences, Karachi, 75020, Pakistan National University of Computer and Emerging Sciences, Karachi, 75020, Pakistan Faculty of Computer and Information Systems, Islamic University of Madinah, Madinah 42351, Saudi Arabia 4 Department of AI & DS, National University of Computer and Emerging Sciences, Karachi, 75020, Pakistan 2
3
arXiv:2609.21599v1 [cs.AI] 18 Sep 2026
This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.
CORRESPONDING AUTHORS: SYED ALI AHMED (email: [email protected]) and MUHAMMAD SHOAIB SIDDIQUI (email: [email protected]). Muhammad Shoaib Siddiqui’s ORCID is 0000-0002-5656-0416. The authors extend their appreciation to the Deanship of Scientific Research, Islamic University of Madinah, Saudi Arabia, for funding this research work.
ABSTRACT Fraudulent job posting detection aims to identify job advertisements that are corrupted either
through fake content, misleading information, or negative intent, disrupting the online eco-system of jobseekers and employers. Existing studies in this domain lack effective methods to simultaneously achieve high accuracy and meaningful structure of latent-space representations that capture subtleties among fake posts. To this end, we propose Centroid-Guided Contrastive Loss (CGCL), a loss function which unifies classification with densely formulated clustering to consistently reshape latent-space through a centroid-driven top-k push-and-pull mechanism. The complementary nature of CGCL enables the model to enforce accurate decision boundaries and maintain high clustering compactness, effectively capturing both class separability and latent structure. Extensive experiments demonstrate the state-of-the-art (SOTA) performance of our method on EMSCAD, a public benchmark dataset. The code associated with this work is available at: https://github.com/ali-ahmed925/CGCL code/tree/main INDEX TERMS Contrastive learning, Centroid-guided loss, Representation learning, Word embeddings,
GloVe, Word2Vec, TF-IDF, Text classification, Clustering metrics, Fraudulent postings, Latent space structuring.
I. INTRODUCTION
I
N the current age of Intelligence, where internet has deeply transformed our modern day life and social media platforms are straightforwardly accessible, companies nowadays, increasingly rely on electronic means to advertise their job postings and recruitment windows. This digitized approach has not only simplified the application process for job seekers but also accelerated recruitment operations for employers. However, malicious attempts to corrupt the recruitment ecosystem have emerged due to the prevalence of fake job postings surfacing on popular job-hunting platforms. These fraudulent postings are a direct attack on applicant’s personal information, exposing them to a range of cyber threats, such as identity theft, financial scams, and privacy breaches [9]. According to a report published by the Better
Business Bureau, employment scams ranked as the second most risky scam type in 2023, with a 5.2% increase in reported incidents and an average reported loss of $1,995 per victim, up from $1,500 in 2022 [12]. Therefore, detecting such fraudulent postings is of utmost importance to ensure compliance and integrity, and also to safeguard job-seekers from financial and identity-related harms. Machine learning in recent years has emerged as a highly promising technique across a wide range of domains, from healthcare [13]–[15] and finance [16]–[19] to video surveillance [20]–[22] and cybersecurity [23]–[26]. A very impactful application of machine learning is in the field of Natural Language Processing (NLP), which is well-suited for tasks involving unstructured textual data. This makes it a viable approach for problems like fake job post detection where
TABLE 1. Comparison of fake job postings detection methods showing feature extraction techniques, imbalance handling approaches, and performance metrics.
Ref
Year
Feature Encoding
Imbalance Handling
Best Performing Model
Accuracy
[1] [2] [3] [4]
2020 2021 2021 2021
Categorical encoding TF-IDF -
✗ ✗ ✗ ✗
Random Forest Classifier BiLSTM MLP Classifier Decision Forests
98.27 98.0 71.0 95.4
[5] [6] [7] [8] [9] [10]
2022 2022 2023 2023 2023 2024
TF-IDF BiLSTM One-hot encoding Word2Vec TF-IDF
✓ ✗ ✗ ✓ ✓ ✓
Extra Tree Classifier BiLSTM DNN BiLSTM LSTM SGD Classifier
99.9 97.21 98.0 98.71 97.18 98.6
[11]
2024
TF-IDF
✓
Extra Tree Classifier
99.76
data is available in natural language format. Several machine learning models have been employed to detect fraudulent samples ranging from traditional models such as Naive Bayes Classifier (NBC), K-Neighbors Classifier (KNNs), Decision Tree Classifier (DTC) and more to advanced deep learning architectures like Long Short Term Memory (LSTMs), and Gated Recurrent Units (GRUs). These models capture temporal context and patterns from sequences of textual input, enabling them to classify anomalous samples as ”fraudulent postings”. While choosing the right model is undoubtedly a critical aspect of solving the problem, it is only one part of the broader learning scheme. An often unnoticed and equally essential component of any learning framework is the loss or cost function. A loss function directly influences the model to learn subtle patterns in the data by computing the difference between original target labels and predicted outputs, thereby optimizing the model’s parameters during training. Most existing studies rely on standard loss functions like Cross-entropy Loss or Hinge Loss, which while effective in many cases, may not be able to fully capture nuanced class boundaries and inter-class separation in such complex tasks where normal and fraudulent samples’ features share subtle similarities. To this end, we propose a novel Centroid-Guided Contrastive Loss (CGCL), which meaningfully reshapes the high-dimensional latent embedding space by incorporating a centroid-based push-and-pull mechanism that enhances intraclass compactness and inter-class separability. We equip CGCL with a top k strategy to select only the k farthest in-class samples during the pull phase and the closest crossclass samples during the push phase, ensuring a focused and effective feature refinement. Additionally, CGCL leverages a class-balanced Cross-Entropy Loss that guides the classifier towards more robust decision boundaries, maximizing the discrimination between unique classes for improved generalization and classification performance. We train a Multilayer Perceptron with six hidden layers, each followed by a ReLU
activation function— except for the final hidden layer which omits the activation. We conducted extensive experiments on a publicly available dataset from Kaggle [27] to validate the effectiveness of our approach. Our contributions can be summarized as follows: 1) We propose a novel Centroid-Guided Contrastive Loss (CGCL) that ensures intra-class compactness and interclass separability through top k push-pull mechanism. 2) We integrate CGCL with class-balanced Cross-Entropy Loss, resulting in more robust and discriminative decision boundaries. 3) We evaluate our complete framework on a benchmark dataset, demonstrating improved performance in classimbalance scenarios. II. RELATED WORK
In this section, we review some of the novel works previously done by seasoned researchers that have contributed to advancing fake job post detection. These studies will serve as the context for our current approach and highlight key trends, challenges, and gaps that our work aims to address. Researchers have extensively explored classical machine learning approaches for identifying fraudulent job listings. Anita et al. in [2] employ several machine learning models to detect fake job posts. The data cleaning process is specifically emphasized as a critical step in their pipeline. Among the various models evaluated, Bidirectional LSTMs achieved the best results, reporting an accuracy of 98%. Similarly, Anbarasu et al. [10] trained multiple machine learning algorithms, including Naive Bayes Classifier (NBC), Decision Tree Classifier (DTC), Multilayer Perceptron (MLP), and Stochastic Gradient Descent (SGD), with SGD yielding the best results, achieving an impressive 98.6% overall accuracy. Another similar work evaluates a broad spectrum of data mining and classification algorithms for predicting fake job posts. Singh et al. in [7] developed a fraud detection model using existing ML models on the EMSCAD dataset, achieving 97.4% accuracy in identifying fake job postings. Their
TABLE 2. Job Dataset Schema: Column Specifications and Data Types
Column Name
Data Type
Non-Null Count
job id title
int64 object
17880 17880
Unique identifier for job Job position or role name
location department salary range company profile description requirements
object object object object object object
17534 6333 2868 14572 17871 15146
City, state, or country posted Hiring department within company Offered compensation range Company background or description Job duties and responsibilities Skills or qualifications needed
benefits telecommuting has company logo has questions employment type required experience
object int64 int64 int64 object object
10637 17880 17880 17880 14409 10830
Perks or advantages offered Remote work allowed (binary) Company logo presence (binary) Screening questions included (binary) Full-time, part-time, contract, etc. Level of experience required
required education industry function fraudulent
object object object int64
9775 12977 11425 17880
Minimum education qualification needed Sector of employment Primary role or specialization Legitimate or fake posting label
approach involved data preprocessing, feature selection, and ensemble classification. TF-IDF is a widely used feature-extraction technique that transforms text into weighted vectors based on token importance. Keerthana et al. [3] applied a TF-IDF vectorizer during preprocessing to encode tabular data, which was then fed into various ML models, with the MLP classifier achieving the highest accuracy of 71%. Similarly, Dutta et al. [1], after applying necessary data preprocessing steps, fed the extracted features to several ML models, with the Random Forest Classifier outperforming others in terms of accuracy. In a novel work that explores ensemble-based learning for fake job detection, Shibly et al. [4] leverage boosted decision trees and two-class decision forest algorithms to detect fake job postings. In the first algorithm, decision trees are arranged in an ensemble manner, where the errors of previous trees are corrected by subsequent ones before arriving at the final prediction. Two class decision forests, on the other hand, rely on aggregated outcomes from grouped decision trees. Beyond classical methods, several studies have shifted towards deep learning architectures to better capture semantic patterns in job descriptions. In a detailed approach presented by Pillai et al. in [8], they propose a training framework built on several BiLSTM layers. First the numeric and textual data is converted into fixed-size numerical vectors using separate tokenization layers, followed by an embedding layer, multiple BiLSTM layers, and a merging mechanism that fuses textual and numerical features for downstream classification. Rathudi et al. in [9] leverage bidirectional LSTMs with Word2Vec [28], a vectorizer that converts textual tokens into low dimensional vector representations based on their
Description
semantic context. Coupling these two approaches allows the framework to capture both the semantic meaning of individual words and the sequential dependencies within the text, resulting in an overall accuracy of 97.1%. Some studies have also emphasized the importance of data imbalance mitigation and feature selection. Amaar et al. in [5] employ both TF-IDF vectorizer and BoW (Bag of Words) techniques for feature extraction. The extracted features are then fed into six machine learning models to evaluate the best performing combination. Additionally, they achieve over 99% accuracy by incorporating oversampling in their framework. In a distinct study, Afzal et al. [11] argue the reason that existing works often fell short is because feature selection and class imbalance are mostly overlooked. They leverage Chi-Square and PCA techniques to select the most relevant features, and apply SMOTE for minority class oversampling to effectively address the issue of class imbalance. While the above studies rely primarily on traditional feature engineering methods, a recent work explored deep contextual representations, Qayyum et al. [6] propose a novel feature extraction technique called Deep Contextualized Word Representation (DCWR), that employs a two-layered BiLSTM to generate context-aware encoded representation of words by modeling the likelihood of word sequences in both forward and backward directions. Furthermore, they apply PCA as a feature reduction technique to select the principal components that capture maximum variance in the data.
FIGURE 1. Analysis of job postings: (a) Distribution of fraudulent vs. genuine postings, (b) Geographic distribution of postings, (c) Educational requirement analysis, and (d) Distribution by employment type
III. METHODOLOGY A. DATA COLLECTION
To carry out experiments on our proposed approach, we employed The Employment Scam Aegean Dataset (EMSCAD) [27], a publicly available benchmark dataset for detection of fraudulent job postings. The dataset is highly imbalanced, comprising 17,014 genuine job advertisements and 866 fraudulent ones, collected between 2012 and 2014. The tabular dataset contains both numerical and categorical columns, providing sufficient information to train complex neural networks capable of modeling intricate relationships between features. Table 2 presents description of each feature along with their datatypes. B. DATA ANALYSIS
Data Analysis is a crucial step in building a machine learning model as it offers comprehensive insights related to our underlying dataset and aids in determining the selection of an appropriate modeling workflow. We first examined the distribution of target label (Figure 1) (a), which reveals a
severe class imbalance, with legitimate postings significantly exceeding fraudulent ones. The handling of such a problem is of utmost importance as it can bias the model towards underfitting and can lead to poor generalization over unseen examples of fraudulent postings. The solution to this problem shall later be discussed in the paper. Additionally, analyzing geographical aspect of the dataset disclosed that majority of the advertisements came from cities like London, New York, and Athens that are considered the employment hubs of their respective countries as shown in Figure 1 (b). This shows us that the distribution of postings is not random, but is sophistically linked to the regions that are employment-concentrated. We also investigated the educational qualifications required by the postings (Figure 1) (c), which ranged from school-level or equivalent to bachelor’s and master’s degrees, with the majority of jobs demanding at least an undergraduate degree, reflecting the diversity in posted advertisements across employment sectors. Furthermore, we inspect what kind of employment types are more prone to get advertised in fraudulent postings,
which revealed that Full-time positions exhibit a higher share of fraudulent advertisements compared to other positions as shown in Figure 1 (d). Finally, the data analysis provides a clear picture of all the challenges that need to be resolved before proceeding with effective preprocessing and model development. C. DATA PREPROCESSING 1) DATA CLEANING
We first cleaned our data through a preprocessing pipeline. Data cleaning involves addressing inaccuracies, missing entries, duplicates, and inconsistently formatted data. We began by handling null values in the data which posed a significant challenge as there were a total of 70183 void entries in the dataset. However, this number is aggregated across all columns and therefore can be misleading. To obtain a more accurate assessment, we only selected the columns that contained some amount of null values, and computed the average number of nulls per affected column which resulted in approximately 5848 null entries per column which still is a very concerning number. To handle this, we removed the integer columns from our dataset, retaining only the categorical features that were sufficient for modeling complex patterns. Following this, all the null entries were then replaced with empty strings. These empty strings individually seem to appear meaningless and insignificant, however, when concatenated with other categorical columns, we obtain a combined column named ”text”. This methodology preserves the structure of the data, as it eliminates the need to remove rows with missing values, thereby maintaining the overall size and integrity of the dataset. Furthermore, we removed all the duplicated rows from our dataset to maintain consistency, and performed additional pre-processing on the remaining textual data. In order to do this, we designed a pipeline that carried out several cleaning tasks including converting text to lowercase, removing emails, URLs, and HTML tags, eliminating punctuation and numbers, removing stopwords, and finally lemmatizing the tokens using spaCy, ensuring that our data is suitable for feature engineering and machine learning modeling.
here and instead present its comparative performance in Section IV. Algorithm 1: Centroid-Guided Contrastive Loss (CGCL) Input: Embeddings E, Logits Z, Labels Y , hyperparameters: k, α, β, margin m Output: Loss L Step 1: Classification Loss (Weighted CE) Compute class counts: nc = count(Y = c) for each class c Compute weights: wc = nc1+ϵ PN LCE ← − N1 i=1 wyi log p(yi | xi ) Step 2: Normalize Embeddings E E ← ∥E∥ 2 Step 3: Contrastive Loss (Centroid-Guided) Initialize Lcontrastive ← 0, C ← 0 foreach class c ∈ unique(Y ) do Ec ← {ei ∈ E | yi = c} E¬c ← {ei ∈ E | yi ̸= c} if |Ec | < k + 1 or |E¬c | < k then continue P Compute centroid: µc = |E1c | e∈Ec e // Pull: hardest positives d+ = {∥e − µc ∥2 | e ∈ Ec } + Select top-k Pfarthest: Hk 2 1 Lpull = k e∈H + ∥e − µc ∥2 k
// Push: hardest negatives d− = {∥e − µc ∥2 | e ∈ E¬c } Select top-kPclosest: Hk− Lpush = k1 e∈H − max(0, m − ∥e − µc ∥2 )2 k
// Combine push-pull Lcontrastive ← Lcontrastive + (1 − α)Lpull + αLpush C ←C +1 if C > 0 then Lcontrastive ← Lcontrastive /C Step 4: Unified Loss L ← βLCE + (1 − β)Lcontrastive return L
2) FEATURE EXTRACTION
Once the dataset was structurally cleaned, we employed several feature extraction methods such as TF-IDF Vectorizer, Word2Vec, and Glove to obtain feature embeddings. Among these, the method that we finally adopted was Glove as it showed improved and consistent results with our clustering approach. Since Glove is a well-established embedding technique, that captures global context by estimating the cooccurance probability between two words, we were able to obtain nth -dimensional embeddings (where n = 300) for each word in the vocabulary with minimal parameter tuning. Therefore, we focus less on its implementation details
3) IMBALANCE HANDLING
Imbalance handling is a necessary step when the dataset is largely inclined towards the distribution of a certain class samples. Class imbalance hinders the generalization capability of machine learning models as they get biased towards majority class samples leading to the model overfitting. To handle class imbalance, we leveraged two complementary strategies, first, at the data level, we applied ADASYN (Adaptive Synthetic Sampling) to generate synthetic minority
FIGURE 2. The pipeline consists of two phases: (1) Data Preprocessing & Augmentation, where raw text from the EMSCAD dataset is cleaned, vectorized using Word2Vec, and balanced using the ADASYN oversampling technique. (2) Deep Metric Learning, featuring a hierarchical Feed-Forward Neural Network (Input to FC-16) optimized by the proposed Centroid-based Geometric Contrastive Loss (CGCL). The CGCL formulation (bottom center) minimizes intra-class variance by pulling samples toward their respective centroids (C1, C2) while maximizing inter-class separability.
samples with a sampling ratio of 0.4, enabling the model to capture intricate representations of minority class for reliable classification. The choice of ADASYN over SMOTE (Synthetic Minority Oversampling Technique) was because ADASYN oversamples the minority class by focusing on its local context, thereby generating minority class samples that are harder to classify, whereas SMOTE produces uniformly distributed synthetic samples that may not sufficiently capture minority complexity. Second, at algorithmic level, we used a class-weighted cross-entropy loss, where we assigned weights to the class based on their frequencies. The class having majority samples was given a lower weight while the one having lesser samples was allocated a higher weight, ensuring that misclassifications of the minority class were penalized more heavily during training.
D. LOSS FORMULATION
In this section, we propose CGCL (Centroid-Guided Contrastive Loss), a custom loss function that unifies both classification and clustering within a single optimization objective. While many existing works treat fraudulent postings detection as merely a classification task, we argue that relying solely on Cross-Entropy loss is suboptimal, as it does not modulate the structure of latent-embedding space explicitly. Additionally, leveraging only the classic Contrastive loss would be ineffective due to its uniform pairwise distance computations, which can dilute the focus on harder examples. CGCL on the other hand, addresses these limitations by incorporating a classification loss that enforces discriminative decision boundaries and a centroiddriven contrastive loss to reshape the embedding space, ensuring intra-class compactness and inter-class separability.
1) WEIGHTED CROSS-ENTROPY COMPONENT
3) UNIFIED LOSS FUNCTION
To handle class-imbalance, we adopt a weighted variant of the standard Cross-Entropy loss. We do this by assigning weights to the classes in inverse proportion to their frequency in the dataset, unlike the vanilla CE loss which treats all classes equally. This approach gives more importance to the minority class while significantly reducing the dominance of majority classes. Formally, the loss is given by:
Finally, we integrate the class-balanced Cross-Entropy component (1) with the centroid-guided contrastive objective (4) into a unified loss formulation. This combined loss not only enforces robust class separation through discriminative decision boundaries but also explicitly reshapes the latent embedding space via the push-pull mechanism. Formally, the complete loss is given by:
LCE = −
N 1 X wyi log p(yi | xi ) , N i=1
(1)
where N is the total number of samples, p(yi | xi ) is the predicted probability of the ground-truth class yi for input xi , and wyi denotes the class weight.
LCGCL = β · LCE + (1 − β) · Lcontrastive ,
(5)
where β ∈ [0, 1] controls the balance between the classification term and the contrastive objective. Moreover, all embeddings are ℓ2 -normalized to maintain consistent scale across samples. For clarity, the step-by-step procedure of the proposed CGCL is summarized in Algorithm 1. E. MODEL ARCHITECTURE
2) CENTROID-GUIDED PUSH AND PULL MECHANISM
This is the core of our proposed algorithm. Taking inspiration from the well-known Contrastive loss [29], we encourage embeddings of samples from the same class to be pulled closer together, while embeddings of samples from different classes are pushed apart in the latent space. However, unlike traditional contrastive loss, where pairwise distances are computed exhaustively between samples, we design a novel centroid-driven strategy where centroids for unique classes are computed dynamically to represent their respective clusters. For every centroid, we identify the farthest top-k samples (hard-positives) belonging to the same class as the active centroid, and pull them closer to the centroid (2), thereby reinforcing intra-class compactness. Similarly, we locate top-k from other classes (hard-negatives) and push them away from the active centroid by a margin m (3), which can be adjusted during training, ensuring stronger inter-class separability. We then combine this push-pull mechanism in a unified framework, yielding a novel variant of contrastive loss, as formulated in (4). Lpull =
K 1 X
K
2
f (xhard k ) − µc 2 ,
(2)
k=1
where xhard represents the k-th hardest positive sample k (farthest from the centroid µc ). K
Lpush =
2 1 X max 0, m − ∥f (xneg , k ) − µc ∥2 K
(3)
For performing classification, we adopt a straightforward Multi-layer Perceptron (MLP) architecture. First the embeddings are obtained via Glove, which are then passed through the feed-forward neural network. The model consists of six fully connected layers, each followed by a ReLU activation function which enables the model to capture nonlinearity in the data. These deep layers significantly reduce the dimensionality of input features from 300-dimensional GloVe embeddings to a compact latent representation of size 16. The resultant embeddings are then sent into a linear layer for classification, formally the model can be expressed as: z = fembed (x)
(6)
logits = Wcls z + bcls
(7)
z is the latent representation, Wcls are the classifier weights, and bcls is the classifier bias. IV. EXPERIMENTAL RESULTS
We validate our proposed approach on a benchmark dataset in comparison with well-known studies previously done in the field of fraudulent posts detection. We perform extensive experiments to justify the effectiveness of our method, including model-level assessments through hyperparameter tuning and training configurations, and data-level evaluations to analyze the influence of various feature extractors such as GloVe, Word2Vec, and TF–IDF. on the performance of our model.
k=1
where xneg represents the hard negatives and m is the k margin parameter. Lcontrastive = (1 − α) · Lpull + α · Lpush ,
(4)
where α ∈ [0, 1] balances the contribution of push loss with respect to the pull loss.
A. IMPLEMENTATION DETAILS
Every experiment was carried out in PyTorch on a system with 64GB of RAM and an NVIDIA RTX 4090 GPU. We assessed three representations for feature extraction: TF-IDF, Word2Vec, and GloVe (300-dimensional pretrained embeddings). The proposed MLP embedder was trained using the Adam optimizer with an initial learning rate of
TABLE 3. Performance of Word2Vec, GloVe, and TF-IDF embeddings under identical hyperparameters (α = 0.6, β = 0.9, k = 3, m = 1.0). Values are reported as Macro / Micro. Best results are in bold.
Input Layer
Embedding Epochs Accuracy
300-dim GloVe Word Embeddings
F1-Score
1500 2000 3000
0.981 0.985 0.986
0.97 / 0.98 0.98 / 0.98 0.98 / 0.98 0.98 / 0.99 0.99 / 0.99 0.98 / 0.99 0.98 / 0.99 0.99 / 0.99 0.98 / 0.99
GloVe
1500 2000 3000
0.985 0.981 0.990
0.98 / 0.99 0.99 / 0.99 0.98 / 0.98 0.97 / 0.98 0.98 / 0.98 0.98 / 0.98 0.98 / 0.99 0.99 / 0.99 0.98 / 0.99
TF-IDF
500 1000
0.989 0.991
0.98 / 0.99 0.99 / 0.99 0.98 / 0.99 0.99 / 0.99 0.99 / 0.99 0.99 / 0.99
FC: 256 + ReLU Feature Extraction
Recall
Word2Vec
R300
Hidden Layer 1
Precision
R256 TABLE 4. Top hyperparameter configurations for GloVe and Word2Vec
Hidden Layer 2 FC: 128 + ReLU Dimensionality Reduction
embeddings.
Embedding
α
β
K
Accuracy
F1
GloVe
0.5 0.5 0.5 1.0 1.0
0.7 0.5 0.5 0.3 0.5
5 10 7 10 7
0.992 0.987 0.986 0.986 0.984
0.987 0.978 0.977 0.976 0.973
0.5 1.0
0.3 0.5
7 5
0.989 0.988
0.981 0.980
1.0 1.0 0.5
0.7 0.7 0.3
7 5 10
0.988 0.987 0.986
0.979 0.978 0.976
R128
Hidden Layer 3 FC: 64 + ReLU Feature Compression R64
Word2Vec
Hidden Layer 4 FC: 40 + ReLU Abstraction Layer R40
Hidden Layer 5 FC: 32 + ReLU Semantic Encoding R32
Embedding Layer
1 × 10−3 , batch size of 128, and early stopping based on validation loss. The training was carried out across various epochs such as 500, 1500, 2000 and 3000. For our CGCL loss, we explored different settings of the hyperparameters α, β, k, and m, while keeping all embeddings ℓ2 -normalized. The dataset was split into 80% training and 20% testing, with results reported on the held-out test set. TABLE 5. Best performing models for each embedding method. Listed are the optimal hyperparameters and corresponding performance scores.
FC: 16
Embedding
α
β
k
Accuracy
F1-Score
Latent Representation
Word2Vec
0.5
0.3
7
0.989
0.981
TF-IDF
0.6
0.9
3
0.991
0.990
GloVe
0.5
0.7
5
0.992
0.987
R16
Classifier 2 Classes Binary Classification FIGURE 3. Enhanced MLP Embedder Architecture. This diagram illustrates a seven-layer fully connected neural network used for feature extraction and latent representation learning.
B. RESULTS
The first set of experiments evaluated different feature extraction techniques under identical hyperparameters across multiple epochs (1500, 2000, and 3000 for Word2Vec and GloVe). For the TF-IDF vectorizer, experiments were conducted at 500 and 1000 epochs due to its faster convergence. Beyond 1000 epochs, no significant performance gains were observed with TF-IDF. To ensure a fair comparison, the
FIGURE 4. t-SNE (left) reveals the continuous internal structure of fraud templates, while UMAP (right) confirms the global separability and distinct sub-clustering of legitimate job categories.
TABLE 6. Clustering evaluation metrics for the final selected model (GloVe, α = 0.5, β = 0.7, k = 5).
Metric
Score
Silhouette Score
0.701
Adjusted Rand Index (ARI)
0.953
Normalized Mutual Information (NMI)
0.903
Homogeneity Score
0.909
Completeness Score
0.898
V-Measure
0.903
Calinski-Harabasz Score
5248.366
Davies-Bouldin Score
0.777
hyperparameters were kept fixed across all experimental settings. Details are shown in Table 3. Furthermore, since our designed loss function CGCL is highly sensitive to hyperparameter settings, in order to obtain the best performing model for each technique, we conducted several experiments with various combinations of parameters. In total, 27 possible combinations were evaluated per technique. For clarity and conciseness, we present only the top 5 results for GloVe and Word2Vec, summarized in structured form in Table 4. Across these experiments, we observed that α within the range of 0.5-1.0 exhibited consistent results, while α = 1.5 consistently failed across both techniques. Additionally, the β parameter displayed technique-specific behavior: The top results were obtained when β was set to 0.5 on GloVe-based embeddings, whereas Word2Vec performed better with β = 0.3, showing that the optimal setting of β is embedding-dependent. Table 5 presents the best-performing model for each technique, making it convenient to identify the GloVe-based embedding model as the final selected model for our study. It achieves an impressive 99.2% overall accuracy and an F1-score of 98.7% on our test set. Finally, since our loss function is inherently incomplete without the contrastive component to enforce tight clustering in the latent space and thereby complement the classification
FIGURE 5. Evolution of Latent Space. The progression from epoch 500 to 3000 shows the ’push’ mechanism creating a gradient of confidence for legitimate jobs (blue) while compacting fraudulent jobs (red) into a dense anomaly cluster. Hyperparameters: (α = 0.5, β = 0.7, k = 5)
TABLE 7. Test-set metrics (mean ± std, five seeds, fraud class). Best per column in bold.
Variant
Accuracy
Precision
Recall
CE only
0.974±.002
0.763±.036
0.677±.015
0.718±.021
0.727±.022
0.706±.023
CE weighted
0.976±.003
0.772±.043
0.719±.022
0.744±.026
0.740±.030
0.732±.028
CE ADASYN
0.989±.000
0.967±.002
0.995±.003
0.981±.001
0.982±.003
0.973±.001
CE CenterLoss
0.990±.001
0.968±.005
0.997±.002
0.982±.002
0.980±.006
0.975±.003
CGCL full
0.983±.012
0.947±.035
0.997±.002
0.971±.020
0.981±.011
0.960±.028
objective, we therefore present the evaluations of our final model on established clustering metrics in Table 6. In conclusion, our proposed approach shows astounding classification performance with GloVe-based model achieving an overall accuracy of 99.2% and an F1-score of 98.7%. In addition, the clustering performance metrics further validate our model’s robustness in a complex classification task, thereby proving its representational and discriminative efficacy.
C. VISUALIZATIONS AND LATENT SPACE ANALYSIS
To prove the effectiveness and authority of CGCL and to understand how latent space evolves during training, we visualized the high-dimensional embeddings (d = 16) projected into 2D space. These visualizations validate that the custom dual-objective loss function successfully enforces both intraclass compactness and inter-class separability.
F1
PR-AUC
MCC
2) TOPOLOGICAL STRUCTURE (T-SNE AND UMAP)
We further analyze the learned embeddings using t-SNE and UMAP projections (Figure 4). The t-SNE visualization reveals elongated and continuous structures within the fraudulent class, suggesting that fraudulent postings may vary along a spectrum of scam templates rather than forming a single homogeneous cluster. Dense regions correspond to frequently occurring patterns, while the continuous trajectories indicate gradual transitions between related fraud types. UMAP produces a similar overall structure while providing clearer global separation between classes [30]. Both legitimate and fraudulent postings exhibit internal substructures, reflecting fine-grained semantic variations within each class while maintaining a clear distinction between classes. These observations complement the quantitative clustering results and provide visual evidence that CGCL promotes both intra-class compactness and inter-class separability in the learned latent space. D. ABLATION STUDY 1) Introduction
1) TEMPORAL EVOLUTION OF CLASS SEPARATION AND INTER-CLASS COMPACTNESS
Figure 5 illustrates the evolution of the latent space over 3000 training epochs. The blue points represent Non-Fraudulent (Class 0) job postings, while the red points represent Fraudulent (Class 1) postings. At earlier epochs, there is substantial overlap between the two classes, particularly at epoch 500. As training progresses, the embeddings gradually organize into two distinct regions, and by epoch 3000 the separation becomes considerably clearer. The legitimate job postings form a tailed distribution resembling a comet-like shape. Samples near the decision boundary share lexical and structural characteristics with fraudulent postings, whereas higher-confidence legitimate examples are pushed farther away, forming the tail. At the same time, the non-fraudulent cluster becomes increasingly compact, indicating that the pull mechanism is successfully reducing intra-class variation. A similar trend is observed for fraudulent postings, which evolve from a scattered distribution into a tighter and more coherent cluster as the push mechanism separates them from the legitimate manifold.
Fraud detection suffers from severe class imbalance, which biases standard cross-entropy classifiers toward the majority (legitimate) class and suppresses fraud recall. This report evaluates five ablation variants on an identical MLP backbone, isolating the contribution of loss re-weighting, synthetic oversampling (ADASYN), center loss, and the full CGCL model. All variants share the same architecture, features, and optimizer; only the loss objective and sampling strategy differ. Results are averaged over five random seeds {7, 13, 21, 42, 99}.
2) EXPERIMENTAL SETUP a: Model and Training
Each variant uses a three-hidden-layer MLP (256–128–64, ReLU + BatchNorm) trained with Adam (lr = 10−3 , weight decay 10−4 ) for up to 500 epochs. b: Ablation Variants
• CE only — Standard cross-entropy; no imbalance handling. • CE weighted — Cross-entropy with inverse-frequency class weights.
TABLE 8. Per-seed test-set results (F1 / PR-AUC / MCC). For each seed, best F1, best PR-AUC, and best MCC across variants are individually underlined.
Seed 7 Variant
F1
PR-
Seed 13 MCC
F1
AUC
PR-
Seed 21 MCC
F1
PR-
AUC
Seed 42 MCC
F1
AUC
PR-
Seed 99 MCC
F1
AUC
PR-
MCC
AUC
CE only
.727
.741
.716
.742
.749
.732
.689
.702
.675
.734
.744
.724
.696
.699
.681
CE weighted
.751
.701
.741
.782
.780
.773
.732
.719
.718
.753
.770
.741
.702
.733
.688
CE ADASYN
.980
.982
.972
.981
.982
.973
.980
.976
.972
.981
.986
.974
.982
.982
.974
CE CenterLoss
.978
.975
.970
.980
.974
.972
.984
.981
.977
.983
.981
.977
.984
.991
.978
CGCL full
.984
.990
.978
.980
.988
.972
.978
.984
.970
.981
.984
.974
.931
.961
.905
• CE ADASYN — Cross-entropy on ADASYNoversampled training data. • CE CenterLoss — Cross-entropy + center loss (λ = 0.5) to compact intra-class embeddings. • CGCL full — Full proposed model: contrastive graph objective + cross-entropy.
3) RESULTS a: Quantitative Summary
Table 7 reports mean ± std of all metrics across five seeds. CE only and CE weighted achieve ∼97% accuracy yet score below F1 = 0.75 on the fraud class—a result of majority-class bias. Every variant explicitly addressing imbalance exceeds F1 = 0.97. b: Confusion Matrices (Seed 7)
Figure 6 shows confusion matrices for seed 7. CE only and CE weighted produce 56 and 51 false negatives respectively. Enhanced variants improve sharply: CE ADASYN yields 13 false negatives, CE CenterLoss 5, and CGCL full only 2.
4) DISCUSSION a: Class imbalance as the primary bottleneck.
The 26-point F1 gap between CE only (0.718) and CE ADASYN (0.981) with no architectural change confirms that the MLP backbone is not the bottleneck. Loss reweighting alone yields only marginal improvement (+0.026 F1), consistent with prior findings that synthetic oversampling is more effective than static weights for highly skewed distributions. b: Center loss stability.
CE CenterLoss matches CE ADASYN in F1 and MCC while achieving the lowest seed variance across all metrics. The center-loss regularizer tightens the fraud-class embedding cluster, reducing false negatives (5 vs. 13 for ADASYN on seed 7) and stabilising the decision boundary across initializations.
c: CGCL graph context and variance.
CGCL full achieves the joint-best Recall (0.997) and highest PR-AUC (0.981), and virtually eliminates false negatives (2 on seed 7). However, its F1 variance (±0.020) and MCC variance (±0.028) are notably higher than CE-based variants. Seed 99 converged significantly more slowly, suggesting sensitivity to the interaction between graph topology and the contrastive temperature hyper-parameter—an open problem for future work. V. CONCLUSION
In this paper, we introduced a Centroid-Guided Contrastive Loss (CGCL) loss function that is capable of enforcing accurate decision boundaries while consistently restructuring the latent space into dense and well-separated clusters. Our loss function dynamically updates unique clusters during training using centroid-guided push and mechanism where top-k hard-positives are pulled toward the active cluster while topk hard-negatives are pushed farther away. We also integrate Cross-Entropy loss to complement the contrastive objective, ensuring strong class discrimination alongside compact cluster formation. Extensive experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance in both classification and clustering metrics, highlighting the effectiveness of CGCL as a unified loss function. REFERENCES [1] S. Dutta and S. K. Bandyopadhyay, “Fake job recruitment detection using machine learning approach,” International Journal of Engineering Trends and Technology, vol. 68, no. 4, pp. 48–53, 2020. [2] C. Anita, P. Nagarajan, G. A. Sairam, P. Ganesh, and G. Deepakkumar, “Fake job detection and analysis using machine learning and deep learning algorithms,” Revista Geintec-Gestao Inovacao e Tecnologias, vol. 11, no. 2, pp. 642–650, 2021. [3] B. Keerthana, A. R. Reddy, and A. Tiwari, “Accurate prediction of fake job offers using machine learning,” in Machine Intelligence and Soft Computing: Proceedings of ICMISC 2020. Springer, 2021, pp. 101–112. [4] F. Shibly, S. Uzzal, and H. Naleer, “Performance comparison of two class boosted decision tree snd two class decision forest algorithms in predicting fake job postings,” Association of Cell Biology Romania, 2021. [5] A. Amaar, W. Aljedaani, F. Rustam, S. Ullah, V. Rupapara, and S. Ludi, “Detection of fake job postings by utilizing machine learning and natural language processing approaches,” Neural Processing Letters, vol. 54, no. 3, pp. 2219–2247, 2022.
FIGURE 6. Confusion matrices for all five variants (seed 7). Top row (L→R): CE only, CE weighted, CE ADASYN. Bottom row: CE CenterLoss, CGCL full.
[6] H. Qayyum, F. Ali, M. Nawaz, and T. Nazir, “Frd-lstm: a novel technique for fake reviews detection using dcwr with the bi-lstm method,” Multimedia Tools and Applications, vol. 82, no. 20, pp. 31 505–31 519, 2023. [7] V. R. Singh, P. Sampras, A. Dhage et al., “Fake job post prediction using data mining,” Journal of Scientific Research and Technology, pp. 39–47, 2023. [8] A. S. Pillai, “Detecting fake job postings using bidirectional lstm,” arXiv preprint arXiv:2304.02019, 2023. [9] A. D. Rathudi, “Fake job post prediction,” Ph.D. dissertation, Dublin, National College of Ireland, 2023. [10] V. Anbarasu, S. Selvakani, and M. K. Vasumathi, “Fake job prediction using machine learning,” ubiquity, vol. 13, no. 1, pp. 06–14, 2024. [11] H. Afzal, F. Rustam, W. Aljedaani, M. A. Siddique, S. Ullah, and I. Ashraf, “Identifying fake job posting using selective features and resampling techniques,” Multimedia Tools and Applications, vol. 83, no. 6, pp. 15 591–15 615, 2024. [12] Better Business Bureau, “Employment scams risk report 2022,” 2022, accessed: 2025-07-26. [Online]. Available: https://www.bbb. org/article/news-releases/26980-bbb-scam-tracker-risk-rsingeport [13] H. Habehh and S. Gohel, “Machine learning in healthcare,” Current genomics, vol. 22, no. 4, pp. 291–300, 2021. [14] K. Shailaja, B. Seetharamulu, and M. Jabbar, “Machine learning in healthcare: A review,” in 2018 Second international conference on electronics, communication and aerospace technology (ICECA). IEEE, 2018, pp. 910–914. [15] J. Wiens and E. S. Shenoy, “Machine learning for healthcare: on the verge of a major shift in healthcare epidemiology,” Clinical infectious diseases, vol. 66, no. 1, pp. 149–153, 2018. [16] M. F. Dixon, I. Halperin, P. Bilokon et al., Machine learning in finance. Springer, 2020, vol. 1170. [17] F. Rundo, F. Trenta, A. L. Di Stallo, and S. Battiato, “Machine learning for quantitative finance applications: A survey,” Applied Sciences, vol. 9, no. 24, p. 5574, 2019. [18] S. Ahmed, M. M. Alshater, A. El Ammari, and H. Hammami, “Artificial intelligence and machine learning in finance: A bibliometric review,” Research in International Business and Finance, vol. 61, p. 101646, 2022.
[19] B. Kelly, D. Xiu et al., “Financial machine learning,” Foundations and Trends® in Finance, vol. 13, no. 3-4, pp. 205–363, 2023. [20] W. Sultani, C. Chen, and M. Shah, “Real-world anomaly detection in surveillance videos,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 6479–6488. [21] R. Chalapathy and S. Chawla, “Deep learning for anomaly detection: A survey,” arXiv preprint arXiv:1901.03407, 2019. [22] G. Pang, C. Shen, L. Cao, and A. V. D. Hengel, “Deep learning for anomaly detection: A review,” ACM computing surveys (CSUR), vol. 54, no. 2, pp. 1–38, 2021. [23] I. H. Sarker, A. Kayes, S. Badsha, H. Alqahtani, P. Watters, and A. Ng, “Cybersecurity data science: an overview from machine learning perspective,” Journal of Big data, vol. 7, no. 1, p. 41, 2020. [24] Y. Xin, L. Kong, Z. Liu, Y. Chen, Y. Li, H. Zhu, M. Gao, H. Hou, and C. Wang, “Machine learning and deep learning methods for cybersecurity,” Ieee access, vol. 6, pp. 35 365–35 381, 2018. [25] K. Shaukat, S. Luo, V. Varadharajan, I. A. Hameed, and M. Xu, “A survey on machine learning techniques for cyber security in the last decade,” IEEE access, vol. 8, pp. 222 310–222 354, 2020. [26] G. Apruzzese, P. Laskov, E. Montes de Oca, W. Mallouli, L. Brdalo Rapa, A. V. Grammatopoulos, and F. Di Franco, “The role of machine learning in cybersecurity,” Digital Threats: Research and Practice, vol. 4, no. 1, pp. 1–38, 2023. [27] S. Bansal, “Real or fake fake job posting prediction,” 2018, accessed: 2025-07-27. [Online]. Available: https://www.kaggle.com/ datasets/shivamb/real-or-fake-fake-jobposting-prediction [28] T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” arXiv preprint arXiv:1301.3781, 2013. [29] P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan, “Supervised contrastive learning,” Advances in neural information processing systems, vol. 33, pp. 18 661–18 673, 2020. [30] L. McInnes, J. Healy, and J. Melville, “Umap: uniform manifold approximation and projection for dimension reduction. arxiv,” arXiv preprint arXiv:1802.03426, vol. 10, 2018.