Peer-reviewed and presented at 2025 International Conference on Advancement in Communication and Computing Technology (INOACC-2025); self-published by the author due to a sustained 13-month indexing delay by the organizers.
An Efficient Machine Learning-based Framework for Detection and Prevention of Frauds in Telecom Networks Praveen Hegde Associate Director-Emerging Commercial Platforms Verizon Atlanta, USA
Mishal Shah Senior Software Engineer Bloomberg LP Jersey City, USA
Abstract—Telecommunication fraud is an acute problem that leads to substantial material losses and compromises the reliability of telecom systems worldwide. Only effective and efficient detection mechanisms can help to deal with these threats, though there are certain shifts in the approaches to fraud detection. This paper evaluates the performance of AIdriven models for fraud detection in telecommunication networks using Call Detail Record (CDR) datasets. This study focuses on fraud detection in telecom networks using the Telecom CDR dataset, which contains 101,174 customer records with 17 attributes, including 8,830 fraud cases. In feature preprocessing, missing values were dealt with, followed by data scaling using Min-Max scaling and data balancing using the SMOTE technique. The dataset was trained for predictive analysis using Random Forest (RF) and XGBoost models. F1score, ROC AUC, recall, accuracy, time, and precision were used as indicators with which to compare performance of the two models. RF recorded a high level of accuracy at 99.9% while XGBoost at 99.7%. Findings show that the suggested framework successfully detects fraud with few misclassifications. Several machine learning models were evaluated and contrasted, such as RF, XGBoost, DBSCAN, RoBERTa, and K-means. Among all the models, RF was seen to give the highest performance with an accuracy of 99.9% and precision of 99.9%, recall of 99.9% and F1-score of 99.9%, XGBoost, GNN and BERT. The findings emphasize RF as the most effective model for detecting fraudulent activities in telecom networks, ensuring robust and reliable prevention of fraud.
created a dynamic landscape where mobile and digital transactions are increasingly prevalent[6]. However, this growth is not without issues on cybersecurity risks, particularly in the aspect of fraud prevention [7][8]. Since financial transactions have become closely connected with telecommunication networks, instances of fraud activity have increased[9], thus requiring the use of increased security for protecting significant financial data and ensuring the purity of financial systems[10].
Keywords—Telecommunication fraud, Machine learning, Call Detail Records (CDR), Fraud detection, Telecom security.
I. INTRODUCTION Online banking, e-commerce credit card transactions, and telecommunication fraud are just a few examples of the many fraudulent activities that have flourished in recent years, thanks to the proliferation of new technology. These scams cost businesses and consumers across the globe over a billion dollars annually [1]. In truth, fraud is quite expensive for all telecom operators when it comes to capacity and revenue lost [2]. A study by Neural Technologies estimated that in 2016, the telecom sector incurred an average loss of $249 billion due to fraud [3]. Fraud may be described as the unlawful utilization of telecommunications infrastructure, such as mobile communications, with the goal of not paying for services, abusing voice calls (or data, SMS, or MMS), deceiving subscribers, and unlawfully accessing telecommunications provider networks[4][5]. The convergence of telecommunications and financial services has
Telecommunication network fraud may present itself in several forms such as breaking into the devices that form part of the network,[11] interfering with the routing protocols, exploiting the weakness in the networking and so on[12]. These activities not only incur tangible financial but also lead to erosion of confidence and certainty in the telecommunication services [13]. Thus, telecom fraud can be performed in a wide range of ways with reference to the capabilities and opportunities offered to the fraudster [14][15]. A wide number of techniques is used for fraud detection. Fortunately, Artificial intelligence based Machine Learning (ML) technologies have become an effective weapon in the fight against telecom fraud, nowadays[16][17]. ML techniques have made it possible to develop smart methodology using data, models for learning from data and making predictions for automated inspection-free processes [18][19]. ML algorithms have the ability to analyze enormous amounts of data, and using ML-based classification allows to assess malicious actions and protect against fraud. A. Motivation and Contribution of the Study Telecom fraud poses significant financial and operational challenges, necessitating advanced solutions to identify and mitigate fraudulent activities effectively. Traditional methods often struggle with scalability and accuracy in handling large, imbalanced datasets. This research introduces a novel ML framework that combines robust preprocessing techniques, such as SMOTE for data balancing, with powerful predictive models like Random Forest and XGBoost. By benchmarking against emerging methods like BERT and GNNs, the framework not only enhances detection capabilities but also provides insights into model suitability, addressing gaps in existing research. The aim is to develop an efficient and reliable system for detecting and preventing telecom fraud, offering a significant improvement over existing approaches, through comprehensive evaluation and benchmarking. The following are the primary benefits of this research: • The study leverages a substantial and realistic Telecom CDR dataset with diverse attributes, providing a strong
•
•
•
•
•
basis for developing fraud detection models tailored to the telecom industry. The adoption of preprocessing techniques, including missing value handling, Min-Max scaling, and SMOTE-based data balancing, ensures a robust and unbiased dataset for model training and evaluation. By implementing and analyzing multiple ML approaches, including RF and XG Boost, the study contributes to understanding their relative suitability for telecom fraud detection. To evaluate the efficacy of the model in identifying telecom fraud, the research used a comprehensive set of performance indicators, including accuracy, precision, recall, F1-score, and ROC AUC. The proposed methodology demonstrates scalability and adaptability, making it a viable solution for deploying fraud detection systems in real-world telecom networks. The research compares its models to modern methods like DBSCAN, RoBERTa, and K-means, shedding light on their relative merits and shortcomings when it comes to detecting fraud.
B. Organization of the paper Here is the structure of the paper: The literature study on telecom fraud and the obstacles in detecting it is covered in Section II. The methods and models used in the research are detailed in Section III. The experimental results and performance analysis are detailed in Section IV. The paper concludes in Section V, which explores possible future study directions. II. LITERATURE REVIEW This section of the research summarises previous efforts that have addressed the topic of telecom network fraud categorization and detection. The majority of the publications under consideration concentrate on machine learning methods and frameworks that have been created to better identify and stop telecom fraud. Some reviews are: H. palivela et al. (2024), research aims to identify transaction patterns to help algorithms recognize and flag scammers. The proposed approach incorporates methods of ensemble learning, including Gradient Boost, Random Forest, Logistic Regression, and Voting Classifiers, along with TABLE I. Reference H. Palivela et al. (2024) Jun Li et al. (2024)
hyperparameter tuning to prevent overfitting. It effectively manages false positives through careful review, achieving an accuracy of 99.59%, surpassing other innovative models. [20]. Jun Li et al. (2024) proposed RoBERTa-MHARC, a model for text-based telecom fraud detection that combines RoBERTa with residual connections and a multi-head attention feature. The approach begins by sorting the CCL2023telecom fraud dataset into many groups. Experimental results demonstrated that our model attains better results than the benchmarks on several datasets include an F1 score97.65 on the FBS dataset, 98.10 on our own dataset, and 93.69 on the news dataset [21]. V. Chang et al. (2024), this research looked at how well the SMOTE performed when compared to two other methods: random under-sampling and over-sampling. The results show that random under-sampling achieves good recall (92.86%) but poor precision, whereas SMOTE makes a better accuracy (86.75%) and more consistent F1 score (73.47%), but its recall is somewhat lower[22]. R. Li et al. (2023), provide a method for reducing attributes that takes into account MCIR in order to enhance the quality of the data. They proceed to address partial or missing data by developing a MCIR-RGAD based on the concept of maximum consistent blocks. Finally, the MCIR-RGAD method is shown to be an efficient solution for lowering computing time, enhancing data quality, and processing incomplete data in experimental findings using actual telephony fraud data and UCI data[23]. S. K. Hashem et al. (2022), to manage the relative importance of valid and fraudulent transactions, you may want to think about class weight-tuning hyperparameters. The findings display that LGBM and XGBoost reach the best level requirements with ROC-AUC=0.95, precision0.79, recall0.80, F1 score0.79, and MCC0.79. Through the use of DL and the Bayesian optimization technique, we are also able to achieve the following criteria: ROC-AUC=0.94, precision=0.80, recall=0.82, F1 score=0.81, and MCC=0.81[24]. Table 1 presents the comparative analysis of linked studies according to their Key Findings, limitations, future work, and Focus On.
SUMMARY OF RELATED WORK FOR TELCO FRAUD DETECTION USING MACHINE LEARNING TECHNIQUES
Methodology Ensemble learning methods: Gradient Boost, Random Forest, Logistic Regression, and Voting Classifiers with hyperparameter tuning. The RoBERTa-MHARC paradigm encompasses RoBERTa, a multi-head attention mechanism, and residual connections.
V. Chang et al. (2024)
Comparison of random undersampling and SMOTE for class balancing in fraud detection models.
R. Li et al. (2023)
Attribute reduction algorithm (MCIR) and rough-gain anomaly detection (MCIR-RGAD) for handling incomplete data and improving data quality.
Dataset Sizable dataset of anonymized credit card transactions CCL2023 telecom fraud dataset and benchmark datasets Fraudulent transaction datasets.
Results Achieved an accuracy of 99.59%, surpassing other models.
Telecommunica tion fraud data and UCI dataset.
Reduced computation time, improved data quality, and effective processing of incomplete data.
F1 score of 97.65 (FBS dataset), 98.10 (own dataset), 93.69 (news dataset). Random undersampling: recall of 92.86%. SMOTE: accuracy of 86.75%, F1 score of 73.47%.
Limitations Cannot entirely prevent fraudulent transactions; relies on post-detection manual review. Requires robust preprocessing for varied dataset quality; computationally intensive. Random under-sampling compromises precision; SMOTE introduces synthetic data, which may not reflect realworld cases. Limited scalability to large datasets; focused mainly on incomplete data scenarios.
Future Work Enhance scalability and real-time detection to further reduce false positives. Extend the model to real-time and multilingual fraud detection scenarios. Investigate advanced balancing techniques and their integration with deep learning models. Extend to more diverse datasets and improve scalability for largescale implementations.
S. K. Hashem et al. (2022)
Class weight-tuning hyperparameters, majority voting ensemble learning method, LightGBM, XGBoost, and Bayesian optimization for hyperparameter tuning.
Fraudulent transaction datasets.
III. METHODOLOGY This research aims to assess models powered by artificial intelligence for the purpose of detecting fraud in telecommunications networks. The flow diagram in Figure 1 shows the research methodology. The process begins with the collection of a telecom Call Detail Record (CDR) dataset containing key attributes such as account length, call minutes, charges, and fraud status. Initially, the dataset underwent preprocessing, including handling missing values, noise reduction, and balancing data using the SMOTE technique. Following this, feature scaling was performed using Min-Max normalization to the dataset. The data that had already been preprocessed was then divided into two parts: training and testing. At last, we put a number of ML models into action, including XGBoost and RF, and we measured their efficacy using measures including AUC, F1-score, recall, precision, and accuracy. The outcomes provide a comprehensive comparison of these models with DBSCAN, RoBERTa and K-means in detecting fraudulent activities effectively. Data Preprocessing
Handling Missing Values
Telecom CDR Dataset
Data Scaling with Min-Max Scaling
Data balancing with SMOTE
Data splitting Classification with RF and XGBoost
Training
Testing
LightGBM and XGBoost: ROC-AUC = 0.95, precision = 0.79, recall = 0.80, F1 score = 0.79, MCC = 0.79.
Requires extensive computational resources; ensemble methods may increase complexity and reduce interpretability.
international use, total income from SMS, and current status of fraud.
Fig. 2. Telecom Fraud Data Distribution
B. Data Preprocessing The concept of ‘‘data preprocessing’’ is used to describe any action taken on unprocessed data before it is used. Data preparation is the process of transforming raw data into a more usable format [25][26]. Raw data undergoes a series of treatments known as data preparation to make it more useable and processed. Below is a list of the pre-processing steps: • Handling Missing Values: Missing values were identified and removed to maintain the integrity of the dataset. Transactions with missing critical features were excluded. • Data Scaling with Min-Max Scaling: is that the features will be rescaled to ensure the[27] mean and the standard deviation to be 0 and 1, respectively. The Min-Max Scaling formulate as equ.1. 𝑥′ = 𝑥 −
RF and XGBoost compare with DBSCAN, RoBERTa and K-means
Fig. 1. Flowchart for Detection of Frauds in Telecom
The overall steps of the flowchart for of machine learning models for fraud detection are provided in below: A. Data Collection This study used the Telecom CDR Dataset1. The dataset used for this study includes 17 attributes from 101,174 customers, with 8,830 fraud, shows in figure 2. The following variables are included in it: state, account duration, phone number, international plan, mail plan, amount of voice mail messages, call minutes, costs for day, evening, and night,
1
(𝑥𝑚𝑖𝑛 ) (𝑥𝑚𝑎𝑥 −𝑥𝑚𝑖𝑛 )
(1)
where X is an initial value, X' is a transformed value, and Xmax and Xmin are a highest and lowest points for that characteristic, respectively.
Performance matrix including accuracy, precision, recall, f1score, and AUC
Results
Apply these methods to real-time fraud detection systems and reduce computational overhead.
https://github.com/jamesrawlins1000/Telecom-CDR-Dataset-
C. Data balancing with SMOTE The models were made more accurate by using the SMOTE. To improve the representation of under-represented groups in datasets, SMOTE employs an oversampling technique that, given an existing class sample, creates additional samples from that sample [28]. The technique generates non-duplicative minority class samples by convexly combining two or more neighboring data samples in the feature space that are chosen at random. Figure 3 shows the after SMOTE techniques applied balanced classes non-fraud and fraud.
Fig. 3. Telecom Fraud Data Distribution
D. Data splitting Data splitting involves separating a dataset into smaller parts so that each phase may be trained, validated, and tested independently. The training set is used to train the model, while the test set is used to evaluate its performance. Two portions of the dataset were separated: 80% for training and 20% for testing.
𝑓(𝓍) = 𝑎𝑟𝑔𝑚𝑎𝑥𝑃(𝑌 = 𝑦|𝑋 = 𝓍) (5) 2) Extreme gradient boosting (Xgboost) GBM are among the most effective algorithms in supervised learning, and Xgboost is an implementation of this method [31][32]. This tool is versatile enough to tackle both classification and regression issues. Because of its fast execution speed out of core computing, Xgboost is favored among data scientists [33] The Xgboost operates in the following manner: Consider the following scenario: we have a datasetDS with m features and n examplesDS={(Xi, Yi ): i=1……n, ∈ ℝ𝑚 , 𝑦𝑖 ∈ ℝ}. An ensemble tree model may be constructed using equation 6 to forecast the output, which is denoted as yi: Å. 𝑖 = ∅(𝑥𝑖 ) = ∑𝐾 (6) 𝑘=1(𝑥𝑖 ), 𝑓𝑘 ∈ 𝔉 The number of trees in the model is represented by K. Finding the optimal collection of functions by minimising the loss and regularisation goal in accordance with equation 7 is necessary to answer the aforementioned equation, where fk denotes the (k-th tree). ℒ(∅) = ∑𝑖 1(𝑦𝑖 , Å. 𝑖) + ∑̇ 𝑘 Ω(𝑓𝑘 )
E. Prediction with ML-based RF and XGBoost model In this study, have utilized both RF and XGBoost models for predictive analysis. RF leverages multiple DT to provide robust predictions for regression and classification, optimizing loss functions like squared error and zero-one loss. XGBoost, known for its efficiency, enhances accuracy through iterative loss and regularization minimization. Together, these models ensure high-performance predictions across the dataset.
(7) The loss function, denoted as l, is a difference between an actual output yi and a projected output yˆi. this helps prevent the model from being over-fit, where O is a measure of the model's complexity. together with the following equations 8:
1) Random Forest Trees are the backbone of the RF ensemble, and each tree's performance is affected by a collection of random variables [29]. To put it another way, we may think of the input or predictor variables X= (X1,..., Xp)T as having actual values and the response variable Y as having real values. Then, we can imagine an unknown joint distribution PXY (X, Y) for these variables. The aim is to find a function (X) that can predict Y [30] The prediction function is chosen by the loss function L(Y, f(X)) with the aim of minimizing the anticipated value of the loss equation 2.
F. Performance Metrics These performance measures are used to assess the efficacy and precision of the suggested framework for telecom network fraud detection and prevention based on machine learning, with the use of the Telecom-CDR dataset:
𝐸𝑋𝑌 (𝐿(𝑌, 𝑓(𝑋))) (2) If X and Y are joint distributions and the subscripts indicate expectations with regard to that distribution. It seems to reason that L(Y,f(X)) would be a measure of the proximity of f(X) to Y; it would penalise f(X)values that are far by Y. Options for Lare's squared error loss In regression, L(Y,f(X)) = (Y−f(X))2 and in classification, zeroone loss is equal to equation 3.
𝐿(𝑌, 𝑓(𝑋)) = 𝐼(𝑌 ≠ 𝑓(𝑋)) = {
0 𝑖𝑓 𝑌 = 𝑓(𝑋) 1 𝑜𝑡ℎ𝑒𝑟𝑤𝑖𝑠𝑒
(3) Equation 4 is a conditional expectation obtained by minimising EXY (L(Y,f(X))) with respect to squared error loss.
𝑓(𝓍) = 𝐸(𝑌|𝑋 = 𝓍) (4) sometimes called a regression function. Assuming that Y stands for the set of possible Yi values in a classification situation, the Bayes rule equ.5 may be written as the minimum of EXY(L(Y,f(X))) for a loss of zero.
1
Ω(𝑓𝑘 ) = 𝛾𝑇 + 𝜆 ∥ 𝑤 ∥2
(8) In equation 8, T stands for the total number of tree leaves, and w for the weight of every leaf. 2
• TP (True Positive): Positive sample count that was accurately determined. • TN (True Negative): The amount of negative samples that were accurately labelled as negative. • FP (False Positive): The sum of all samples that were labelled as negative but were really positive. • FN (False Negative): Amount of positive samples that were mistakenly labelled as negative. 1) Accuracy: Accuracy is a ratio of a number of correctly classified samples to a total number of samples. The corresponding equation (9) is given below: (TP)+(TN) Accuracy = (9) (TP + TN + FP + FN)
2) Precision: Accurate positive predictions inside a positive class are measured as precision. Precision is best understood as follows, according to mathematical equation (10): (𝑇𝑃) 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛 = (10) (𝑇𝑃)+ (𝐹𝑃)
3) Recall: The imbalance issue relies heavily on the recall metric, which reliably measures the percentage of significant categories that are discovered[34]. Below are the matching equations (11): (𝑇𝑃) 𝑅𝑒𝑐𝑎𝑙𝑙 = (11) (𝑇𝑃)+(𝐹𝑛)
4) F1-Score: F1 is another all-encompassing indicator for assessing unbalanced issues; it is described as equation (12): 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛×𝑅𝑒𝑐𝑎𝑙𝑙 𝐹1 − 𝑆𝑐𝑜𝑟𝑒 = 2 × (12) 𝑃𝑟𝑒𝑐𝑖𝑠𝑖𝑜𝑛+𝑅𝑒𝑐𝑎𝑙𝑙 5) Receiver operating characteristics curve (ROC): A method for representing the efficacy of a model graphically is the ROC. This index fully captures the sensitivity and specificity factors, which are continuous. The correlation among recall and precision is seen by the curve. IV. RESULT & DISCUSSION
Fig. 5. Random Forest ROC Curve
The results of the study's classification techniques are examined here, with a particular emphasis on the RF and XGBoost models' capabilities for detecting telecom fraud. These models were benchmarked and validated on the Telecom CDR dataset using precision, accuracy, recall, F1score and ROC AUC. A Python-based simulation environment using Jupyter Notebook running on Google Colab was used to carry out the experimental results. In the analysis, the required libraries where utilized and these included; Scikit-learn, NumPy, Pandas, Keras, Matplotlib. A local machine with an Intel i7 processor, 16 GB of RAM, and an NVIDIA GTX 1660 Ti GPU for best performance was used to run the simulation. The table II provide the performance of parameters for telco fraud detection. TABLE II. Matrix Accuracy precision Recall F1-Score ROC AUC
Figure 5 displays the ROCcurve, which demonstrates how well the Random Forest Classifier differentiates among the classes. The curve closely follows the top-left corner, and the AUC is 1.0, indicating perfect classification.
RF AND XGBOOST MODEL PERFORMANCE ON THE TELECOM CDR DATASET RF 99.9 99.9 99.9 99.9 99.9
XGBoost 99.7 99.8 99.8 99.8 99.8
The table II shows the performance of RF and XGBoost model on the Telecom CDR Dataset. The models achieved 99.9% and 99.7% accuracy respectively and other measures performance for fraud detection.
Fig. 6. Confusion Matrix for XGBoost
In figure 6, the confusion matrix represents the performance of an XGBoost classifier. It shows 273,562 true negatives and 25,833 true positives, indicating correct predictions for classes 0 and 1, respectively. There are 309 FP and 296 FN, which are misclassifications. The model works as intended, however there is room for improvement when compared to the ideal categorisation situation due to the occurrence of false positives and false negatives.
Fig. 4. Random Forest Confusion Matrix
Figure 4 illustrates the confusion matrix, which indicates the Random Forest Classifier's performance. It shows the TP (26127) and TN (273871), indicating the correct predictions for classes 1 and 0, respectively. There are no false positives (0), and only 2 false negatives, showcasing excellent predictive performance.
Fig. 7. ROC Curve for XGBoost
Figure 7 shows the ROC curve, which shows how well the XGBoost model performed. The dashed diagonal line stands for random guessing, whereas the orange curve towards the top-left corner signifies good classification performance. The model's exceptional capability to differentiate among positive and negative classes is shown by an AUC of 0.99. TABLE III.
PERFORMANCE COMPARISON OF MACHINE LEARNING MODELS FOR FRAUD DETECTION
Matrix Accuracy precision Recall F1-score
DBSCAN [35] 0.9039 0.9938 0.8989
RoBERTa 98.10 98.11 98.12 98.12
Kmeans[36] 0.961 0.471 0.727 0.533
RF 99.9 99.9 99.9 99.9
XGBoost 99.7 99.8 99.8 99.8
The table below compares the efficiency of several ML models trained on a dataset consisting of call detail recordings in order to identify instances of fraud in telecom networks. XGBoost achieved exceptional performance with an accuracy99.9%, precision99.9%, recall99.9%, and an F1score99.9%, making it the most effective model for this task. RF followed closely with an accuracy99.7%, precision99.8%, recall99.8%, and an F1-score99.8%, showcasing its robustness. RoBERTa demonstrated high accuracy at 98.10%, precision at 98.11%, recall at 98.12%, and an F1-score of 98.12%, performing well but slightly below tree-based models. K-means clustering, while effective in some scenarios, struggled with a lower accuracy0.961, precision0.471, recall0.727, and an F1-score0.533, highlighting its limitations in detecting fraud. DBSCAN, another clustering technique, showed a recall of 0.9938 and precision of 0.9039 but suffered from a lower F1-score of 0.8989, indicating challenges in handling imbalanced datasets. Overall, XGBoost and RF emerged as the most efficient and reliable models for fraud detection in telecom networks. V. CONCLUSION & FUTURE WORK Fraud detection in telecommunication networks is an ongoing challenge that requires innovative approaches to combat the increasing sophistication of fraudulent activities. The proposed machine learning-based framework for telecom fraud detection shows exceptional predictive performance. The Random Forest classifier achieved near-perfect results with an AUC of 1.0, while XGBoost delivered comparable performance with an AUC of 0.99. This study affirms the effectiveness of ensemble learning approaches, including RF and XGBoost, in managing class imbalance and producing optimal accuracy in detecting fraud ready for further largescale application in the telecoms networks. As the research shows the feasibility of using Random Forest and XGBoost models with good performance metrics for the detection of telecom frauds, the study has some drawbacks, especially due to the relative imbalance of the given dataset, which often necessitates oversampling tools like SMOTE, which may introduce synthetic biases. Further, due to the dependence on computational resources the models might possess limited supply in larger real-world applications with larger datasets. Future research could consider deep learning architectures, advanced ensemble techniques, and real-time fraud detection systems. To improve interpretability and model flexibility for a wider range of deployment, more advanced data set holding and explainable AI could be implemented.
[4]
[5]
[6]
[7]
[8]
[9]
[10]
[11]
[12]
[13]
[14]
[15]
[16]
[17]
[18]
[19]
[20]
[21]
REFERENCES [1]
[2]
[3]
H. Sinha, “An examination of machine learning-based credit card fraud detection systems,” Int. J. Sci. Res. Arch., vol. 12, no. 01, pp. 2282–2294, 2024, doi: https://doi.org/10.30574/ijsra.2024.12.2.1456. Suhag Pandya, “A Machine and Deep Learning Framework for Robust Health Insurance Fraud Detection and Prevention,” Int. J. Adv. Res. Sci. Commun. Technol., pp. 1332–1342, Jul. 2023, doi: 10.48175/IJARSCT-14000U. A. Chouiekh and E. H. I. El Haj, “ConvNets for fraud detection analysis,” Procedia Comput. Sci., vol. 127, pp. 133–138, 2018,
[22]
[23]
[24]
doi: 10.1016/j.procs.2018.01.107. M. R. S. and P. K. Vishwakarma, “The Assessments Of Financial Risk Based On Renewable Energy Industry,” Int. Res. J. Mod. Eng. Technol. Sci., vol. 06, no. 09, pp. 758–770, 2024. R. Tandon, “The Machine Learning Based Regression Models Analysis For House Price Prediction,” Int. J. Res. Anal. Rev., vol. 11, no. 3, pp. 296–305, 2024. N. Ahmadzai, H. Mohammadi, and N. Mangal, “Data Mining Techniques in Telecommunication Company,” J. Res. Appl. Sci. Biotechnol., 2023, doi: 10.55544/jrasb.2.1.12. Olajide Soji Osundare, Chidiebere Somadina Ike, Ololade Gilbert Fakeyede, and Adebimpe Bolatito Ige, “x,” Comput. Sci. IT Res. J., vol. 4, no. 3, pp. 458–477, 2023, doi: 10.51594/csitrj.v4i3.1499. R. Tandon, “Prediction of Stock Market Trends Based on Large Language Models,” J. Emerg. Technol. Innov. Res., vol. 11, no. 9, pp. a615–a622, 2024. L. X. Liu, Y. Liu, X. Ruan, and Y. Zhang, “Big Data Analysis with No Digital Footprints Available: Evidence from Cyber-Telecom Fraud,” SSRN Electron. J., 2021, doi: 10.2139/ssrn.3991369. F. E. Onuodu and S. B. Nnaa, “An Enhanced Fraud Detection Model using Neural Networks for Telecommunications and Smart Cards in Nigeria,” London J. Res. …, vol. 20, no. 2, 2020. M. Bajpai, “Available online www.jsaer.com Fraud Detection and Prevention in Telecommunication Data and Voice Networks,” vol. 7, no. 6, pp. 302–305, 2020. V. Kolluri, “Revolutionizing Healthcare Delivery: The Role of AI and Machine Learning in Personalized Medicine and Predictive Analytics,” Well Test. J., vol. 33, no. 02, 2024. R. Li, Y. Zhang, Y. Tuo, and P. Chang, “A Novel Method for Detecting Telecom Fraud User,” in Proceedings - 2018 3rd International Conference on Information Systems Engineering, ICISE 2018, 2018. doi: 10.1109/ICISE.2018.00016. D. S. Terzi, Ş. Sağıroğlu, and H. Kılınç, “Telecom fraud detection with big data analytics,” Int. J. Data Sci., vol. 6, no. 3, p. 191, 2021, doi: 10.1504/ijds.2021.121090. B. Boddu, “The Future of Database Administration: AI Integration and Innovation,” J. Sci. Eng. Res., vol. 11, no. 1, pp. 312–316, 2024. A. J. Rahul Dattangire, Ruchika Vaidya, Divya Biradar, “Exploring the Tangible Impact of Artificial Intelligence and Machine Learning: Bridging the Gap between Hype and Reality,” 2024 1st Int. Conf. Adv. Comput. Emerg. Technol., pp. 1–6, 2024. B. A. Yehya and N. Salhab, “Telecommunications Fraud Machine Learning-based Detection,” 2023 4th Int. Conf. Data Anal. Bus. Ind. ICDABI 2023, no. August, pp. 656–661, 2023, doi: 10.1109/ICDABI60145.2023.10629612. J. Thomas, H. Volikatla, J. Vummadi, and R. Shah, “AI-Enhanced Demand Forecasting Dashboard Device Having Interface for Optimal Inventory Management.” 2024. E. M. D. Djomadji, K. I. Basile, T. T. Christian, F. V. K. Djoko, and M. E. Sone, “Machine Learning-Based Approach for Identification of SIM Box Bypass Fraud in a Telecom Network Based on CDR Analysis: Case of a Fixed and Mobile Operator in Cameroon,” J. Comput. Commun., 2023, doi: 10.4236/jcc.2023.112010. H. Palivela et al., “Optimisation of Deep Learning based Model for Identification of Credit Card Frauds,” IEEE Access, vol. 12, no. September, pp. 125629–125642, 2024, doi: 10.1109/ACCESS.2024.3440637. J. Li, C. Zhang, and L. Jiang, “Innovative Telecom Fraud Detection: A New Dataset and an Advanced Model with RoBERTa and Dual Loss Functions,” Appl. Sci., vol. 14, no. 24, p. 11628, Dec. 2024, doi: 10.3390/app142411628. V. Chang, B. Ali, L. Golightly, M. A. Ganatra, and M. Mohamed, “Investigating Credit Card Payment Fraud with Detection Methods Using Advanced Machine Learning,” Information, vol. 15, no. 8, p. 478, Aug. 2024, doi: 10.3390/info15080478. R. Li, H. Chen, S. Liu, K. Wang, B. Wang, and X. Hu, “TFD-IISCRMCB: Telecom Fraud Detection for Incomplete Information Systems Based on Correlated Relation and Maximal Consistent Block,” Entropy, 2023, doi: 10.3390/e25010112. S. K. Hashemi, S. L. Mirtaheri, and S. Greco, “Fraud Detection in Banking Data by Machine Learning Techniques,” IEEE Access,
[25]
[26]
[27]
[28]
[29]
[30]
[31]
[32]
[33]
[34]
[35]
[36]
2023, doi: 10.1109/ACCESS.2022.3232287. V. Agarwal, “Research on Data Preprocessing and Categorization Technique for Smartphone Review Analysis,” Int. J. Comput. Appl., 2015, doi: 10.5120/ijca2015907309. K. Patel, “Quality Assurance In The Age Of Data Analytics: Innovations And Challenges,” Int. J. Creat. Res. Thoughts, vol. 9, no. 12, pp. f573–f578, 2021. L. Guanyu, “Leveraging Machine Learning for Telecom Banking Card Fraud Detection : A Comparative Analysis of Logistic Regression , Random Forest , and XGBoost Models,” vol. 1, no. 1, pp. 13–27, 2024. N. Abid, “Empowering Cybersecurity : Optimized Network Intrusion Detection Using Data Balancing and Advanced Machine Learning Models,” TIJER, vol. 11, no. 12, 2024. H. Sinha, “Analysis of anomaly and novelty detection in time series data using machine learning techniques,” Multidiscip. Sci. J., vol. 7, no. 06, 2024, doi: https://doi.org/10.31893/multiscience.2025299. A. Cutler, D. R. Cutler, and J. R. Stevens, “Ensemble Machine Learning,” Ensemble Mach. Learn., no. January, 2012, doi: 10.1007/978-1-4419-9326-7. M. Gopalsamy, “Evaluating the Effectiveness of Machine Learning (ML) Models in Detecting Malware Threats for Cybersecurity,” Int. J. Curr. Eng. Technol., vol. 13, no. 06, 2023, doi: : https://doi.org/10.14741/ijcet/v.13.6.4. M. Gopalsamy, “Identification and Classification of Phishing Emails Based on Machine Learning Techniques to Improvise Cybersecurity,” Int. J. Sci. Adv. Res. Technol., vol. 10, no. 10, pp. 47–57, 2024. A. Ibrahem Ahmed Osman, A. Najah Ahmed, M. F. Chow, Y. Feng Huang, and A. El-Shafie, “Extreme gradient boosting (Xgboost) model to predict the groundwater levels in Selangor Malaysia,” Ain Shams Eng. J., vol. 12, no. 2, pp. 1545–1556, 2021, doi: https://doi.org/10.1016/j.asej.2020.11.011. X. Hu, H. Chen, H. Chen, X. Li, J. Zhang, and S. Liu, “Mining Mobile Network Fraudsters with Augmented Graph Neural Networks,” Entropy, vol. 25, no. 1, pp. 1–16, 2023, doi: 10.3390/e25010150. M. A. Jabbar and S. Suharjito, “Fraud Detection Call Detail Record Using Machine Learning in Telecommunications Company,” Adv. Sci. Technol. Eng. Syst. J., vol. 5, no. 4, pp. 63– 69, Jul. 2020, doi: 10.25046/aj050409. Z. Aziz and R. Bestak, “Insight into Anomaly Detection and Prediction and Mobile Network Security Enhancement Leveraging K-Means Clustering on Call Detail Records,” Sensors, vol. 24, no. 6, p. 1716, Mar. 2024, doi: 10.3390/s24061716.