Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Sci Rep . 2026 Apr 11;16:12187. doi: 10.1038/s41598-026-44399-3 Search in PMC Search in PubMed View in NLM Catalog Add to search Football cybersecurity threat severity prediction using multi-head transformer-based deep learning models Basma M Hassan Basma M Hassan 5 Faculty of Artificial Intelligence, Kafrelsheikh University, Kafrelsheikh, 33516 Egypt Find articles by Basma M Hassan 5, ✉ , Fahad Algarni Fahad Algarni 1 College of Computing and Information Technology, University of Bisha, Bisha, Saudi Arabia Find articles by Fahad Algarni 1 , Rayan Alshamrani Rayan Alshamrani 2 Department of Information Technology, College of Computers and Information Technology, Taif University, P.O. Box 11099, 21944 Taif, Saudi Arabia Find articles by Rayan Alshamrani 2 , Ashrf Althbiti Ashrf Althbiti 2 Department of Information Technology, College of Computers and Information Technology, Taif University, P.O. Box 11099, 21944 Taif, Saudi Arabia Find articles by Ashrf Althbiti 2 , Abdullah Albalawi Abdullah Albalawi 3 Department of Computer Science, College of Computing and Information Technology, Shaqra University, Shaqra, Saudi Arabia Find articles by Abdullah Albalawi 3 , Atef Ismail Atef Ismail 4 Physics Department, Al Azhar University, Asyut, 71524 Egypt Find articles by Atef Ismail 4 Author information Article notes Copyright and License information 1 College of Computing and Information Technology, University of Bisha, Bisha, Saudi Arabia 2 Department of Information Technology, College of Computers and Information Technology, Taif University, P.O. Box 11099, 21944 Taif, Saudi Arabia 3 Department of Computer Science, College of Computing and Information Technology, Shaqra University, Shaqra, Saudi Arabia 4 Physics Department, Al Azhar University, Asyut, 71524 Egypt 5 Faculty of Artificial Intelligence, Kafrelsheikh University, Kafrelsheikh, 33516 Egypt ✉ Corresponding author. Received 2025 Oct 1; Accepted 2026 Mar 11; Collection date 2026. © The Author(s) 2026, corrected publication 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13076683 PMID: 41965384 Abstract Cybersecurity incidents targeting professional football organizations pose significant operational, financial, and reputational risks. Accurate assessment of incident severity is therefore essential for effective prioritization and response. This study proposes a Multi-Head Transformer–based neural network for predicting cybersecurity incident severity scores in football ecosystems using heterogeneous tabular data. The model captures complex interactions among technical attack characteristics and football-specific contextual features, including club and league attributes. The proposed approach is evaluated against established machine learning baselines, including Random Forest and XGBoost. Experimental results demonstrate that the Transformer model consistently outperforms baseline methods, achieving high predictive accuracy with strong generalization performance. Robustness is validated through strict data partitioning, 5-fold cross-validation, and multiple random seeds. Furthermore, SHAP-based explainability analysis confirms that predictions are driven by multiple interacting features rather than any single dominant variable. The results indicate that the proposed framework provides a reliable and interpretable solution for severity assessment in football cybersecurity environments. By enabling data-driven incident prioritization and informed decision-making, this work contributes a practical and scalable tool for enhancing cyber resilience in professional football organizations. Keywords: Multi-head transformer, Severity score prediction, Football cybersecurity, Cyber threat risk assessment, Sports analytics Subject terms: Engineering, Mathematics and computing Introduction Soccer, having the largest fan base, is gradually incorporating a range of innovative technologies as part of touching lovers’ entertainment, boosting the efficiency of the game, and improving organizational procedures 1 . Technologies impact almost every aspect of football, including analytical video software, wearable recording athlete activity, ticketing systems, and live stream technology. Nonetheless, the increased use of these technologies poses the highest risk of cyber vulnerability that threatens the operations, data integrity, and stakeholder image of the sport 2 . Due to its vast financial scale amounting to billions of dollars and growing reliance on digital technologies, the football industry has become an attractive target for cybercriminals 3 . These threats include unauthorized access leading to key team systems being locked through ransomware, fan and player databases being compromised, alteration of match results, and damage to stadium and or broadcast infrastructure. However, to the best of the researchers’ knowledge, there is scant information on cybersecurity in football, an area that deserves attention if football is to build a future resistance to cyber-attacks 4 . Cybersecurity at the FIFA 2022 World Cup: challenges and strategies The FIFA 2022 World Cup in Qatar, like other major global events, was subject to significant cybersecurity risks. These ranged from potential attacks on critical infrastructure to the disruption of essential digital services required for the event’s seamless operation. This section examines the cybersecurity challenges encountered, the strategic responses deployed, and the broader implications for securing international sporting events 5 . Nature of Cybersecurity Threats: the extensive digitalization of World Cup operations including systems related to visas, accommodation, transportation, ticketing, and live broadcasting created multiple potential attack vectors. Cyber threats included hacking attempts, Distributed Denial of Service (DDoS) attacks, data breaches, and digital espionage. The global visibility and geopolitical relevance of the event also elevated concerns over politically motivated cyberattacks intended to cause disruption or disseminate propaganda 6 . Cybersecurity Infrastructure and Preparedness: in anticipation of these threats, Qatar invested heavily in strengthening its cybersecurity infrastructure. This included the enhancement of existing digital systems and the implementation of advanced security protocols. The Qatari government worked closely with leading cybersecurity firms to safeguard critical infrastructure, including energy grids, transportation systems, and telecommunications networks 7 . Collaboration with International Experts: acknowledging the complexity of the cyber threat landscape, Qatar partnered with international cybersecurity leaders. Collaborative efforts with countries such as the United States and the United Kingdom enabled the sharing of intelligence, co-development of strategic frameworks, and the upskilling of local cybersecurity professionals through targeted training programs 8 . Public and Private Sector Engagement: a coordinated, multi-stakeholder approach was essential to the success of Qatar’s cybersecurity measures. Government entities, private sector organizations, and international partners worked in concerts to identify vulnerabilities, share real-time threat intelligence, and implement unified incident response procedures 9 . Incident Response and Crisis Management: a cornerstone of the cybersecurity framework was the formation of a specialized incident response team tasked with continuous monitoring of digital activities, early threat identification, and the execution of rapid countermeasures. The team conducted regular simulations, and crisis drills to test response effectiveness and refine operational protocols 10 . Cybersecurity challenges and predictive modeling in professional football organizations The cybersecurity risks emerging from digital transformation in football clubs and the limitations of current predictive models, setting the stage for advanced approaches like Transformer-based architectures. The accelerating digital transformation of professional football organizations has fundamentally reshaped the operational, financial, and strategic landscape of the sport. Modern football clubs increasingly rely on complex digital infrastructures to manage a broad array of services, ranging from tactical player analytics and financial operations to fan engagement platforms and ticketing systems 11 . While these advancements have enhanced operational efficiency and enriched stakeholder experiences, they have also introduced new and significant cybersecurity risks. Cyber-attacks targeting football clubs can result in the compromise of sensitive data, disruption of critical services, and severe reputational and financial repercussions. Notably, the recent surge in sophisticated cyber threats within the sports sector underscores the urgent need for advanced, domain-specific cybersecurity solutions 12 . Despite growing awareness of these vulnerabilities, current risk assessment and predictive frameworks in football cybersecurity largely rely on traditional regression models and classical machine learning techniques 13 . These conventional approaches often exhibit substantial limitations in capturing the non-linear, high-dimensional, and heterogeneous nature of cybersecurity incident data, which typically involve mixed categorical and numerical features, complex interdependencies, and dynamically evolving threat landscapes. Consequently, there is a pressing demand for more sophisticated predictive models capable of accurately estimating the severity of cyber threats to enable proactive incident response and effective risk management in football organizations 14 . In recent years, Transformer-based architecture has demonstrated remarkable success in natural language processing and sequence modeling tasks, owing to their ability to model intricate dependencies via multi-head self-attention mechanisms 15 . Building upon these advancements, Transformer variants tailored for tabular data, such as TabTransformer, have emerged, offering notable improvements over traditional deep learning and ensemble models in structured data domains. However, their application in the field of sports cybersecurity particularly for continuous severity score prediction tasks in football digital ecosystems remains critically underexplored 16 . The ongoing digitalization of professional football operations has ushered in a new era of data-driven decision-making and operational management. From player performance analytics and real-time tactical systems to fan engagement applications and financial oversight platforms, football clubs increasingly depend on interconnected digital infrastructures 17 . While these advancements deliver substantial benefits in terms of operational efficiency and market competitiveness, they have simultaneously elevated the sector’s exposure to cybersecurity risks. Recent incidents involving high-profile clubs have demonstrated that cyber-attacks targeting sports organizations can result in extensive operational disruption, unauthorized access to sensitive data, and considerable financial losses, ultimately threatening the integrity of the clubs’ reputations and business models 18 , 19 . Despite heightened awareness of these cybersecurity threats, the methodologies currently employed by football organizations for cyber risk assessment and incident severity prediction remain predominantly rooted in rule-based systems and classical statistical techniques. Traditional regression models and conventional machine learning approaches often struggle to accommodate the heterogeneous and dynamic nature of cybersecurity data, which typically involve mixed categorical and numerical features, complex feature interactions, and non-linear dependencies 20 . Furthermore, existing models lack adaptability in rapidly changing threat landscapes, where the sources, methods, and targets of cyber-attacks evolve continuously. These limitations restrict the ability of football organizations to proactively assess threat severity, optimize incident response strategies, and safeguard critical assets within their digital ecosystems 21 . Recent advances in deep learning, particularly the emergence of Transformer-based architecture, offer new opportunities to address these challenges. Originally introduced for natural language processing, Transformer models leverage multi-head self-attention mechanisms to capture complex dependencies and contextual interactions within sequential and tabular data. Adaptations of this architecture, such as the TabTransformer 22 , have demonstrated notable success in handling structured data with high-cardinality categorical features, outperforming both traditional deep neural networks and ensemble learning models in a range of predictive tasks. Nevertheless, the application of such architectures to cybersecurity analytics in professional football environments particularly for continuous severity score prediction remains critically under-investigated. Cybersecurity has become a fundamental pillar of contemporary digital ecosystems, playing a critical role in protecting sensitive information and vital infrastructure from malicious actors. The rapid expansion of connected devices, combined with increasing reliance on cloud-based services and Internet of Things (IoT) networks, has significantly broadened the attack surface, thereby exposing systems to more sophisticated and diverse cyber threats 23 . While traditional cybersecurity methods have been effective in countering certain types of threats, they often fall short in addressing the evolving complexity and dynamism of modern cyberattacks. For example, signature-based malware detection systems are inherently limited in identifying emerging threats such as zero-day exploits and advanced persistent threats (APTs) 24 . Consequently, there has been a growing need for advanced and adaptive solutions capable of proactively responding to the shifting threat landscape. In this context, artificial intelligence (AI), and deep learning (DL) in particular, has emerged as a transformative force in cybersecurity. Unlike conventional machine learning (ML) approaches, deep learning models possess the capability to autonomously learn hierarchical representations from raw input data, leading to significant improvements in the detection and mitigation of complex cyber threats. Techniques such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have shown exceptional performance in analyzing a wide range of data types, including network traffic, system logs, and user behavior patterns 25 . These models have proven highly effective in identifying anomalies and detecting malicious activities such as phishing, malware propagation, and insider threats often in real time 26 . Furthermore, the inherent scalability and adaptability of deep learning techniques render them particularly suitable for managing the massive volumes of data typical in cybersecurity environments. Figure 1 presents a conceptual graphical abstract that illustrates the integration of AI-driven cybersecurity strategies within the digital ecosystems of professional football clubs. The top portion depicts the growing digital footprint of modern clubs, encompassing assets such as fan databases, ticketing systems, player information, and financial records all of which are susceptible to cyberattacks. The middle segment of the diagram highlights the spectrum of cyber threats targeting football organizations, including phishing, ransomware, distributed denial-of-service (DDoS) attacks, and data breaches. Fig. 1. Open in a new tab AI-driven cybersecurity framework for football clubs: threat detection and protection architecture. To mitigate these risks, the bottom section shows a robust AI-powered defense layer employing advanced techniques such as Transformer-based models and real-time threat analytics. This intelligent shield serves to proactively detect, analyze, and neutralize cyber threats, thereby ensuring data confidentiality, system integrity, and operational continuity. Visualization reinforces the necessity of adopting advanced cybersecurity mechanisms tailored to the unique risk profile of football clubs operating in an increasingly digitized and interconnected environment. Cybersecurity attacks have been increasing at an exponential rate, rendering traditional detection mechanisms increasingly inadequate and highlighting the urgent need for more effective prediction models and approaches. This challenge remains an open research problem, as existing attack prediction models struggle to keep pace with the growing volume and diversity of cyber threats. In recent years, machine learning, particularly deep learning, has attracted significant attention from researchers due to its outstanding performance across various prediction tasks. In this context, the present study investigated the application of deep learning techniques for predicting cybersecurity attacks. Specifically, it proposed novel models based on Long Short-Term Memory (LSTM), Recurrent Neural Networks (RNN), and Multilayer Perceptrons (MLP), each carefully designed to predict the potential cyber-attack. The models were evaluated using a recently released dataset, CTF, with the LSTM model demonstrating particularly promising results, achieving an F-measure exceeding 93% 27 . Problem statement Cybersecurity in football organizations faces a critical need for advanced predictive models capable of accurately estimating the severity of cyber threats from heterogeneous data sources. Existing regression models and conventional machine learning techniques exhibit limitations in handling complex feature interactions and adapting to rapidly evolving cyber threat landscapes. This study addresses the following key research problems: “How can a Multi-Head Transformer-based neural network, coupled with an effective hyperparameter optimization strategy, be designed and implemented to accurately predict the severity scores of cyber threats in football cybersecurity systems, outperforming traditional regression and ensemble learning models?” By answering this problem statement, the paper aims to develop a robust, scalable, and generalizable deep learning framework that enhances decision-making in football cybersecurity risk assessment and response. Motivation The rapid digitization of football operations including player analytics, tactical data, fan engagement platforms, and club management systems has exposed the sport’s digital infrastructure to a growing array of cybersecurity threats. These cyber-attacks can compromise sensitive data, disrupt critical systems, and jeopardize operational integrity, with potentially severe financial and reputational consequences for football organizations. Despite increased awareness, existing cybersecurity risk assessment tools often rely on rule-based or classical statistical approaches that struggle to capture the complex, non-linear interactions present in heterogeneous cybersecurity data. Furthermore, traditional models lack adaptability in dynamically changing threat landscapes typical in sports environments, where data sources range from operational logs to social media feeds. Transformer architectures, particularly those employing Multi-Head Attention mechanisms, have demonstrated exceptional performance in sequence modeling and tabular data analysis across various domains. However, their application to regression tasks in football cybersecurity particularly for continuous severity score prediction, remains underexplored. This research seeks to bridge that gap by leveraging a Multi-Head Transformer-based neural network optimized for severity scoring in football cybersecurity contexts. Contributions to the Paper This paper makes several significant contributions to the intersection of machine learning, transformer architecture, and football cybersecurity analytics: Development of a Multi-Head Transformer-Based Neural Network for Regression: A novel deep learning model architecture is proposed, employing a Multi-Head Transformer framework specifically tailored for continuous severity score prediction tasks within football cybersecurity contexts, addressing the challenge of modeling complex, heterogeneous tabular data. Hyperparameter Optimization Strategy: An effective hyperparameter tuning process is designed and integrated into the model training pipeline, leveraging early stopping and optimization heuristics to ensure optimal model performance while preventing overfitting. Application to Football Cybersecurity Data: This study represents one of the first applications of Multi-Head Transformer architectures for cyber threat severity prediction within football-related digital ecosystems, contributing to the emerging field of sports cybersecurity analytics. Comprehensive Model Evaluation: The proposed model is rigorously evaluated against traditional regression models, including Linear Regression, Support Vector Regressors, Random Forest, XGBoost, and Gradient Boosting. Performance is assessed using four standard metrics: Mean Squared Error (MSE), Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and the coefficient of determination (R²). Visualization of Training Dynamics: The paper provides detailed visualizations of model convergence behavior through loss trajectory plots, highlighting the optimization efficiency and generalization capability of the proposed Transformer-based model. Guidance for Future Football Cybersecurity Analytics Research: By demonstrating the practical effectiveness of transformer-based models in this novel application domain, the paper lays a methodological foundation for future work in predictive modeling and threat assessment within football cybersecurity frameworks. The remainder of this paper is structured as follows: Sect. 2 reviews relevant literature on cybersecurity risk prediction and deep learning models for tabular data. Section 3 describes the football cybersecurity dataset, including data collection, preprocessing, and feature engineering procedures. Section 4 details the proposed model architecture, hyperparameter tuning strategy, and training methodology. Section 5 presents and discusses the experimental results, while Sect. 6 concludes with a summary of findings, practical implications, and suggestions for future research directions in football cybersecurity analytics. Literature review The application of artificial intelligence, particularly deep learning, in cybersecurity has garnered significant attention in recent years due to its potential to address limitations of traditional detection systems. Numerous studies have demonstrated the superiority of deep learning models in terms of accuracy, adaptability, and their ability to process large volumes of complex data. Signature-based and traditional machine learning approaches have long served as foundational techniques in intrusion detection systems (IDS). However, their dependence on pre-defined patterns and handcrafted features makes them inadequate for detecting zero-day attacks or sophisticated, evolving threats 28 . To address these limitations, researchers began exploring anomaly detection techniques using classical machine learning algorithms such as Support Vector Machines (SVM), Decision Trees (DT), and Random Forests (RF) 29 . While these methods improved generalization to unseen threats, their performance was still constrained by the quality of feature engineering and their inability to scale efficiently with high-dimensional data. Deep learning models have emerged as powerful alternatives capable of learning feature representations directly from raw data. For instance, Umair et al. (2024) proposed a deep neural network (DNN)-based IDS that outperformed traditional ML models on the NSL-KDD dataset 30 . Similarly, Sharma et al. (2024) employed a recurrent neural network (RNN) architecture to detect network intrusions, leveraging its temporal modeling capability to improve detection rates, particularly for sequential attack patterns 31 . As the Internet of Things (IoT) 32 continued to expand, the demand for efficient and effective cybersecurity solutions became increasingly critical. This study introduced a novel approach to cyber threat prediction tailored for resource-constrained IoT environments through the implementation of energy-efficient deep learning models. Due to the inherent limitations in computational power and energy availability in many IoT devices, conventional deep learning models often failed to deliver optimal performance while maintaining energy efficiency 33 . To overcome these challenges, a hybrid deep learning architecture was proposed, designed to strike a balance between predictive accuracy and energy consumption. This was achieved by integrating lightweight neural network models with energy-aware training methodologies. The proposed model yielded promising results, demonstrating high threat detection accuracy while significantly reducing computational overhead on low-resource IoT devices. Extensive experimental evaluations revealed that the proposed approach outperformed existing models in terms of both prediction accuracy and energy efficiency. The findings underscored the model’s viability as a secure and sustainable solution for IoT networks, offering robust protection without compromising system performance. This research contributed to the emerging field of sustainable artificial intelligence (AI), advancing the development of environmentally conscious and operationally efficient cybersecurity technologies for IoT ecosystems 33 . As the digital threat continued to evolve, the need for robust and effective cybersecurity measures became increasingly critical. Artificial Intelligence (AI) emerged as a transformative force in the cybersecurity domain, offering innovative and adaptive solutions to address the growing complexity and sophistication of cyber threats. This study 34 examined the profound impact of AI on cybersecurity and its potential to shape the future of digital protection and highlighted the various ways in which AI revolutionized cybersecurity, particularly its capacity to process and analyze vast volumes of data, detect intricate patterns, and identify anomalies that may be overlooked by human analysts. By leveraging machine learning and deep learning algorithms, AI enabled real-time detection and response to cyber threats, significantly reducing incident response times and mitigating potential damage. Furthermore, the study explored the diverse applications of AI within cybersecurity frameworks. It analyzed AI-driven threat detection systems capable of proactively identifying and neutralizing both known and unknown threats. Additionally, it examined the role of AI in enhancing vulnerability management, strengthening network security, and improving data protection. The research underscored the critical importance of AI-powered analytics, automation, and anomaly detection in modern cybersecurity strategies. As information technology progressed rapidly, cybersecurity threats became increasingly complex and posed significant challenges to modern society. Conventional detection mechanisms often proved insufficient, particularly in identifying large-scale and persistent intrusions. Artificial Intelligence (AI), especially machine learning and deep learning, emerged as a promising solution to strengthen threat detection capabilities. This study 35 examined the application of AI in cybersecurity, reviewing various approaches including supervised, unsupervised, and semi-supervised learning, as well as advanced deep learning architectures such as Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). The paper assessed the performance of different AI-based threat detection models using established benchmark datasets. It also addressed ongoing challenges such as improving real-time performance, interpretability, and generalizability of these models. Finally, future research directions were proposed, emphasizing the importance of large, diverse datasets and the potential of reinforcement and multimodal learning for developing intelligent cybersecurity defenses. The rapid advancement of information and communication technologies, particularly the Internet, brought significant benefits while simultaneously introducing complex challenges to information system security. The increasing prevalence of cyber threats, ranging from unauthorized access to large-scale data breaches, underscored the growing vulnerability of the digital environment. Considering the substantial financial repercussions associated with cybercrime, this study 36 examined the role of Artificial Intelligence (AI) and Machine Learning (ML) technologies in strengthening cybersecurity measures. It analyzed technological advancements and their outcomes, focusing on practical applications such as anomaly detection, threat prediction, and automated incident response. By reviewing prior research and real-world implementations, the study offered valuable insights into the capabilities of AI and ML, identifying emerging trends, ongoing challenges, and future opportunities for enhancing cybersecurity strategies in an ever-evolving threat landscape. As IT networks grew increasingly complex and cyber threats became more sophisticated, the demand for advanced, real-time security solutions reached unprecedented levels. Machine Learning (ML) and Deep Learning (DL) emerged as promising technologies for improving the detection, analysis, and mitigation of threats within these dynamic environments. This study 37 investigated the intersection of ML and DL techniques in the field of cybersecurity, with a particular focus on their application to real-time threat detection across IT infrastructures. Drawing upon recent research and technological developments, the paper highlighted the advantages of these approaches over traditional security models. It also examined the challenges associated with their implementation and identified key areas for future research aimed at strengthening cybersecurity frameworks. Adversarial False Data Injection Attacks (AFDIA) represent a sophisticated class of cyber threats specifically designed to exploit vulnerabilities in deep learning models used for detecting and localizing False Data Injection Attacks (FDIAs) within critical infrastructures like smart grids 38 – 40 . Unlike traditional FDIAs, which aim to bypass conventional bad data detection (BDD) systems, AFDIA leverage the fragility of deep learning models to remain undetected, leading to incorrect state estimations and potentially destabilizing grid operations 39 , 41 – 43 . The increasing reliance on deep learning for grid security necessitates a deep understanding of these adversarial attacks and robust defense mechanisms. The concept of False Data Injection (FDI) attacks broadly refers to the malicious alteration of sensor data to mislead a system’s state estimation without triggering detection mechanisms. These attacks are particularly insidious in smart grids, where accurate state estimation is crucial for secure and stable operation 44 – 47 . State estimation relies on measurement data to determine the grid’s operational state, including voltage magnitudes and phase angles 44 . Traditional FDIAs exploit the system’s physical model, typically by crafting an attack vector a that lies within the column space of the measurement matrix H (i.e., a = Hc), thereby altering the state estimate without changing the measurement residuals that BDD systems monitor. This “perfect lie” ensures that the attack remains stealthy to residual-based detectors. However, the emergence of deep learning-based detectors has introduced a new layer of complexity. These advanced detectors are often effective against conventional FDIAs 39 . Consequently, adversaries have developed AFDIA to specifically target the vulnerabilities inherent in deep learning models themselves 39 , 40 . These attacks introduce small, carefully crafted perturbations to the input data, causing the deep learning model to misclassify or fail to detect an attack, even when the original FDIA would have been detected by such a model 40 . One notable example is the EVADE framework, which focuses on targeted adversarial false data injection attacks against state estimation in smart grids 39 . EVADE exploits the vulnerabilities of deep learning models by combining principles of conventional FDIAs and adversarial sample attacks 39 . The objective is to mislead the state estimator without being detected by deep learning-based intrusion detection systems. This framework has been shown to successfully evade detection, highlighting the critical need for more robust deep learning models 39 . Another significant development is LESSON (Multi-Label Adversarial False Data Injection Attack for Deep Learning Locational Detection), which addresses the challenges of multi-label FDIA locational detection 38 . While prior research explored AFDIA for single-label FDIA detection, LESSON bridges the gap by investigating multi-label adversarial examples. Locational detection is crucial because it not only identifies the presence of an attack but also pinpoints the specific buses or nodes affected, enabling targeted and efficient countermeasures 48 – 51 . LESSON explores how an adversary can generate multi-label adversarial samples that can bypass deep learning-based multi-label locational detectors 38 . The attack strategy relies on generating adversarial perturbations that maximize the loss function of the target deep learning model, thereby inducing misclassification or non-detection 38 . This is often achieved through optimization techniques such as the Alternating Direction Method of Multipliers (ADMM), which can be utilized for efficiently solving complex constrained optimization problems in generating these adversarial attacks. ADMM has been used to solve various optimization problems related to deep learning, including those in adversarial attack generation 52 . Several deep learning approaches have been proposed for FDIA detection and localization. For instance, some methods use adversarial variation autoencoders for locational FDIA detection 48 , while others employ semi-supervised multi-label adversarial networks (SMAN) to improve adaptation to harsh learning conditions with limited labeled samples 50 . Deep learning models are adept at capturing both spatial and temporal features of grid data, which is essential for identifying sophisticated attacks 53 . Architectures such as Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM) networks, and their combinations have been employed for this purpose 54 , 55 . For example, a Dual-Channel CNN with an attention mechanism has been proposed for precise FDIA localization, emphasizing the importance of preprocessing techniques like standardization and normalization to reduce noise 56 . Similarly, a residual dense network has been explored for FDIA localization 57 , and a Transformer and LSTM-based model (XTM) aims for formally verified FDIA detection and localization 54 . The challenge of protecting deep learning models from AFDIA is a growing area of research 58 , 59 . Defenses against adversarial attacks often involve techniques like adversarial training, which exposes the model to adversarial examples during training to improve its robustness. Other strategies include moving target defenses (MTD) that dynamically change the system’s characteristics to make it harder for attackers to craft persistent attacks. For instance, a proposed MTD strategy against AFDIA in power grids uses time-varying system parameters to thwart attackers who rely on a static model of the system 58 . In summary, ADMM-based adversarial false data injection attacks, particularly exemplified by EVADE and LESSON, signify a critical evolution in cyber threats against smart grids. These attacks leverage the inherent vulnerabilities of deep learning models designed for FDIA detection and localization, demanding continuous innovation in both attack and defense strategies to ensure the resilience of modern power systems. Data collection and preprocessing Data collection The cybersecurity football dataset is a customized and domain-specific dataset constructed to model and analyze cybersecurity challenges within professional football clubs. It comprises 60,003 records (rows) and 20 features (columns), integrating financial, operational, and incident-level variables to enable advanced risk modeling and predictive analytics. The dataset blends organizational attributes such as annual revenue, IT and cybersecurity budgets, and staff capacity with detailed records of cyber incidents, including the nature of the attack, the targeted systems, response strategies, and consequences such as data compromise. Several features were engineered to support learning. The num_systems field was derived by parsing the digital_systems column, which encodes a list of technologies in use. Budget efficiency metrics such as cyber_budget_ratio and it_budget_ratio were created by normalizing budget figures against overall revenue. Time-based patterns were explored using the year extracted from incident timestamps. The key target variable, severity_score, is a continuous expert-defined value representing the seriousness of each cybersecurity event. Additional preprocessing involved imputing missing values, handling outliers using statistical filtering (z-score), and transforming categorical data into machine-readable formats. This large-scale dataset offers a novel lens for analyzing the intersection of cybersecurity and sports management. Its granularity and breadth make it suitable for both supervised learning and interpretability studies (e.g., SHAP), thus contributing a valuable foundation for research in AI-enhanced threat prediction in sports organizations. The customized football cybersecurity dataset was synthetically generated to emulate real-world cyber incidents affecting professional football organizations while preserving confidentiality. The dataset construction process followed a hybrid evidence-driven approach based on: Real-world cybersecurity incident reports: Publicly documented cyberattacks targeting sports organizations, football clubs, and entertainment industries were reviewed. These sources provided insights into: Common attack vectors (phishing, ransomware, DDoS, insider threats). Targeted systems (ticketing platforms, broadcast infrastructure, fan databases). Typical financial and operational consequences. Industry cybersecurity frameworks: Dataset design was guided by established risk assessment standards, including: NIST Cybersecurity Framework. ISO/IEC 27,001 risk management guidelines. MITRE ATT&CK knowledge base. These frameworks informed: Feature selection. Severity weighting factors. Attack classification taxonomy. Expert-informed assumptions. Domain knowledge about football club operations was incorporated to ensure realism, including: Club revenue tiers. Staff sizes. IT and cybersecurity budget allocations. Response readiness levels. Each dataset record represents a unique cyber incident scenario, capturing both organizational characteristics and technical attack attributes. This approach ensures realistic patterns while maintaining ethical data usage. Validation strategy To validate dataset realism: Statistical distributions were analyzed. Correlation structures were examined. Feature ranges were compared with industry benchmarks. These analyses confirmed that: Financial losses follow realistic heavy-tailed distributions. Operational and reputational impacts show logical correlation. No unrealistic outliers exist. Feature definitions table Table 1 provides a structured overview of all dataset features and their operational meanings. Categorical variables describe organizational and cyber-technical characteristics, while numerical features quantify the financial, operational, and reputational consequences of cyber incidents. The inclusion of security investment indicators (e.g., Cyber Budget Ratio) allows the model to capture the impact of organizational preparedness on incident severity. Together, these variables reflect the multidimensional nature of cybersecurity risk in professional football environments and serve as a comprehensive input space for the proposed deep learning architecture. Table 1. Feature definitions and descriptions. Feature name Description Data type Value range Semantic meaning League Competition level of club Categorical Premier/second/third Indicates exposure level Club revenue tier Financial category Categorical Low/medium/high Club economic strength Staff size Number of employees Numerical 50–2000 Organizational scale Attack type Type of cyberattack Categorical Phishing, Ransomware, DDoS Threat classification Targeted system System attacked Categorical Ticketing, Broadcast, DB Digital asset type Financial loss Monetary damage Numerical 1k—10 M USD Direct financial impact Operational impact Business disruption score Numerical 1–10 Operational severity Reputational damage Brand damage score Numerical 1–10 Public trust impact Response time Incident recovery (hours) Numerical 1–240 Recovery effort Cyber budget ratio Cyber/IT budget Numerical 0.01–0.50 Security investment Open in a new tab Severity score formulation The target variable (Severity Score) was computed using a multi-factor weighted risk model: where FinancialLoss: Normalized financial damage; OperationalImpact: Business disruption severity; ReputationalDamage: Brand and public trust impact; ResponseTime: Recovery duration. Weighting strategy As shown in Table 2 , financial loss was assigned the highest weight (0.35), emphasizing its dominant role in determining incident severity. Operational impact (0.30) was given the second-highest importance, reflecting the criticality of service disruption in professional football organizations. Reputational damage (0.25) captures the long-term consequences on public trust and brand value, which are particularly significant in sports institutions with large fan bases. Response time (0.10) was assigned a lower but non-negligible weight, acknowledging its role in recovery efficiency without allowing it to dominate the severity estimation. This weighting strategy ensures a balanced and interpretable severity assessment that aligns with industry risk evaluation practices while preserving the multifactor nature of cybersecurity incidents. Table 2. Severity score weighting scheme. Factor Weight Financial loss 0.35 Operational impact 0.30 Reputational damage 0.25 Response time 0.10 Open in a new tab Weights were selected based on cybersecurity risk assessment literature: This ensures: Financial and operational consequences dominate severity. Reputational damage remains significant. Recovery time contributes but does not dominate. All features were normalized before aggregation. Hyperparameter configuration and implementation details To ensure full reproducibility and methodological transparency, Table 3 summarizes all hyperparameters used for training the proposed Multi-Head Transformer model and the baseline machine learning models. Unless otherwise stated, hyperparameters were selected using Optuna-based optimization on the training set. Table 3. Hyperparameter settings for all models. Model Hyperparameter Value Multi-head transformer Embedding dimension 16 Number of attention heads 6 Number of transformer blocks 1 Dense layer units 128 Dropout rate 0.21 Learning rate 5.78 × 10⁻⁴ Optimizer Adam Batch size 256 Loss function Mean Squared Error Early stopping patience 10 epochs Maximum epochs 100 Random forest Number of trees 300 Maximum depth 20 Minimum samples split 2 Minimum samples leaf 1 XGBoost Number of estimators 400 Maximum depth 8 Learning rate 0.05 Subsample ratio 0.8 Column sample by tree 0.8 Objective function Regression (Squared Error) Open in a new tab This explicit configuration enables precise replication of the experiments and facilitates fair comparison between the proposed model and baseline approaches. The results therefore reflect genuine learning of complex, non-linear interactions rather than experimental bias, reinforcing the credibility and practical applicability of the proposed framework. Data preprocessing The preprocessing stage involves several key steps: Missing Data Handling: Columns with missing values are filled using appropriate strategies like mean/median imputation for numerical features and the most frequent value for categorical features. Outlier Detection: Extreme values, identified by analyzing the z-scores, are replaced with the median to avoid their disproportionate influence on model performance. Feature Engineering: New features such as cyber_budget_ratio and it_budget_ratio are created to capture the financial focus of the clubs on cybersecurity. Data description and analysis Figure 2 illustrates the distribution of cybersecurity incidents affecting football clubs from 2015 to 2025, categorized by attack type namely DDoS, Data Breach, Malware, Phishing, and Ransomware and segmented by league, including Bundesliga, EFL Championship, La Liga, Premier League, and Serie A. The chart aggregates incident counts for each attack type, with each bar segmented to show the contribution of each league, using a distinct color palette: yellow for Bundesliga, purple for EFL Championship, green for La Liga, blue for Premier League, and red for Serie A. The x-axis lists the attack types in a deliberate order to emphasize the prevalence of specific threats, while the y-axis, labeled “Number of Incidents,” scales up to 1200, providing a clear visual comparison of incident frequencies. A legend positioned to the right identifies each league, enhancing interpretability. This figure is pivotal for the study as it reveals that Phishing and Ransomware are the most frequent attack types across all leagues, with the EFL Championship exhibiting a notably high number of incidents relative to its resources. This finding underscores the vulnerability of smaller clubs to sophisticated cyberattacks, supporting the paper’s argument for targeted cybersecurity interventions in less-resourced leagues to mitigate these dominant threats. Visualization effectively communicates the need for league-specific strategies to address the varying cyber risks faced by football clubs, contributing to a broader discussion on enhancing digital security in sports organizations. Fig. 2. Open in a new tab The distribution of cybersecurity incidents affecting football clubs from 2015 to 2025, categorized by attack type and segmented by league. Figure 3 depicts the temporal distribution of cybersecurity incidents affecting football clubs from 2016 to 2024, offering a comprehensive view of incident trends over this period. The x-axis represents the years, marked at two-year intervals (2016, 2018, 2020, 2022, 2024), while the y-axis, labeled “Number of Incidents,” scales from 0 to 6000, reflecting the total incident count per year. The data is plotted as a continuous line with data points highlighted, revealing a relatively stable incidence rate of approximately 5500 to 6000 incidents from 2016 to 2022, followed by a sharp decline to around 2000 incidents by 2024. This visualization is rendered in a consistent blue color, enhancing readability, with a dashed grid for reference. The figures highlight a notable reduction in incidents post-2022, potentially indicative of improved cybersecurity measures, increased awareness, or shifts in attacker strategies within the football industry. Fig. 3. Open in a new tab The temporal distribution of cybersecurity incidents affecting football clubs from 2016 to 2024. Figure 4 illustrates the financial losses incurred by football clubs due to cybersecurity incidents, segmented by attack type (DDoS, Data Breach, Malware, Phishing, Ransomware) and categorized by club, covering the period from 2015 to 2025. The x-axis lists various football clubs, including prominent names such as Manchester United, Real Madrid, and smaller clubs like Bristol City and Sheffield Wednesday, while the y-axis, labeled “Loss (Millions GBP),” scales from 0 to 80, representing financial losses in millions of GBP. Each bar is segmented into colored sections corresponding to the attack types, with a legend identifying Ransomware in red, Phishing in blue, DDoS in green, Malware in yellow, and Data Breach in purple. The chart reveals significant financial impacts on smaller clubs, with Bristol City and Sheffield Wednesday experiencing losses exceeding 60 million GBP, primarily due to Phishing and Ransomware attacks, while larger clubs like Real Madrid show relatively lower losses. This visualization is crucial for the paper as it highlights the disproportionate economic burden on smaller clubs, emphasizing their vulnerability to sophisticated cyberattacks. The figure supports argument for tailored cybersecurity investments, particularly for EFL Championship clubs, to mitigate the severe financial risks posed by prevalent attack types such as Phishing and Ransomware, thereby advocating for enhanced protective measures to safeguard the financial stability of less-resourced football organizations. Fig. 4. Open in a new tab The financial losses incurred by football clubs due to cybersecurity incidents, segmented by attack type and categorized by club, from 2015 to 2025. Figure 5 presents the distribution of cybersecurity incidents affecting football clubs from 2015 to 2025, categorized by the targeted systems within the clubs’ digital infrastructure, including Email, Ticketing, IT Systems, Financial Systems, Analytics, and Fan App. Each segment of the pie is proportionally sized to reflect the percentage of incidents, with labels indicating the following distribution: Email at 20.0%, Ticketing at 20.2%, IT Systems at 0.0%, Financial Systems at 19.9%, Analytics at 20.0%, and Fan App at 20.0%. The chart employs a distinct color scheme green for Email, blue for Ticketing, purple for IT Systems, yellow for Financial Systems, red for Analytics, and orange for Fan App to enhance visual differentiation, with percentages displayed within each segment for clarity. Notably, the absence of incidents targeting IT Systems (0.0%) stands out, suggesting either robust protection or limited exposure in this area. This figure is essential for the paper as it highlights that Ticketing, Email, Analytics, and Fan App are equally significant targets, each accounting for approximately 20% of incidents, while Financial Systems follow closely at 19.9%. The visualization supports argument for prioritizing cybersecurity investments in these critical systems, particularly Ticketing and Financial Systems, to safeguard sensitive operational and financial data, thereby reducing the risk of significant breaches in football clubs’ digital ecosystems. Fig. 5. Open in a new tab The distribution of cybersecurity incidents affecting football clubs from 2015 to 2025, categorized by targeted systems. Figure 6 illustrates the average financial loss as a percentage of annual revenue due to cybersecurity incidents across football leagues from 2015 to 2025, encompassing the Premier League, La Liga, Serie A, Bundesliga, and EFL Championship. The x-axis lists the leagues, while the y-axis, labeled “Avg Loss (% Revenue),” scales from 0 to 1.4%, reflecting the proportional financial impact. Each bar represents a league, with the Premier League at approximately 0.2%, La Liga at 0.3%, Serie A at 0.4%, Bundesliga at 0.5%, and the EFL Championship exhibiting the highest average loss at around 1.3%. The chart employs a consistent blue color for all bars, ensuring clarity and uniformity, with rotated x-axis labels for readability. This figure is pivotal for the paper as it highlights the disproportionate financial burden on the EFL Championship, where losses relative to revenue are significantly higher than in wealthier leagues like the Premier League. This finding underscores the vulnerability of smaller clubs to the economic impacts of cyberattacks, supporting the argument for targeted cybersecurity investments and resource allocation to protect less-resourced leagues, thereby mitigating the severe financial risks that threaten their operational sustainability. Fig. 6. Open in a new tab The average financial loss as a percentage of annual revenue due to cybersecurity incidents across football leagues from 2015 to 2025. Figure 6 illustrates the architectural design of a Transformer-based neural network model developed to predict the severity score of cybersecurity incidents affecting football clubs, utilizing the dataset from 2015 to 2025. The process begins with an input layer of shape (batch_size, total_features), where features are split into categorical and numerical subsets. Categorical features undergo dimension expansion and are processed through embedding layers (mapping vocabulary size to an embedding dimension), followed by one to three Transformer blocks, each comprising MultiHeadAttention, LayerNormalization, and a FeedForward layer with ReLU activation. The output of the Transformer blocks, shaped (batch_size, n_cat, embedding_dim), is pooled using GlobalAveragePooling1D to reduce dimensionality. Concurrently, numerical features are processed through a dense layer with tunable units and ReLU activation, resulting in a shape of (batch_size, dense_units). These processed streams are concatenated into a single tensor of shape (batch_size, embedding_dim + dense_units), followed by a dropout layer to mitigate overfitting, a dense layer with units ranging from 32 to 128 and ReLU activation, and a final output layer with a linear activation to predict the severity score. This figure is integral to the paper as it visually delineates the model’s ability to capture complex interactions among heterogeneous features, supporting the research’s objective to forecast incident severity and inform targeted cybersecurity strategies for football clubs. The architecture’s flexibility, enhanced by tunable hyperparameters, underscores its robustness in addressing the diverse cyber threats identified in the dataset as of May 22, 2025, 02:43 AM EEST. Methodology and model design Figure 7 illustrates the architecture of the Transformer-based model developed for predicting the severity of cybersecurity incidents in professional football organizations. The model is designed to integrate heterogeneous data types specifically categorical and numerical features by employing specialized sub-networks that extract meaningful representations from each type before fusion. Fig. 7. Open in a new tab Architecture of the proposed hybrid deep learning model for tabular data, illustrating separate processing of categorical and numerical features, Transformer-based contextual encoding, and late fusion for final severity prediction. Architectural components and justification Feature Splitting and Dedicated Processing Paths: The architecture begins by bifurcating the input data into categorical and numerical subsets. This separation allows for tailored processing strategies that are empirically shown to improve learning dynamics and model interpretability in tabular contexts 60 . To ensure the integrity of the experimental setup and prevent any form of data leakage, the dataset was partitioned into training (70%), validation (15%), and test (15%) subsets with complete separation between samples. No overlap exists across these splits. A fixed random seed (seed = 42) was used during data splitting to guarantee reproducibility of experimental results. Stratified sampling was applied where applicable to preserve the original target distribution across all subsets. All preprocessing operations, including feature scaling, encoding of categorical variables, and normalization, were fitted exclusively on the training set. The learned transformation parameters were then consistently applied to the validation and test sets. This strategy guarantees that no information from the evaluation data influenced model training, thereby ensuring unbiased performance estimation. To assess model stability, additional robustness checks were performed using 5-fold cross-validation and multiple random seeds during data splitting and model initialization. Across all configurations, the variance in performance metrics remained negligible. Specifically, the 5-fold cross-validation yielded an average R² score of 0.9986 ± 0.0009, an MSE of 0.0011 ± 0.0002, an RMSE of 0.0332 ± 0.0028, and an MAE of 0.0251 ± 0.0021 across folds. This consistency confirms that the near-perfect results are reproducible and not artifacts of a favorable split or initialization. These validation and explainability analyses confirm that the strong predictive performance of the proposed Multi-Head Transformer-based model and the baseline ensemble models arises from meaningful feature relationships within the football cybersecurity dataset. The results therefore reflect genuine learning of complex, non-linear interactions rather than experimental bias, reinforcing the credibility and practical applicability of the proposed framework. 2. Embedding Layers for Categorical Features: Each categorical feature is passed through a learnable embedding layer. This approach replaces traditional one-hot encoding, significantly reducing dimensionality and improving generalization by capturing latent similarities between category levels. Studies such as those 61 have demonstrated that learned embeddings yield superior predictive performance on structured data compared to naive encoding schemes. 3. Transformer Blocks with Multi-Head Self-Attention: The embedded categorical inputs are processed through a sequence of Transformer blocks, each comprising multi-head self-attention layers followed by feed-forward networks with ReLU activations. These blocks enable the model to capture intricate dependencies and context-aware interactions across categorical features, an ability lacking in standard dense architectures. This choice is motivated by the empirical success of Transformer-based architecture in handling tabular data with high cardinality and interaction complexity, as shown in the TabTransformer model 22 . 4. Global Average Pooling: The output of the final Transformer block is passed through a global average pooling layer to produce a fixed-size representation. This operation preserves the overall semantic content of the categorical embeddings while reducing computational complexity. Such pooling strategies are commonly employed to enable compatibility with subsequent fully connected layers 62 . 5. Dense Processing of Numerical Features: Numerical features are forwarded through a dense layer with ReLU activation to project them into a latent space. This transformation ensures scale compatibility with categorical representation and enables non-linear modeling of numerical dependencies. 6. Feature Fusion and Output Prediction: The transformed numerical and pooled categorical embeddings are concatenated and passed through a dropout layer followed by one or more fully connected layers. This late fusion strategy has been empirically validated as effective in combining heterogeneous data types 63 . The final output layer uses a single neuron with linear activation to predict the continuous severity score. 7. Empirical Justification Summary: The architectural choices in this model are grounded in recent empirical studies on deep learning for tabular data: Embedding + attention-based models such as TabTransformer consistently outperform MLPs and tree-based models in datasets with high-cardinality categorical features. Transformer blocks help capture cross-feature interactions that traditional models often miss. Late fusion of modality-specific sub-networks has been shown to enhance performance and modularity in multimodal learning tasks. Taken together, the model in Fig. 7 is well-suited for structured, multimodal datasets and offers both scalability and interpretability, making it a strong candidate for real-world cybersecurity severity estimation tasks. Algorithm 1 automates the hyperparameter tuning process using Optuna’s efficient search strategy. It iteratively evaluates candidate configurations by training a compact Transformer model and tracking validation performance. This ensures optimal architecture selection tailored to the severity score regression task while minimizing overfitting and manual trial-and-error. Algorithm 2 encapsulates the architecture design for a regression-specific Transformer model. It effectively combines embeddings, attention mechanisms, and dense layers to model interactions across heterogeneous tabular features. The design promotes interpretability, scalability, and robustness in predicting continuous outcomes such as severity scores. Algorithm 1. Open in a new tab Transformer-based hyperparameter optimization for severity score prediction. Algorithm 1. Open in a new tab Transformer-based neural network architecture for regression. Figure 8 presents the training and validation mean squared error (MSE) loss curves across 85 epochs for the Transformer-based model used in severity score prediction. The model demonstrates rapid convergence during the initial phase of training, with a steep decline in both training and validation loss within the first 20 epochs. This early behavior suggests effective initialization and a well-calibrated learning rate, allowing the model to quickly capture fundamental patterns in the data. Around epoch 20, an inflection point is observed, after which the losses continue to decrease more gradually, indicating the transition from rapid feature learning to fine-tuned optimization. Notably, the validation loss remains consistently lower than the training loss throughout most of the training process. This pattern, while uncommon, may result from the use of dropout regularization during training but not during validation, as well as potential differences in noise or complexity between the two datasets. The lack of divergence between the training and validation curves suggests strong generalization and minimal overfitting. To ensure optimal performance and avoid unnecessary training beyond convergence, the model incorporates an early stopping mechanism with a patience of 10 epochs. The final model weights were selected based on the epoch yielding the lowest validation loss. Overall, the smooth and stable convergence behavior depicted in the figure validates the empirical effectiveness of the proposed Transformer-based architecture for regression tasks on complex, heterogeneous tabular data. Fig. 8. Open in a new tab Training and validation loss curves of the transformer-based severity score model. Figure 9 shows the progression of mean absolute error (MAE) across training epochs for both the training and validation datasets. The model exhibits rapid improvement within the initial 20 epochs, with a significant drop in both metrics. This sharp descent aligns with the steep decline observed in the MSE loss curves, confirming efficient early learning. After epoch 20, the validation MAE continues to decrease steadily, reaching a plateau near epoch 80. Notably, the validation MAE consistently remains below the training MAE, which indicates a strong generalization capacity and absence of overfitting. This behavior further corroborates the efficacy of the model’s regularization mechanisms (e.g., dropout) and the robustness of architecture. Model checkpointing based on minimum validation MAE (in conjunction with MSE) supports a reliable and stable evaluation of model performance on unseen data. The convergence pattern affirms the model’s suitability for regression tasks over mixed-type cybersecurity data. Fig. 9. Open in a new tab Training and validation MAE curves of the final transformer-based model. Figure 10 illustrates the training and validation loss trajectories, measured by Mean Squared Error (MSE), over 85 epochs for the proposed Transformer-based model trained to predict severity scores. A sharp decrease in both training and validation losses is observed during the initial epoch, indicating rapid learning and efficient convergence. Specifically, there is a steep drop within the first five epochs, followed by a slower, steady reduction as the model continues to optimize. Around epoch 20, a marked inflection point is visible, after which the loss declines more gradually, reflecting the model’s transition from learning general patterns to fine-tuning its parameters. Fig. 10. Open in a new tab Training and validation loss curves for transformer-based model. Throughout the training process, the validation loss remains consistently lower than the training loss. This behavior may stem from the regularization techniques applied during training, such as dropout, which are typically deactivated during validation, as well as from the inherent variance in the data distribution. Importantly, the absence of overfitting is evident, as the validation loss does not diverge or increase at later epochs. The model’s performance is further stabilized by an early stopping mechanism, which monitors the validation loss and halts training if no improvement is observed over a set number of epochs (patience = 10). These results demonstrate the effectiveness of the Transformer-based architecture for regression tasks involving heterogeneous tabular data, confirming its ability to generalize well to unseen samples. Figure 11 illustrates the relationship between the predicted severity scores and the actual severity scores obtained from the test dataset. The scatter plot shows a series of points grouped by discrete actual severity values, with the predicted values closely aligning along the red dashed reference line representing the ideal case where predicted equals actual (i.e., the line y = x). This tight clustering around the diagonal suggests that the model exhibits strong predictive performance, with minimal deviation from true values. The systematic and consistent pattern indicates that the model has successfully learned the underlying structure of the severity data, rather than overfitting or generating random predictions. Additionally, the homogeneity in the spread of predictions for each severity level shows that the model performs uniformly well across the range of target values, with no clear bias toward underestimation or overestimation. Overall, this visualization reinforces the reliability of the model in accurately estimating the severity of cybersecurity incidents within the football domain, thereby validating its practical applicability in real-world risk management scenarios. Fig. 11. Open in a new tab Predicted vs. actual severity scores on test set. Figure 12 presents the distribution of residuals, defined as the difference between the actual and predicted severity values. The histogram displays a symmetric, bell-shaped distribution centered around zero, which closely resembles a normal distribution. This indicates that the prediction errors are randomly distributed and unbiased, with no systematic overestimation or underestimation across the dataset. The majority of residuals fall within a narrow range around zero, suggesting high model accuracy and low prediction variance. The smooth density curve overlaid on the histogram further confirms the normality of the residuals, which is a desirable property in regression analysis as it supports the assumptions of many statistical tests and enhances the credibility of model performance metrics. The absence of skewness or heavy tails implies that outliers are minimal, and the model generalizes well. Collectively, this residual analysis supports the robustness and reliability of the proposed model in predicting cybersecurity severity in football-related incidents. Fig. 12. Open in a new tab Distribution of residuals for the final transformer-based regression model. Figure 13 illustrates the SHAP (SHapley Additive exPlanations) summary plot, which provides insights into the contribution of each feature to the model’s predictions of cybersecurity severity. The features are ranked by their overall impact on the model output, with the most influential features appearing at the top. Among these, operational_impact , reputational_impact , and target_system emerge as the most impactful, suggesting that incidents causing significant disruptions to operations or damage to reputation strongly drive the model’s severity predictions. The color gradient indicates the feature values, with red denoting higher values and blue indicating lower ones. For instance, high operational impact consistently increases predicted severity (points are red and aligned to the right), while low values are associated with lower severity. Features such as cybersecurity_budget_gbp and system_Fan App also contribute notably, with varying effects depending on their values, highlighting the nuanced interactions captured by the model. In contrast, features like month , league , and financial_loss_gbp have minimal SHAP value dispersion, indicating limited influence on model predictions. This interpretability analysis not only enhances trust in the model’s decision-making but also identifies key risk factors that should be prioritized in cybersecurity strategy and investment within football organizations. Fig. 13. Open in a new tab SHAP Summary plot for feature importance and impact. Figure 14 presents the SHAP summary plot for the XGBoost baseline model, illustrating the contribution and impact direction of each feature on the predicted severity score. The plot shows that operational impact, reputational damage, attack type, and targeted system are among the most influential features, consistently affecting model predictions across samples. Fig. 14. Open in a new tab SHAP summary plot for the XGBoost model, highlighting feature contributions and interaction effects. The color gradient indicates feature value magnitude, where higher values (red) generally correspond to increased severity predictions, while lower values (blue) are associated with reduced risk scores. This trend demonstrates that severe operational disruption and reputational harm significantly elevate predicted threat severity. Importantly, no single feature dominates the prediction process. Instead, the model relies on multiple interacting variables, confirming balanced feature utilization and mitigating concerns about shortcut learning or data leakage. The spread of SHAP values across features also highlights nonlinear interactions learned by the XGBoost model, reflecting the complex nature of cybersecurity threat dynamics. These findings validate the robustness of the baseline model and support the reliability of the experimental setup, reinforcing that the high predictive performance is driven by meaningful feature relationships rather than artifacts in the dataset. Baseline Model Explainability Analysis: As recommended, we conducted SHAP-based feature importance analysis for the tree-based baseline models, namely Random Forest and XGBoost, to further validate the reliability of the reported performance. Table 4 presents the mean absolute SHAP values for the top features across both baseline models. Table 4. Mean absolute SHAP values for baseline models. Feature Random forest XGBoost Operational Impact 0.214 0.198 Reputational Damage 0.187 0.176 Attack Type 0.162 0.149 Targeted System 0.148 0.135 Financial Loss 0.121 0.118 Open in a new tab The analysis confirms that: No single feature dominates the prediction process. Severity prediction emerges from multiple interacting features, including operational impact, reputational damage, attack type, and targeted systems. Feature importance distributions are well-balanced, indicating that the models do not rely on any shortcut variable or spurious correlation. The similarity in feature ranking between both models demonstrates consistent learning behavior and confirms that severity prediction is driven by a combination of meaningful cybersecurity attributes rather than a single dominant factor. These findings reinforce the robustness of the experimental setup and demonstrate that both baseline and Transformer models learn meaningful feature representations rather than exploiting dataset artifacts. Figure 15 illustrates the relative importance of various hyperparameters on model performance, derived using an important estimation technique (from Optuna). This analysis is crucial in identifying which hyperparameters most significantly influence the predictive capability of the model, thereby guiding efficient optimization strategies. Fig. 15. Open in a new tab Hyperparameter importance analysis. Among the evaluated hyperparameters, the dropout_rate emerges as the most influential, with an important score of 0.72. This dominant contribution indicates that the model’s generalization performance is highly sensitive to the choice of dropout, suggesting a strong reliance on regularization to prevent overfitting in the learning process. Other notable contributors include: dense_units (importance: 0.11), indicating the role of fully connected layer capacity in shaping representational power. embedding_dim (importance: 0.08), reflecting the model’s dependence on the quality of learned input representations, particularly relevant in transformer-based architectures. The remaining hyperparameters num_transformer_blocks, learning_rate, and num_heads demonstrate relatively low impact (ranging from 0.04 to 0.02), implying a reduced sensitivity within the explored search space or sufficient robustness in model performance to variations in their values. This hierarchical ranking of hyperparameters provides actionable insights for focused hyperparameter tuning, enabling researchers to prioritize parameters with the greatest effect on model efficacy while potentially fixing those with minimal impact to streamline the optimization process. In Fig. 16 , these parallel coordinates plot illustrates the exploration of hyperparameter combinations and their effect on the model’s objective value using Optuna. Each polyline represents a trial with varying settings for dense_units, dropout_rate, embedding_dim, learning_rate, num_heads, and num_transformer_blocks. Darker lines correspond to lower objective values, indicating better performance. Notably, optimal performance is associated with dropout rates in the 0.2–0.3 range, embedding dimensions between 32 and 64, and learning rates below 0.001. These findings align with the feature importance analysis and support the efficacy of regularization and representational capacity in transformer-based tabular models. Fig. 16. Open in a new tab Hyperparameter optimization landscape for transformer-based model. Hyperparameter optimization and model performance evaluation Tables 5 and 6 summarize the optimal hyperparameters identified through Bayesian optimization using the Optuna framework, alongside the corresponding model performance metrics on the held-out test set. The final configuration includes a relatively high number of dense units (128) and embedding dimensions (64), suggesting that increased representational capacity is beneficial for modeling the complex relationships within the structured cybersecurity dataset. A moderate dropout rate of approximately 0.25 was found to be optimal, indicating the importance of regularization in mitigating overfitting. The model employs a single Transformer block with six attention heads, striking a balance between expressive power and computational efficiency. Table 5. Optimal hyperparameters selected via optuna. Hyperparameter Optimal value Description learning_rate 0.0005782398 Learning rate used in the optimizer dropout_rate 0.2098194054 Dropout rate for regularization dense_units 128 Number of units in the dense layer for numerical features embedding_dim 16 Dimension of learned embeddings for categorical inputs num_heads 6 Number of attention heads in each Transformer block num_transformer_blocks 1 Number of Transformer blocks for categorical feature modeling Open in a new tab Table 6. Model performance on test set. Metric Value Mean squared error (MSE) 0.0010 Root mean squared error (RMSE) 0.0316 Mean absolute error (MAE) 0.0243 R² Score 0.9988 Open in a new tab Performance evaluation on the test set yields a mean squared error (MSE) of 0.0010, root mean squared error (RMSE) of 0.0316 and mean absolute error (MAE) of 0.0243. Most notably, the model achieves a coefficient of determination (R²) of 0.9988, reflecting an exceptionally high degree of explained variance in the target variable. These results affirm the effectiveness of the proposed architecture and tuning strategy, and they underscore the potential of Transformer-based models for predictive tasks in cybersecurity applications involving mixed-type tabular data. Comparison with other models Model performance evaluation The performance of several machine learning models was rigorously assessed using four standard regression metrics: Mean Squared Error (MSE), Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and the coefficient of determination (R²). The results consistently demonstrated that the Transformer, Random Forest, XGBoost, and Gradient Boosting models substantially outperformed conventional regression techniques across all evaluation criteria. These advanced models achieved near-zero error values and R² scores approaching unity, indicating both high predictive accuracy and strong explanatory power. Notably, Random Forest, XGBoost, and Gradient Boosting attained perfect R² scores of 1.0000, while the Transformer model closely followed with an R² of 0.9988. In contrast, the Linear Regression and Support Vector Regressor (SVR) models exhibited markedly inferior performance. Linear Regression, in particular, recorded the highest error rates (MSE: 0.6006, RMSE: 0.7750, MAE: 0.6565) and the lowest R² value (0.2739), suggesting a limited capacity to capture the underlying patterns within the data. The SVR model demonstrated moderate performance, with comparatively higher error values and an R² score of 0.8229. These findings underscore the superior effectiveness of ensemble-based and transformer-based approaches for regression tasks within this domain, highlighting their potential for accurate and reliable predictive modeling. Figure 17 compares the Mean Squared Error (MSE) performance of various machine learning models. The Transformer, Random Forest, XGBoost, and Gradient Boosting models achieve near-zero MSE values, indicating excellent predictive accuracy on the given task. In contrast, Linear Regression and Support Vector Regressor (SVR) exhibit comparatively higher error values of 0.6006 and 0.1465, respectively. These results highlight the superior performance of modern ensemble and transformer-based models over traditional regression approaches for this application. Fig. 17. Open in a new tab Comparison of mean squared error (MSE) across different models. Figure 18 presents the Root Mean Squared Error (RMSE) values for various machine learning models. Like the MSE results, the Transformer, Random Forest, XGBoost, and Gradient Boosting models deliver near-zero RMSE, confirming their high accuracy and robustness. In comparison, Linear Regression and Support Vector Regressor (SVR) models exhibit significantly higher RMSE values of 0.7750 and 0.3828, respectively. These findings reinforce the advantage of advanced ensemble and transformer-based models in minimizing prediction errors within this study. Fig. 18. Open in a new tab Root mean squared error (RMSE) comparison of different models. Figure 19 compares the Mean Absolute Error (MAE) values obtained by different machine learning models. The Transformer, Random Forest, XGBoost, and Gradient Boosting models achieve minimal MAE values, effectively reducing average prediction errors to near-zero levels. Conversely, Linear Regression and Support Vector Regressor (SVR) display notably higher MAE values of 0.6565 and 0.2985, respectively. These results further emphasize the strong predictive accuracy and consistency of ensemble and transformer-based models compared to traditional regression techniques in this study. Fig. 19. Open in a new tab Mean absolute error (MAE) comparison across models. Figure 20 illustrates the R² scores achieved by various machine learning models, reflecting the proportion of variance in the target variable explained by each model. Random Forest, XGBoost, and Gradient Boosting attain perfect scores of 1.0000, while the Transformer model follows closely with an R² of 0.9988. In contrast, the Support Vector Regressor (SVR) and Linear Regression yield considerably lower scores of 0.8229 and 0.2739, respectively. These results confirm the superior explanatory power and model fit of advanced ensemble and transformer-based approaches compared to traditional regression techniques. Fig. 20. Open in a new tab Coefficient of determination (R²) comparison across models. Justification of the transformer architecture While tree-based models (Random Forest and XGBoost) achieved very high predictive performance, the primary objective of this study extends beyond maximizing accuracy on a static tabular dataset. The proposed Multi-Head Transformer architecture was selected to address fundamental limitations of classical machine learning models and to provide a scalable framework for future cybersecurity applications. Model expressiveness Tree-based models rely on fixed split rules and handcrafted feature interactions. Although effective for structured tabular data, they lack the ability to explicitly learn: High-order feature interactions. Context-dependent relationships. Long-range dependencies. In contrast, the Transformer model employs self-attention mechanisms that dynamically re-weight features based on context. This allows the network to: Capture complex cross-feature dependencies. Adapt importance weights per incident scenario. Model nonlinear interaction patterns more effectively. Scalability and Future Extensions The proposed architecture is designed to be extensible to: Temporal incident sequences. Streaming cybersecurity logs. Multi-event dependency modeling. This positions the model for real-world deployment in dynamic cybersecurity environments where incidents evolve over time. Interpretability advantages Unlike black-box deep models, the Transformer provides: Attention weight matrices. Feature contribution visualization. Context-aware importance estimation. This improves explainability and aligns with trustworthy AI requirements for security systems. Computational trade-off analysis Although Transformers introduce higher computational overhead, this cost is justified in: Large-scale deployments. Real-time monitoring systems. High-dimensional risk modeling. As shown in Table 7 , classical models such as Random Forest and XGBoost exhibit lower training and inference times due to their simpler architectures and CPU-based execution. However, these models lack the representational capacity to capture complex, high-order feature interactions. The proposed Transformer model demonstrates higher computational requirements, with a training time of 62.9 s and an inference latency of 1.21 ms per sample, attributable to its multi-head attention mechanism and deeper network structure. Table 7. Runtime and model complexity comparison. Model Training time (s) Inference time (ms/sample) Parameters Hardware Random forest 18.4 0.32 ~ 1.2 M CPU XGBoost 21.7 0.41 ~ 850 K CPU MLP 34.6 0.88 ~ 1.6 M GPU Proposed transformer 62.9 1.21 3.4M GPU Open in a new tab Despite this increased overhead, the Transformer’s computational cost remains within practical limits, particularly when leveraging GPU acceleration. The modest increase in inference latency supports its feasibility for real-time monitoring systems, while the enhanced modeling capacity justifies its deployment in large-scale and high-dimensional cybersecurity environments. These results highlight a favorable trade-off between predictive performance and computational efficiency. While the Transformer requires higher training and inference time, it remains within acceptable limits for real-time systems. GPU acceleration significantly reduces computational overhead, making deployment feasible in production environments. Attention Visualization Figures Figure 21 illustrates the attention weight distribution learned by Head 1 of the Multi-Head Transformer. Strong weights indicate dominant feature interactions, demonstrating how the model dynamically prioritizes cybersecurity factors such as financial loss, attack type, and targeted systems. The model assigns higher attention to impact-related variables during high-severity incidents, confirming domain consistency. Fig. 21. Open in a new tab Attention weight Heatmap—Head 1. Figure 22 Multi-head attention visualization showing complementary learning patterns across different heads. Each head captures distinct interaction structures among features. This demonstrates that: Fig. 22. Open in a new tab Attention comparison across heads. Some heads focus on financial and operational factors. Others prioritize organizational preparedness. Ensemble attention improves representation learning. Figure 23 shows the aggregated attention scores across all heads highlighting the most influential features in severity prediction. This confirms: Fig. 23. Open in a new tab Parameter sensitivity (heads vs. R2). No single feature dominates. Importance is distributed across multiple dimensions. Model learns balanced decision patterns. Practical implications for football cybersecurity The proposed severity prediction framework has direct and actionable implications for cybersecurity management within professional football organizations. Unlike generic cybersecurity environments, football clubs operate under unique constraints, including high public visibility, strong fan engagement through digital platforms, heterogeneous financial capacities, and tight competition schedules. Consequently, cyber incidents can lead not only to operational disruption but also to reputational damage, financial loss, and regulatory consequences. The predicted severity scores generated by the proposed Multi-Head Transformer model can be leveraged as a decision-support mechanism to prioritize cybersecurity incidents based on their anticipated impact. High-severity predictions enable clubs and league administrators to rapidly allocate technical resources, escalate response procedures, and activate contingency plans for critical systems such as ticketing platforms, broadcast infrastructure, player data repositories, and fan engagement services. Football-specific features, including club identity, league affiliation, and budget-related indicators, play a critical contextual role in this process. These attributes capture structural differences across organizations, such as disparities in cybersecurity investment, digital asset exposure, and tolerance to service downtime. As demonstrated by the explainability analysis, severity estimation emerges from interactions between these domain-specific factors and technical attack characteristics, rather than from club identifiers alone. At the league level, aggregated severity predictions can support strategic planning by identifying systemic risk patterns across competitions, enabling governing bodies to design targeted cybersecurity policies, awareness programs, and compliance requirements. At the club level, the framework can assist security teams in aligning incident response strategies with organizational risk profiles and operational priorities. Overall, this subsection positions the proposed model not merely as a predictive tool, but as a practical, interpretable, and scalable framework for enhancing cyber resilience in football ecosystems, bridging advanced machine learning methodology with real-world cybersecurity decision-making. Practical implications, advantages, and deployment challenges The proposed Multi-Head Transformer-based framework offers significant practical value for cybersecurity risk assessment in professional football organizations. One of its primary advantages lies in its ability to dynamically model complex feature interactions through self-attention mechanisms. Unlike traditional machine learning models that rely on fixed decision boundaries, the proposed architecture adaptively learns contextual relationships among financial, operational, reputational, and technical risk indicators, enabling more accurate and robust severity prediction. This capability is particularly beneficial in cybersecurity environments where threat patterns continuously evolve. Another key advantage of the proposed approach is its scalability and extensibility. The model is designed to accommodate future enhancements, including temporal incident modeling, real-time data streams, and multi-event dependency analysis. This makes it well-suited for integration into Security Information and Event Management (SIEM) systems, where continuous monitoring and early threat detection are critical. Additionally, the built-in attention mechanism provides improved interpretability, allowing security analysts to understand which factors contribute most to severity predictions, thereby supporting informed decision-making. Despite these strengths, several challenges must be considered for real-world deployment. The Transformer architecture introduces higher computational and memory requirements compared to classical machine learning models. Practical implementation may therefore require GPU acceleration and optimized hardware resources, particularly for real-time monitoring scenarios. Moreover, the model’s performance depends on the availability of high-quality and up-to-date cybersecurity data, which may vary across organizations. Issues such as data imbalance, noise, and missing values can impact predictive accuracy and must be addressed through continuous data preprocessing and model retraining. From a deployment perspective, additional considerations include system integration, model maintenance, and cybersecurity policy alignment. Periodic retraining is necessary to ensure that the model adapts to emerging threat patterns. Techniques such as model compression, knowledge distillation, and hybrid architectures combining deep learning with lightweight models can help mitigate computational overhead while maintaining performance. Looking ahead, the proposed framework has strong future prospects. It can be extended to support real-time streaming data, predictive threat forecasting, and cross-organizational threat intelligence sharing. Furthermore, the methodology can be adapted to other critical infrastructures such as healthcare, finance, and smart cities, where proactive cybersecurity risk assessment is equally vital. These extensions position the proposed model as a scalable and future-proof solution for intelligent cybersecurity management. Limitations and failure cases Despite the strong predictive performance achieved by the proposed Multi-Head Transformer framework, several limitations should be acknowledged to ensure realistic interpretation and responsible deployment. First, the model exhibits reduced robustness for rare and extreme severity cases, which are underrepresented in the dataset. Although cross-validation results indicate stable overall performance, predictions for highly severe incidents may be associated with increased uncertainty due to class imbalance. This limitation is common in real-world cybersecurity datasets and may be mitigated in future work through targeted resampling strategies or cost-sensitive learning. Second, the framework relies on the availability and quality of structured incident records. Incomplete, noisy, or inconsistently reported features such as missing impact assessments or inaccurate attack categorization can negatively affect prediction accuracy. In operational environments where data collection practices vary across clubs or leagues, this dependency may limit performance. Third, while football-specific contextual features enhance domain relevance, the model’s generalization capability may be constrained when applied to unseen football ecosystems, such as amateur leagues, lower-division competitions, or regions with substantially different financial and digital infrastructures. Domain adaptation or transfer learning techniques may be required to extend applicability beyond the studied context. Fourth, the proposed approach is inherently data-driven and retrospective, learning from historical incident patterns. As a result, it may be less effective against novel or rapidly evolving attack strategies that differ significantly from those observed during training. Incorporating online learning or adaptive retraining mechanisms represents a promising direction for future research. Finally, although explainability analyses indicate balanced feature contributions, the Transformer architecture remains more complex than traditional models. This increased complexity introduces higher computational costs and may limit real-time deployment in resource-constrained environments without further optimization. By explicitly acknowledging these limitations and failure cases, the study aims to provide a transparent assessment of the proposed framework and to outline clear directions for future improvement and extension. Despite its strong predictive performance, the proposed framework is subject to limitations related to rare extreme-severity cases, data quality dependencies, generalization to unseen football contexts, and increased model complexity, which highlight important directions for future research and deployment refinement. Conclusion and future work This study proposed a novel Multi-Head Transformer-based deep learning framework for predicting cybersecurity threat severity in professional football organizations. By leveraging self-attention mechanisms, the model dynamically captures complex feature interactions and contextual dependencies among financial, operational, reputational, and technical risk factors, enabling more expressive representation learning compared to traditional machine learning approaches. Experiments conducted on a customized dataset comprising 60,003 cybersecurity incident records demonstrated that the proposed architecture achieves state-of-the-art predictive performance, attaining an R² score of 0.9988 with consistently low error values. This confirms the model’s strong predictive capability and stability. To ensure methodological rigor and reliability, strict data partitioning strategies were applied to eliminate information leakage, followed by extensive robustness evaluations using k-fold cross-validation and multiple random seed experiments. Furthermore, SHAP-based feature importance analysis for tree-based baseline models and attention visualization for the Transformer confirmed that severity prediction is driven by multiple interacting factors rather than a single dominant feature. Ablation studies and parameter sensitivity analyses further validated the architectural design, highlighting the importance of multi-head attention in capturing diverse interaction patterns. In addition, computational complexity and memory footprint analyses demonstrated that, despite higher resource requirements, the proposed model remains feasible for real-world deployment when supported by GPU acceleration. The key innovations of this work include the first application of a Multi-Head Transformer for football cybersecurity severity prediction, the development of a realistic synthetic dataset tailored to sports organizations, the integration of explainable AI techniques, and the execution of comprehensive robustness and architectural validation experiments. Collectively, these contributions establish a powerful, interpretable, and extensible AI framework for proactive cybersecurity risk assessment in professional sports environments. The proposed methodology can be readily extended to other critical infrastructures, enabling data-driven security decision-making and enhancing organizational cyber resilience in increasingly complex threat landscapes. Future work will focus on extending the proposed model to multi-class classification tasks for categorizing cyber threats by severity levels, as well as integrating the framework into real-time monitoring systems for continuous threat assessment. Additionally, incorporating external data sources such as social media sentiment, network traffic logs, and IoT telemetry will be explored to enhance prediction contextuality and accuracy. Further investigations will consider the application of explainable AI techniques, such as SHAP and attention visualization, to improve model interpretability and transparency. Cross-domain evaluations in other sports or enterprise cybersecurity environments, alongside comparative studies with emerging architectures like TabTransformer variants and Performer models, are also planned to validate scalability and generalizability. These directions aim to refine and expand the practical utility of transformer-based models for intelligent, data-driven cybersecurity decision-making in football and beyond. Acknowledgements The authors are thankful to the Deanship of Graduate Studies and Scientific Research at University of Bisha for supporting this work through the Fast-Track Research Support Program. The authors would like to thank the Deanship of Scientific Research at Shaqra University for supporting this work. Author contributions Fahad Algarni: Write the Related Work Rayan Alshamrani: Write the Proposed methodology Ashrf Althbiti: Write the Results sectionAbdullah Albalawi: Write the DiscussionAtef Ismail: Revised the paperBasma M. Hassan: Proposed the article plan, and wrote the methodology, and the supervisor. Funding Not Applicable. Data availability Data and code will be available on request. Declarations Competing interests The authors declare no competing interests. Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Change history 5/12/2026 The original online version of this Article was revised: The original version of this Article contained errors in the Acknowledgements. The Acknowledgements section now reads: “The authors are thankful to the Deanship of Graduate Studies and Scientific Research at University of Bisha for supporting this work through the Fast-Track Research Support Program. The authors would like to thank the Deanship of Scientific Research at Shaqra University for supporting this work.” The original Article has been corrected. References 1. Khalid, S., Bose, M. & Yadav, A. The Multifarious Role of Media in Reflecting and Shaping Soccer’s Global Significance, in Advances in Media, Entertainment, and the Arts , F. P. C. Endong, Ed., IGI Global, 141–180 (2024). 10.4018/979-8-3693-4298-5.ch006 2. Qi, Y., Sajadi, S. M., Baghaei, S., Rezaei, R. & Li, W. Digital technologies in sports: Opportunities, challenges, and strategies for safeguarding athlete wellbeing and competitive integrity in the digital era. Technol. Soc. 77 , 102496. 10.1016/j.techsoc.2024.102496 (Jun. 2024). 3. Anderson, R. et al. Measuring the cost of cybercrime, in The economics of information security and privacy , (ed Böhme, R.) 265–300 (Springer, Berlin, 2013). 10.1007/978-3-642-39498-0_12. [ Google Scholar ] 4. Jenkins, S. & Evans, N. Cybersecurity impact of the growth of data in sports, in Cyber Sensing 2020 , P. Chin and I. V. Ternovskiy, Eds., Online Only, Apr. 18. (SPIE, 2020). 10.1117/12.2557898 5. Seloom, M. Qatar’s security strategy In the 2022 FIFA World Cup. SSRN Electron. J. 10.2139/ssrn.4712971 (2024). [ Google Scholar ] 6. Al-Thani, M. In the liminal realm: Qatar’s world cup struggle between tradition, modernity, and human rights. Front. Sports Act. Living 6 , 1434522. 10.3389/fspor.2024.1434522 (2025). 7. Albakri, M., Bello, M. & Rashdi, S. A. Digital transformation and cybersecurity fears on national security in GCC:, in Advances in Electronic Commerce , M. Albakri, Ed., 25–60 (IGI Global, 2024). 10.4018/979-8-3693-5966-2.ch002 8. Hokmabadi, H., Rezvani, S. M. H. S. & De Matos, C. A. Business resilience for small and medium enterprises and startups by digital transformation and the role of marketing capabilities—A systematic review. Systems 12 (6), 220. 10.3390/systems12060220 (2024). 9. Rehman, F. & Hashmi, S. Enhancing cloud security: A comprehensive framework for real-time detection, analysis and cyber threat intelligence sharing. Adv. Sci. Technol. Eng. Syst. J. 8 (6), 107–119. 10.25046/aj080612 (2023). 10. Chindrus, C. & Caruntu, C. F. Securing the network: A red and blue cybersecurity competition case study. Information 14 (11), 587. 10.3390/info14110587 (Oct. 2023). 11. Merten, S., Schmidt, S. L. & Winand, M. Organisational capabilities for successful digital transformation: a global analysis of national football associations in the digital age, J. Strategy Manag. 17 (3), 408–426. 10.1108/JSMA-02-2022-0039 (2024). 12. Bongiovanni, I., Herold, D. M. & Wilde, S. J. Protecting the play: An integrative review of cybersecurity in and for sports events. Comput. Secur. 146 , 104064. 10.1016/j.cose.2024.104064 (2024). 13. Alshammari, F. H. Design of capability maturity model integration with cybersecurity risk severity complex prediction using bayesian-based machine learning models, Serv. Oriented Comput. Appl. 17 (1), 59–72. 10.1007/s11761-022-00354-4 (2023). 14. Rafy, M. F. Artificial intelligence in cyber security. SSRN Electron. J. 10.2139/ssrn.4687831 (2024). [ Google Scholar ] 15. Choi, S. R. & Lee, M. Transformer architecture and attention mechanisms in genome data analysis: A comprehensive review. Biology 12 , 1033. 10.3390/biology12071033 (2023). 16. Somvanshi, S., Das, S., Javed, S. A., Antariksa, G. & Hossain, A. A survey on deep tabular learning. (2024). arXiv:2410.12034. 17. Yiapanas, G. The application of big data analytics in sports as a tool for personalized fan experience, operations efficiency, and fan engagement strategy. Bus. Manag Theory Pract. 2 (1), 3075. 10.54517/bmtp3075 (2025). 18. Dodiya, K. R., Varayogula, S. N. & Gohil, B. V. Rising threats, silent battles: A deep dive into cybercrime, terrorism, and resilient defenses, in Cases on Forensic and Criminological Science for Criminal Detection and Avoidance , A. Chaussée and L. J. Leonard, Eds. 123–150. (IGI Global, 2024). 10.4018/978-1-6684-9800-2.ch006 19. Watters, P. Exposing the dark side: Scams and cybersecurity risks in Indonesia’s illicit sports streaming scene (2024). 10.2139/ssrn.4954969 20. Ahmed, U. et al. Author Correction: Signature-based intrusion detection using machine learning and deep learning approaches empowered with fuzzy clustering. Sci. Rep. 15 (1), 8033. 10.1038/s41598-025-92132-3 (2025). 21. Tahmasebi, M. Beyond defense: Proactive approaches to disaster recovery and threat intelligence in modern enterprises. J. Inf. Secur. 15 (02), 106–133. 10.4236/jis.2024.152008 (2024). [ Google Scholar ] 22. Mustafa, F. M. et al. May., TabNet and TabTransformer: Novel deep learning models for chemical toxicity prediction in comparison with machine learning. J. Appl. Toxicol. jat.4803. 10.1002/jat.4803 (2025). 23. Whitmore, A., Agarwal, A. & Xu, L. D. The internet of things—A survey of topics and trends. Inf. Syst. Front. 17 (2), 261–274. 10.1007/s10796-014-9489-2 (2015). 24. Razzaq, K. & Shah, M. Advancing cybersecurity through machine learning: A scientometric analysis of global research trends and influential contributions. J. Cybersecur. Priv. 5 (2), 12. 10.3390/jcp5020012 (2025). 25. LeCun, Y., Bengio, Y. & Hinton, G. Deep learning. Nature 521 (7553), 436–444. 10.1038/nature14539 (May 2015). 26. Nobles, C. Offensive artificial intelligence in cybersecurity: Techniques, challenges, and ethical considerations, in Advances in Human Resources Management and Organizational Development , D. N. Burrell, Ed., 348–363 (IGI Global, 2023). 10.4018/978-1-6684-8691-7.ch021 27. Ben Fredj, O., Mihoub, A., Krichen, M., Cheikhrouhou, O. & Derhab, A. CyberSecurity attack prediction: A deep learning approach, in 13th International Conference on Security of Information and Networks , 1–6 (ACM, Merkez, 2020). 10.1145/3433174.3433614 28. Ahmed, U. et al. Signature-based intrusion detection using machine learning and deep learning approaches empowered with fuzzy clustering. Sci. Rep. 15 (1), 1726. 10.1038/s41598-025-85866-7 (2025). 29. Rahman, M. M., Gupta, D., Bhatt, S., Shokouhmand, S. & Faezipour, M. A comprehensive review of machine learning approaches for anomaly detection in smart homes: Experimental analysis and future directions. Future Internet . 16 (4), 139. 10.3390/fi16040139 (Apr. 2024). 30. Umair, M. B. et al. A network intrusion detection system using hybrid multilayer deep learning model. Big Data . 12 (5), 367–376. 10.1089/big.2021.0268 (2024). 31. Sharma, H., Kumar, P. & Sharma, K. Recurrent neural network based incremental model for intrusion detection system in IoT. Scalable Comput. Pract. Exp. 25 (5), 3778–3795. 10.12694/scpe.v25i5.3004 (2024). 32. Arun, M. et al. Internet of things and deep learning-enhanced monitoring for energy efficiency in older buildings. Case Stud. Therm. Eng. 61 , 104867. 10.1016/j.csite.2024.104867 (2024). 33. Sani, G. Energy-efficient deep learning models for cyber threat prediction in constrained IoT infrastructure. 10.31219/osf.io/529qp_v1 (2025). 34. Lakhani, A. AI revolutionizing cyber security unlocking the future of digital protection. 31 10.31219/osf.io/tgz8j (2024). 35. Zheng, K. Next-generation cybersecurity threat detection: integration with artificial intelligence. Highlights Sci. Eng. Technol. 138 , 8–16. 10.54097/nx38v729 (2025). 36. Olakunle, A. A. & Balogun, O. A. Leveraging AI/ML for anomaly detection, threat prediction, and automated response. World J. Adv. Res. Rev. 21 (1), 2584–2598. 10.30574/wjarr.2024.21.1.0287 (2024). 37. Alionsi, D. D. D. AI-driven cybersecurity: Utilizing machine learning and deep learning techniques for real-time threat detection, analysis, and mitigation in complex IT networks. Adv. Eng. Innov. 3 (1), (2023). 10.54254/2977-3903/3/2023036 38. Tian, J. et al. LESSON: Multi-label adversarial false data injection attack for deep learning locational detection (2024). 10.48550/ARXIV.2401.16001 39. Tian, J. et al. Targeted adversarial false data injection attacks for state estimation in smart grid. IEEE Trans. Sustain. Comput. 10 (3), 534–546. 10.1109/TSUSC.2024.3492290 (2025). 40. Saber, A. M., Maheshwari, A., Youssef, A. & Kundur, D. Adversarial attacks on deep learning-based false data injection detection in differential relays (2025). arXiv:2506.19302. 41. Mukherjee, D. Detection of data-driven blind cyber-attacks on smart grid: A deep learning approach. Sustain. Cities Soc. 92 , 104475. 10.1016/j.scs.2023.104475 (May 2023). 42. Wu, Q., Lin, H., Ma, X. & Cai, C. State estimation for smart grid under hybrid cyberattacks: Analysis of attack stealthiness and estimator stability. IEEE Trans. Ind. Inform. 22 (1), 328–335 (2026). 10.1109/TII.2025.3609212 43. Tan, Z. & Li, Z. Digital twins for sustainable design and management of smart city buildings and municipal infrastructure. Sustain. Energy Technol. Assess. 64 , 103682. 10.1016/j.seta.2024.103682 (2024). 44. Paudel, S. & Dhungana, D. False data injection attacks against state estimation in the smart grid, in IEEE EUROCON 2025–21st International Conference on Smart Technologies 1–6 (IEEE, Gdynia, 2025). 10.1109/EUROCON64445.2025.11073195 45. Zhang, G. et al. Detection and localization of false data injection attacks in smart grid based on joint maximum a posteriori-maximum likelihood. IEEE Access 11 , 133867–133878. 10.1109/ACCESS.2023.3336683 (2023). [ Google Scholar ] 46. Tasew Reda, H., Anwar, A., Mahmood, A. & Chilamkurti, N. Attacks in Smart Grid. Data-driven Approach for State Prediction and Detection of False Data Injection. J. Mod. Power Syst. Clean. Energy 11 (2), 455–467. 10.35833/MPCE.2020.000827 (2023). 47. Zhang, G. et al. Detection of false data injection attacks in a smart grid based on WLS and an adaptive interpolation extended Kalman filter. Energies 16 (20), 7203. 10.3390/en16207203 (2023). 48. Wang, Y., Zhou, Y., Ma, J. & Jin, Q. A locational false data injection attack detection method in smart grid based on adversarial variational autoencoders. Appl. Soft Comput. 151 , 111169. 10.1016/j.asoc.2023.111169 (2024). 49. Zhu, J., Meng, W., Sun, M., Yang, J. & Song, Z. FLLF: A fast-lightweight location detection framework for false data injection attacks in smart grids. IEEE Trans. Smart Grid 15 (1), 911–920 (2024). 10.1109/TSG.2023.3274642 50. Kesici, M., Alduhaymi, M., Ishchenko, A., Yang, G. & Pal, B. C. Locational detection of false data injection attacks in smart grids: A cloud–edge framework based on split learning. IEEE Trans. Smart Grid 16 (6), 5378–5391. 10.1109/TSG.2025.3580683 (2025). 51. Shen, K., Yan, W., Ni, H. & Chu, J. Localization of false data injection attack in smart grids based on SSA-CNN. Information 14 (3), 180. 10.3390/info14030180 (2023). 52. Cui, M., Zhang, C., Kang, B., Zhang, J. & Gómez-Expósito, A. Privacy-preserving FDIA detection for radial distribution systems considering attacker preference. IEEE Trans. Power Syst. 41 (1), 270–282. 10.1109/TPWRS.2025.3599851 (2026). 53. Ruan, J. et al. Super-resolution perception assisted spatiotemporal graph deep learning against false data injection attacks in smart grid. IEEE Trans. Smart Grid 14 (5), 4035–4046. 10.1109/TSG.2023.3241268 (2023). 54. Baul, A., Sarker, G. C., Sadhu, P. K., Yanambaka, V. P. & Abdelgawad, A. XTM: A novel transformer and LSTM-based model for detection and localization of formally verified FDI attack in smart grid. Electronics 12 (4), 797. 10.3390/electronics12040797 (2023). 55. Li, J., Lu, H. & Su, Q. Detection and localization of false data injection attacks based on multi-scale feature fusion and attention enhancement network in smart grid. Eng. Appl. Artif. Intell. 163 , 112787. 10.1016/j.engappai.2025.112787 (2026). 56. Jiang, W. et al. Dec., A Dual-Channel CNN for Precise Localization of False Data Injection Attacks in Smart Grids, Iran. J. Sci. Technol. Trans. Electr. Eng. 49 (4), 1693–1710. 10.1007/s40998-025-00895-2 (2025). 57. Gupta, S. et al. A Residual Dense Network Approach for False Data Injection Attack Localization in Power Grid, in IEEE North-East India International Energy Conversion Conference and Exhibition (NE-IECCE) 1–6 (IEEE, Silchar, 2025). 10.1109/NE-IECCE64154.2025.11183432 58. Chen, Y., Lakshminarayana, S. & Vincent Poor, H. Moving target defense against adversarial false data injection attacks in power grids. IEEE Internet Things J. 12 (14), 26315–26327. 10.1109/JIOT.2025.3560127 (2025). 59. Nawaz, R., Akhtar, R., Khan, S. U., Bu, S. & Mahmood, M. H. Deep learning-driven false data injection attack in renewable integrated smart grids. Eng. Appl. Artif. Intell. 156 , 110953. 10.1016/j.engappai.2025.110953 (2025). 60. Rudin, C. et al. Interpretable machine learning: Fundamental principles and 10 grand challenges. Stat. Surv. 16 . 10.1214/21-SS133 (2022). 61. Dahouda, M. K. & Joe, I. A deep-learned embedding technique for categorical features encoding. IEEE Access 9 , 114381–114391. 10.1109/ACCESS.2021.3104357 (2021). [ Google Scholar ] 62. Stankevičius, L. & Lukoševičius, M. Extracting sentence embeddings from pretrained transformer models. Appl. Sci. 14 (19), 8887. 10.3390/app14198887 (2024). 63. Moshawrab, M., Adda, M., Bouzouane, A., Ibrahim, H. & Raad, A. reviewing multimodal machine learning and its use in cardiovascular diseases detection. Electronics 12 (7), 1558. 10.3390/electronics12071558 (2023). Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Data Availability Statement Data and code will be available on request. Articles from Scientific Reports are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (5.5 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top