Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Sci Rep . 2026 Mar 2;16:11769. doi: 10.1038/s41598-026-38518-3 Search in PMC Search in PubMed View in NLM Catalog Add to search Evolutionary reinforcement learning framework for energy-efficient fault resilience and topological stability in WSNs S Lakshmi S Lakshmi 1 Electronics and Communication Engineering, Thirumalai Engineering College, Kilambi, Kanchipuram, 631551 India Find articles by S Lakshmi 1, ✉ , S Aswath S Aswath 2 Department of Electronics and Communication Engineering, Vel Tech Rangarajan Dr. Sagunthala R&D institute of Science and Technology, Chennai, India Find articles by S Aswath 2 , A Swaminathan A Swaminathan 3 Department of Computer Science and Business systems, Panimalar Engineering College, Poonamallee, India Find articles by A Swaminathan 3 , Anandakumar Haldorai Anandakumar Haldorai 4 Department of Computer Science and Engineering, Sri Eshwar College of Engineering, Coimbatore, India Find articles by Anandakumar Haldorai 4 Author information Article notes Copyright and License information 1 Electronics and Communication Engineering, Thirumalai Engineering College, Kilambi, Kanchipuram, 631551 India 2 Department of Electronics and Communication Engineering, Vel Tech Rangarajan Dr. Sagunthala R&D institute of Science and Technology, Chennai, India 3 Department of Computer Science and Business systems, Panimalar Engineering College, Poonamallee, India 4 Department of Computer Science and Engineering, Sri Eshwar College of Engineering, Coimbatore, India ✉ Corresponding author. Received 2025 Oct 29; Accepted 2026 Jan 29; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13065733 PMID: 41771944 Abstract Wireless Sensor Networks (WSNs) form the technological foundation for several modern applications, including smart cities, environmental monitoring, healthcare surveillance, and industrial automation. However, the performance and long-term reliability of WSNs continue to be constrained by persistent challenges such as communication faults, accelerated energy depletion, reduced network lifetime, increased latency, and inconsistent throughput. These challenges are further intensified by the growing scale, density, and dynamic behavior of contemporary WSN deployments. Existing techniques frequently optimize one dimension—such as energy conservation or communication quality—while neglecting equally important aspects like fault recovery, adaptability, and real-time responsiveness. To address these limitations, this study introduces EvoGenRL, an Evolutionary Reinforcement Learning framework that integrates Reinforcement Learning (RL), Differential Evolution (DE), and Generative Adversarial Networks (GANs) for robust WSN management. In EvoGenRL, RL is utilized to learn adaptive routing and fault-tolerant decision policies; DE optimizes hyperparameters to improve convergence stability and policy efficiency; and GANs generate diverse, realistic fault scenarios to enrich the training environment and enhance model generalization. This combined strategy enables WSNs to operate efficiently under uncertain, heterogeneous, and failure-prone conditions. Experimental evaluations confirm that EvoGenRL delivers substantial improvements compared to conventional optimization and routing schemes. The proposed method successfully reduces energy consumption to 2.2 J, extends network lifetime to 1700 cycles, increases packet delivery ratio to 99.7%, lowers latency to 3.2 ms, and boosts throughput to 350 kbps. These advancements demonstrate the capability of EvoGenRL to simultaneously enhance energy efficiency, communication performance, and fault resilience. Overall, this research provides a comprehensive and scalable solution for next-generation WSN optimization. The EvoGenRL framework not only addresses current operational limitations but also establishes a foundation for future developments in intelligent, adaptive, and power-efficient sensor network control. Keywords: Wireless sensor networks, Energy efficiency, Fault resilience, Reinforcement learning, Generative adversarial networks Subject terms: Energy science and technology, Engineering, Mathematics and computing Introduction WSNs have turned out to be ground breaking technologies because now it is possible to monitor the surrounding environment in real-time, engage in accurate farming, automate industries, oversee health care, and even administer city infrastructure. A typical WSN consists of a great amount of sensor nodes which are spatially spread and detect, process and relay essential information to a central base station or cloud platform 1 , 2 . This distributed but networked design encourages the gathering of big data and offers significant advantages to unceasing and fine tuned monitoring. The resource constraints of WSNs are extremely high, particularly, the limited amount of energy and processing resources 3 , 4 . The majority of sensor nodes use non-renewable, battery-powered operation and are commonly operated in highly inaccessible, remote or dangerous locations where they are hard to maintain. This consequently makes it one of the key challenges to extend the network lifetime and maintain its continuous operation 5 . Moreover, WSNs often work under conditions of high dynamics, as uncommon weather, physical interference, and environmental degradation may cause node failure, and link disruption, as well as inconsistent network connection. Such failures can spread throughout the network resulting in higher power usage, diminished data integrity, and poor communication throughput 6 , 7 . Fault resilience combined with a stable and efficient topology is hence the key to successful operation of the WSN. The performance of WSN is also multi-dimensional by its nature and energy efficiency, fault recovery, scalability, and communication reliability should be managed simultaneously 8 , 9 . Nonetheless, conventional methods, including predetermined routing methods and heuristic optimization methods are usually only offered in the case of static situations and are not responsive to any dynamic variations, including changing traffic demands, node outages, and changing environmental factors 10 . Recently, machine learning and especially Reinforcement Learning (RL) have provided new possibilities whereby WSNs can learn and learn new optimal policies through autonomous learning based on real-time interactions 11 , 12 . However, the usefulness of RL is greatly determined by the variety of training situations and efficiency of search strategies. Currently used RL based solutions are plagued with slow convergence and lack of exposure to real world failure modes, which restrict their application to real world deployments 13 , 14 . These disadvantages lead to the necessity to have a unified system capable of creating realistic scenario generation, effective exploration and dynamic policy optimization. This would prepare the WSNs with good learning skills that would enable them confront the prevailing challenges in their operations and achieve higher performance, reliability and resiliency in different fields of operation. Problem statement WSNs are currently an indispensable part of the new technological ecosystem, and they are used in the infrastructure of smart cities, environmental monitoring, automation in industries, and healthcare surveillance. Despite this however, WSNs are never without some critical problems that are always faced by the network relating to energy efficiency, fault resilience and stability of the entire network 15 . The sensor nodes have a limited non-renewable battery and are very vulnerable to malfunctions due to the environmental factors, wear and tear of the hardware, or disruption of the mode of operation 16 . These challenges are further amplified in dynamic environments where traffic patterns, extreme weather, and physical obstructions vary and thus, disruption of communication and data loss are probable. The old methods of solutions like the use of the static routing scheme and heuristic based optimization methods do not have the flexibility needed to react well to evolving network environments. These approaches are typically developed to be used in fixed cases and thus they do not cope with the multi-dimensionality of the WSN operations where energy limitation, fault recovery, and stability of topology have to be addressed all at the same time. Reinforcement Learning (RL) has also become a promising alternative of learning optimal decisions by means of interaction with the environment 17 . However, the existing RL-based frameworks are not effective. They are tortured with ineffective exploration plans, tardiness on convergence and absence of competent and realistic training scenarios capable of accurately reflect authentic defects in the WSN. Moreover, the current practices are geared towards the optimization of energy, fault tolerance as well as the stability of communication as individual objectives rather than a combination. All these shortcomings indicate a massive research gap. More optimistic, dynamic and synthesized optimization architecture providing an integration of intelligent learning, non-weak exploration and modelling of real life situation is much needed. Such form of a structure is essential in ensuring that it delivers reliable, fault tolerant and energy efficient WSNs functionality in volatile and unforeseeable situations. Motivation of the study The growing application of the Wireless Sensor Networks (WSNs) in some of the most crucial missions has informed this research, some of which include disaster recovery, healthcare, environmental monitoring, and industrial control. These areas are the ones where the network reliability, resilience and efficiency are paramount so as to be able to make a timely and accurate decision. However, the existing WSN functions are not applicable in the environment of increasing complexity of modern deployments, when energy constraints, communication failures, unexpected environmental factors, and enormous topological alterations all occur simultaneously. The customized optimization methods are less adaptable to such volatile and unpredictable situations. They are usually designed to be applied in situations which are not dynamic and therefore not capable of dealing with the interconnected problems of energy efficiency, fault tolerance and the stability of the network constantly. This constraint creates a sense of desperate need of self adapting and shrewd approaches which can learn and respond to real time conditions in the network. The recent advances in Reinforcement Learning (RL) and Generative Adversarial Networks (GANs) have some fascinating opportunities of reconsidering the notion of WSN management. One is RL that enables optimal operation policy to be trained autonomously, and GANs that enable the development of different training scenarios with realistic fault occurrence and environmental perturbation. Combined with the powerful optimization algorithms such as the Differential Evolution (DE), they will be able to develop one system that will be able to solve the various problems that are not solvable by the conventional means. The present study, therefore, is motivated by the fact that it needs to develop a comprehensive, dynamic and robust model, which would make the network more energy efficient, more resilient to faults, and more stable on the whole. The research will contribute to the future of intelligent management of WSNs, as the approaches to RL, GANs, and evolutionary optimization will be combined to achieve the successful operation of the sensor networks in a real-life environment and ensure their high scalability, reliability, and proper functionality. Significance of the study The research is of significant importance as it tries to come up with a hybrid framework in order to deal with difficult WSNs challenges and offer more reliable and efficient network operations. High-order machine learning algorithms and evolutionary optimization are implemented including a suggested solution delivering a comprehensive solution of energy efficiency and fault resilience. It is an essential innovation in applications where nonstop flow of information and stability of the network is significant like in medical systems and environmental control. Besides, the outcomes of these studies may be applied as a contextual pool of the optimization of other resource-limited network systems and the effect may spread over several fields of technological progress and social values. Key Contributions of the Study Introduced EvoGenRL, a unified framework combining RL, DE, and GANs for holistic WSN optimization. Developed adaptive RL policies for real-time energy-efficient and fault-tolerant routing. Used GANs to generate realistic fault scenarios, improving RL robustness and training diversity. Applied DE for hyperparameter optimization, enabling faster and more stable RL convergence. Achieved significant performance gains, including lower energy use, longer network lifetime, higher PDR, reduced latency, and improved throughput. Even though a few of the studies integrate the reinforcement learning with metaheuristic optimization in routing, the proposed framework presents the key innovations that are not considered in comparison with the current hybrid methods. To begin with, a fault pattern generator with the help of GAN offers synthetic and high-variability fault scenarios, enabling an RL agent to be trained to exhibit behaviors that are more resilient compared to those which it has been trained to exhibit in the context of a conventional RL metaheuristic model. Second, the evolutionary element has not been applied as the simple parameter tuning, as it is shown in the paper that the Differential Evolution dynamically optimizes the RL hyperparameters regarding real-time performance to achieve quicker convergence and guarantee a better stability. Third, the model clearly brings in the state features and reward factor which are fault aware, thus enabling the agent to specifically aim at recovery and reliability in situations of dynamic network failures, which is not looked at in existing hybrid routing methods. All these components form a single fault-centric optimization framework that is better than existing RL-based routing models. Structure of the Study: Chap. 1 : Introduction, Chap. 2: Literature Review, Chap. 3: Methodology, Chap. 4: Results and Discussion, and Chap. 5: Conclusion and Future Work, detailing framework design and implications. Roadmap of the study: This research paper is based on a systematic research roadmap that is aimed to overcome the shortcomings of traditional methods of optimizing WSNs. The main problems of energy efficiency, fault resilience, and network stability in dynamic WSN environments are identified first with the help of the problem analysis. Then, the EvoGenRL system is designed, which combines the Reinforcement Learning to make adaptive decisions, GANs to create realistic fault-scenario, and Differential Evolution to optimize its hyperparameter. The suggested framework is then put to test and exercised under a variety of simulated WSN environments towards testing its versatility and resilience. After this, there are wide-scale performance experiments that are carried out to compare EvoGenRL with the current methods. Lastly, the results are discussed to reflect the improvement in performance, limitations, and the future work opportunities. Related works WSNs have been one of the most active areas of research in the pursuit of enhancing energy efficiency, fault resilience, and better network performance. Various approaches have been put forward, right from heuristic-based routing to AI- and ML-driven optimization techniques. Energy-Efficient routing and Trust-Based protocols Singh et al. 18 proposed an EEOTSR protocol by integrating ACO-based dynamic routing and a distributed trust model to detect and isolate malicious nodes. This protocol enhanced the lifetime of the network by 23% and the packet delivery ratio by 17%. Besides, it ensured a high security level against blackhole attacks. However, this approach suffers from computational overhead, and performance in dense networks remains questionable. Prasad and Periyasamy 19 proposed edge-based clustering and routing using Tasmanian Devil Optimization and Lightweight Encryption Algorithm, together with GTGAN-based secure multipath routing. Although it outperformed other techniques in energy efficiency and data delivery, the limitations lie in scalability and computational complexity for large-scale networks. Also, Hajar et al. 20 , present 3R: a reliable multi-agent reinforcement learning routing protocol for Wireless Medical Sensor Networks. For faster convergence, 3R utilizes a lightweight RL model with two update methods. With the Trust Management System in its reward function, it also defends against packet-dropping attacks. The approach leverages a local-information-based energy model to improve network lifetime and balance consumption. Experimental studies confirm that 3R is lightweight, resilient against attack, and energy-efficient for WMSN applications. Song et al. 21 , propose HDQN: a Deep Q-Network-based energy efficiency optimization algorithm for heterogeneous wireless sensor networks. Firstly, HDQN selects an optimal transmission path depending on the energies of nodes, distances, the number of relays, and real-time environmental information. Then its reward function and enhanced packet header further help to make accurate routing decisions. The simulations demonstrated that HDQN performed better than EEHCHR, 2LHMGEAR, NCOGA, DEEC, and SEP regarding energy efficiency, network lifetime, and robustness. Optimization in IoT and fanets Hussain et al. 22 have proposed IEE-DLSA for UAV-based FANETs and have used relay mechanisms along with Red-Black trees for improving link stability, which reduces transmission delays. Although effective in improving throughput and energy efficiency, the approach may face scalability issues in large FANET deployments. Udayaprasad et al. 23 proposed an energy-aware routing framework for cloud-based Intelligent IoT (I-IoT) networks using a hybrid of Genetic Algorithm (GA), Particle Swarm Optimization (PSO), and Artificial Bee Colony (ABC). Their approach maximized network lifetime and improved load balancing while reducing data transmission delays. However, high computational complexity limits its real-time applicability in highly dense networks. AI and ML-Driven network management Khan et al. 24 further developed the Low Energy Adaptive Clustering Hierarchy protocol by incorporating digital twins and AI analytics for efficient cluster-head management, leading to enhanced energy conservation. The method decreased non-operational nodes by 83%, but deploying large-scale digital twins is considered challenging. Thangavelu and Rajendran 25 proposed a heterogeneous IoT network approach that combined Parallelized Memetic Algorithm clustering with AlexNet-based anomaly detection. This new hybrid anomaly detection achieved 99.11% accuracy and enhanced network efficiency, though possibly at the cost of computational overhead which may reduce real-time performance in large-scale deployments. Metaheuristic and evolutionary approaches Kumar et al. 26 proposed the use of AGAN, along with PFOA for optimizing WSN in terms of layout, connectivity, coverage, and energy consumption. Simulations demonstrated significant improvement in packet delivery ratio, network lifetime, and routing efficiency; however, the approach results in computational complexity for dense networks. Alharbi 27 presented the implementation of energy imbalance in Industrial WSNs by using the WirelessHART Density-controlled Divide-and-Rule (WDDR) topology along with CMA-ES for graph-based routing. Improvement in energy balancing and packet delivery resulted, but scalability issues persist in large industrial deployment. Gupta et al. 28 presented the usage of a multilayer hierarchical clustering framework, namely RHCER, to reduce energy consumption and enhance network lifetime. Limitations arise in dynamic traffic and topology changes. Rao, Singh, and Nagwanshi 29 integrated an Optimized Secure Clustering-Based Routing Protocol for IoT-oriented WSN. Very high throughput and energy efficiency are achievable through this protocol. However, DES encryption might not be suitable for cryptography in the future, and scalability is a great concern. Most of the existing research works address energy optimization, fault resilience, or secure routing independently of each other. Fewer studies that use multiple techniques together often result in computational overhead, limited adaptability, and scaling issues when network conditions become dynamic. Relatively few approaches offer a holistic solution that simultaneously optimizes energy efficiency, fault tolerance, and network topology stability, thus calling for integrated and adaptive frameworks. Based on this research gap, the EvoGenRL framework will be proposed as a combination of RL, GAN-based scenario generation, and evolutionary optimization to enable robust, scalable, and energy-efficient WSN management under dynamic operational conditions. Summarization of the related works is given in Table 1 . Table 1. Summarization of the related works. Study Proposed Solution Key Results Limitations Singh et al 18 . EEOTSR: ACO-based secure routing for WSNs. + 23% lifetime, > 90% PDR under blackhole attacks. High overhead; struggles in dense networks. Prasad and Periyasamy 19 Secure clustering with LEA encryption and GTGAN routing. Better energy efficiency, delivery, and lifetime. Complexity in dense networks. Hajar et al. 20 , 3R: Multi-agent RL routing for WMSNs with trust-based reward and local energy model Lightweight, fast convergence, attack-resilient, energy-efficient Limited to WMSNs; scalability not tested Song et al. 21 , HDQN: DQN-based routing for heterogeneous WSNs using energy, distance, and relay info Better energy efficiency, network lifetime, and robustness than EEHCHR, 2LHMGEAR, NCOGA, DEEC, SEP Needs accurate environmental data; high DQN computation Hussain et al 22 . IEE-DLSA: R-B tree for FANET routing. Improved stability, throughput, reduced energy. Overhead; scalability issues in UAV networks. Udayaprasad et al 23 . GA + PSO+ABC routing for SDN-IoT. + 63% lifetime, −78% energy, + 95% data loading. High computational complexity. Khan et al 24 . Enhanced LEACH with digital twins and AI analytics. 1.66x energy levels, −83% dead nodes. Scalability and real-time bottlenecks. Thangavelu and Rajendran 25 EEMCM: PMA for clustering, AlexNet for anomaly detection. 99.11% accuracy, better efficiency. High resource demands for large networks. Kumar et al 26 . AGAN-PFOA: Optimized WSN layout with GAN+PFOA. + 61% PDR, + 65% lifetime, −69% delay. Computationally intensive. Alharbi 27 WDDR topology with CMA-ES for IWSNs. Balanced energy, improved delivery. Scalability in dense networks. Gupta et al 28 . EEDC: Tier-based RHCER for IoT-WSNs. −31% energy, + 38% packet drop reduction. Issues with dynamic topologies. Rao, Singh, and Nagwanshi 29 OSCBRP: DES encryption for SA WSNs. + 98% throughput, −34% packet drop, 0.002s latency. DES may be outdated; scalability concerns. Open in a new tab Framework for Energy-Efficient fault resilience and topological stability in WSNs The WSN will not be able to work without modern technological applications such as environmental monitoring, medical systems, automation in industries or smart cities. These networks consist of spatially dispersed sensor nodes that gather and send key information to serve in real time decision making. However, their functioning is governed by the presence of limited sources of energy, uncertain weather circumstances and the risk of failure of nodes which has an actual impact on their stability and efficacy. The current solutions which include the use of the static routing protocols and heuristic optimization procedures, however, cannot handle the dynamic and resource constrained environment of WSNs. This in turn necessitates new designs that do not only increase the energy efficiency and fault tolerance, but also guarantee topological stability to maintain network functionalities. As WSN applications become increasingly complex, real-time adaptiveness to changing conditions (e.g. network topology, fault conditions) is required. These challenges can be addressed by traditional RL frameworks which provide the ability for autonomous optimization of policies. Nonetheless, RL approaches usually have problems of exploration strategies and do not have the ability to generate various and realistic training scenarios similar to real world scenarios. In addition, the problem of optimizing energy consumption while keeping fault tolerance and maintaining stable communication pathways makes for a multi objective problem that is not effectively addressed by the conventional techniques. To drive the concept of the advanced hybrid framework, these challenges call for the integration of state of the art machine learning techniques and evolutionary optimization strategies to redefine WSN performance. In this study, a novel hybrid approach, named EvoGenRL (Evolutionary Generative Reinforcement Learning) that combines RL and GAN with DE) optimization is proposed. Energy efficient fault resilience and topological stability in WSNs have been addressed through the framework. GAN is used to generate training scenarios that are diverse and use them to improve the adaptability and robustness of RL policies. DE therefore optimizes exploration strategies so that the framework is able to explore and efficiently exploit optimal solutions in dynamic environments. The study holistically integrates these approaches in order to propose a complete solution that exceeds the capabilities of the state of the art in supporting reliable, efficient and adaptive WSN operations. The block diagram of the proposed study is given in Fig. 1 . Fig. 1. Open in a new tab Block diagram of the proposed study. Data collection In this study, Wireless Sensor Network Data 30 is used to study energy efficient fault resilience as well as topological stability in WSN. This was a Kaggle dataset containing a rich collection of records of sensor node activity (communication patterns, node mobility, network topology changes, etc.). Because of these features it is especially well suited for studying the critical problems of WSNs such as maintaining operation stability as conditions changes dynamically and optimizing energy to increase network life time. A dataset containing the status of sensor node, communication and energy parameters is sampled, all important parameters for the evaluation of the performance of WSNs in different operating scenarios. But when topology related data is included it leads to modelling of real-world scenarios as node failures, mobility patterns and the dynamic network restructuring. These are the important attributes that are required to achieve realistic environments where the EvoGenRL framework is able to be trained and tested. Feature List. Node ID: Unique identifier for each sensor node. Time: Timestamp of the recorded sensor activity. Energy (mJ): Remaining energy level of the sensor node. Packets Sent: Number of packets transmitted by the node. Packets Received: Number of packets successfully received. Status: Operational state of the node (Active/Failed). Signal Strength (dBm): Received signal strength indicator (RSSI). To pre-process, the data is filtered and organized to ensure relevance and usability of the dataset. Removing anomalies, standardizing values and segmenting dataset into training and testing subsets. Using the dataset, the study makes use of authentic and diverse sensor network activities to adapt to real world settings with the framework. Additionally, the dataset is integrated into a simulation environment (e.g. NS-3) to simulate dynamic fault scenarios, as well as to test energy efficient communication strategies. An advantage of such a structured data collection method is establishment of a solid basis for evaluation of the proposed framework’s capabilities. Sample dataset is given in Table 2 . Table 2. Sample dataset. Node ID Time Energy (mJ) Packets Sent Packets Recv Status Signal (dBm) 001 10:00 85.3 120 115 Active −70.5 002 10:01 72.1 98 90 Active −68.2 003 10:02 65.0 110 102 Failed −75.1 004 10:03 95.2 140 138 Active −67.8 005 10:04 58.6 85 80 Active −71.3 Open in a new tab Data pre-processing Pre-processing of data is the first step in data analysis consisting of cleaning, transforming, and preparing raw data into a workable format. It encompasses processes like filling missing values, normalizing the data, deleting duplicate values, and encoding categorical variables so that the dataset becomes accurate, consistent, and model-ready or ready for analysis. Remove duplicates Duplicate entries can corrupt the integrity of the dataset and bias analysis. In Wireless Sensor Network data, duplicates are recognized as records that have equal values for Node ID, time, and other attributes such as energy levels or packet numbers. Elimination of duplicates guarantees each data point makes a distinct contribution to analysis. Normalize/Standardize features Normalization scales data into a [0,1] range and standardization bring data around the mean but with unit variance. Normalization to fair contributions during model training is ensured for energy levels, packet counts, and signal strength. It is shown in Eq. ( 1 ). 1 Where, : The feature’s initial value, such as the energy level, packet count, or signal intensity. : The feature’s lowest value inside the dataset. : The feature’s maximum value inside the dataset. : A normalized value that has been rescaled to fall between 0 and 1. Energy levels Negative values or excessively high values of energy levels signify errors. Handle such values by capping or imputing them within sensor capacity limits. It is shown in Eq. ( 2 ). 2 where, : The dataset’s initial energy level value. : The lowest permitted energy value (e.g., 0 mJ). : The highest permitted energy value (e.g., 100 mJ, depending on sensor specifications). : Modified energy value upon capping within acceptable bounds. Generative adversarial network (GAN) GAN are an influential family of neural networks that aim to produce novel data samples that are similar to a provided dataset. They comprise two parts: Generator ( ): The generator uses random noise vectors to produce synthetic data. Its objective is to generate data that is identical to the original data. Discriminator ( ): It assesses the veracity of data by differentiating between authentic and artificial data samples produced by . An adversarial framework is produced by the interplay of and , in which the generator aims to deceive the discriminator while the discriminator increases its precision in differentiating between actual and fake data. Structure of the GAN is given in Fig. 2 . Fig. 2. Open in a new tab GAN structure. Figure 2 shows GAN-based framework for sensor fault detection in WSNs. The Generator accepts random noise and generates synthetic sensor data, which emulates possible fault scenarios. On the other hand, the Discriminator receives both the real sensor fault dataset and the data generated by the generator, learning to differentiate between real and fake samples. During training, the Generator adapts to generate realistic data in order to try to fool the Discriminator, while the latter continuously updates itself in order to be able to detect more accurately the fake data. This adversarial process will allow the model to create diverse fault scenarios that provide robust training to downstream tasks such as reinforcement learning of fault resilience and energy optimization in WSNs. Generator ( ) The generator creates a data instance that replicates actual fault conditions in WSNs by taking a random noise vector ( ) from a latent space. It is shown in Eq. ( 3 ). 3 Where: : A Gaussian distribution sampled from a noise vector . : The generator’s trainable parameters. The created synthetic fault data is known as . The generator has been trained to reduce the discriminator’s capacity to accurately identify as synthetic. The generator’s loss function is shown in Eq. ( 4 ). 4 Where z: Random noise vector (latent input), : Generated (fake) data sample, : Generator network, : Discriminator network (outputs probability of real data), : Penalty when discriminator correctly identifies fake data. By iteratively minimizing this loss, the generator learns to create increasingly realistic data. Discriminator (D) To ascertain their legitimacy, the discriminator assesses both synthetic data ( ) and actual data ( ). The probability score, , which indicates the likelihood that is real, is the output it produces. The goal of the discriminator is to increase the probability that and will be accurately classified. Its loss function is shown in Eq. ( 5 ). 5 Where, the distribution of actual WSN fault scenarios is shown in . : The output probability of the discriminator that is real. To optimize this loss and raise its classification accuracy, the discriminator modifies its parameters ( ). GAN are useful in generating datasets to perform fault detection and energy optimization to WSN. The Generator and the Discriminator are two neural networks that make up GANs which interact with each other in an adversarial setting to generate realistic artificial data that resembles the actual fault conditions in WSNs. The Generator uses a mapping of random noise vectors to fault-like examples to create artificial fault data. These examples include errors like node malfunctioning, energy consumption, and failure of transmission. Meanwhile, the Discriminator determines whether the data is genuine or not authentic to provide feedback to the Generator in order to optimize its performance. Through repeated training of these networks, GANs produce realistic and diverse fault datasets. With this augmented dataset, the RL model can generalize well across a large set of scenarios, enhancing its capacity to manage unforeseen faults. GAN integration therefore increases the reliability and resilience of WSNs by emulating sophisticated real-world environments for training. Reinforcement learning (RL) RL allows systems to learn adaptive decision-making policies by interacting with the environment. In WSN, RL is applied to learn energy consumption and fault recovery policies dynamically. This method formulates the problem as an MDP, a mathematical model for sequential decision-making under uncertainty. Markov decision process (MDP) components State ( The WSN’s current state is represented by the state . It captures essential network characteristics to give the RL agent enough data to make decisions. It is shown in Eq. ( 6 ). 6 Where, : The network nodes’ energy levels, which represent their present battery condition. Packet Delivery Ratio, or PDR, is a crucial indicator of communication dependability. : Current fault states, which show the kind and existence of faults (e.g., communication failure, node failure). Together, these characteristics reflect the WSN’s dynamic environment. RL structure is given in Fig. 3 . Fig. 3. Open in a new tab RL structure. Action ( ) The agent selects actions to optimize network performance based on the current state. It is shown in Eq. ( 7 ). 7 Where, : To get around errors, it reroutes data packets to alternate paths. In order to save energy, puts nodes in a low-power state. : To guarantee dependable communication, lost or malformed packets are sent again. Every activity has a trade-off, whether it be in terms of energy usage, latency, or the efficiency of fault recovery. Reward ( The success of an activity in achieving goals like fault recovery and energy efficiency is measured by the reward function. At time , the prize is provided by Eq. ( 8 ). 8 Where, : Lowering energy use and encouraging energy-saving behaviors. Packet Delivery Ratio, or PDR, promotes activities that increase the dependability of communication. : The amount of time required to recover from a fault, punishing delayed answers. Weights assigned to balance the relative relevance of the metrics are denoted by , , and . The RL agent is directed by this reward to do behaviors that optimize network performance as a whole. Policy learning The optimal policy , which specifies the likelihood of choosing action an in state , is learned by the RL agent. This policy is updated iteratively using advanced RL algorithms like Deep Q-Learning and Actor-Critic techniques. Q-Learning update rule The agent estimates the value of state-action pairs using a Q-function. It is shown in Eq. ( 9 ). 9 Where, : The expected cumulative reward expressed as the Q-value of the current state-action pair. : Instantaneous compensation for action an in state . Discount factor, weighing immediate against long-term benefits ( ). : Learning rate, which establishes the update step size. The agent explores and exploits the environment to continuously update . Actor-critic methods Two neural networks are combined in actor-critical algorithms: Actor: To enhance action selection, the policy is directly updated. Critic: Assesses the policy by calculating the expected reward for a state , represented by a value function . The actor uses gradients obtained from the critic’s assessment to adjust its policy parameters ( ). It is shown in Eq. ( 10 ). 10 Where, is the temporal difference (TD) error. RL plays a key role in designing an adaptive decision-making system for energy-efficient and fault-tolerant WSN. Through modeling the system as a Markov Decision Process (MDP), RL facilitates dynamic optimal policy learning from real-time network operational environment feedback. The RL agent takes input in the form of the present WSN state, such as node energy levels, packet delivery ratio, and fault states. It then chooses actions such as route optimization, sleep mode, or retransmission to attain desired results. A reward function scores the performance of such actions, weighing goals like energy saving, better fault recovery time, and high-quality communications. Sophisticated RL algorithms such as Deep Q-Learning or Actor-Critic methods guarantee the agent effectively explores and exploits its domain to acquire the optimal actions under different scenarios. This leads to a strong and scalable WSN system to make real-time energy-aware and fault-tolerant decisions. Differential evolution (DE) An effective evolutionary approach for optimizing complex functions is called DE. DE is used to adjust important hyperparameters in RL, including the exploration rate ( ), discount factor ( ), and learning rate ( ). These hyperparameters have a direct impact on the RL agent’s efficacy and rate of convergence. Differential Evolution Flowchart is given below in Fig. 4 . Fig. 4. Open in a new tab Differential evolution flowchart. Initialization The optimization procedure starts with initializing a population of potential solutions, each referred to as a candidate. A candidate is a possible set of hyperparameters and is represented as a vector. It is shown in Eq. ( 11 ). 11 Here: : The learning rate that establishes the Q-function update step size. : The discount factor that governs the agent’s inclination toward upcoming incentives. : The exploration rate that controls the trade-off between investigating novel approaches and taking advantage of established tactics. Values within predetermined limits for each hyperparameter are randomly sampled to establish an initial population that is sufficiently diverse. Mutation In order to efficiently explore the solution space and escape local optima, mutation adds variation to the population. A mutant vector is created for every candidate solution . It is shown in Eq. ( 12 ). 12 Where, , , : Independent solutions chosen at random from the population ( ). : Mutation scaling factor, which regulates the perturbation’s amplitude ( ). By ensuring that is impacted by the variations among current solutions, this operation encourages diversity and investigation. Crossover DE uses a crossover operation to merge the initial candidate solution, , with the mutant vector, , in order to strike a balance between exploration and exploitation. A trial vector is the end result. It is shown in Eq. ( 13 ). 13 Where, : Crossover rate ( ), which calculates the likelihood of choosing elements from the mutant vector. For each dimension , is a randomly generated value ( ). By enabling to inherit traits from both and , the crossover operation makes it easier to take advantage of promising solutions. Selection Following the creation of the trial vector , a fitness function is used to evaluate its quality. The cumulative reward the RL agent receives after training with the specified hyperparameters could serve as the fitness function in the context of RL hyperparameter optimization. It is shown in Eq. ( 14 ). 14 Where, : The reward received at time . The next generation keeps the candidate solution with the greater fitness value (either or ). By doing this, the population is guaranteed to evolve toward hyperparameter combinations that perform better. Iterative optimization The DE process (mutation, crossover, selection) is iterated over many generations. At every iteration, the population converges to optimal or near-optimal hyperparameter values. The solution at the end is an aggregate of hyperparameters that: Maximizes the RL agent’s cumulative reward. Reduces the agent’s convergence time. Improve the overall performance in terms of learning adaptability and stability. DE is crucial for optimizing RL agent hyperparameters in the energy-efficient fault-tolerant WSN framework to enable efficient training and improved decision-making. It works by iteratively improving a population of candidate solutions, where every solution corresponds to RL hyperparameters such as learning rate ( ), discount factor ( ), and exploration rate ( ). DE utilizes three primary steps: mutation, to generate mutant vectors by mixing solutions; crossover, to blend mutants and current solutions; and selection, to preserve the fittest solutions as per a fitness function, e.g., cumulative reward. By repeatedly optimizing these hyperparameters, DE facilitates quicker convergence of the RL agent, maximum energy efficiency, and improved fault recovery. Its capacity to explore large search spaces renders it invaluable in the pursuit of resilient WSN performance. The algorithmic complexity of the proposed framework is still moderate, since both RL policy updates and evolutionary searches operate in linear time with respect to the number of state features and candidate solutions, correspondingly. A brief representation will be utilized in the RL component and will consist of residual energy, packet-delivery ratio, neighbor count, queue load, and delay to ensure that it has all information required about the network. A balanced reward system encourages energy conservation, consistency of delivery and low latency and adds slight punishments of unsteadiness. This is permitting predictable and consistent learning conduct. All the RL hyperparameters are fixed and clearly defined: learning rate, discount factor, and exploration rate. For the evolutionary component, Differential Evolution (DE) is applied for fine-tuning RL parameters using a fixed population size, mutation factor, crossover rate, and generation count. Such an approach is able to ensure a good compromise between performance and computational overhead. Algorithm 1 Algorithm for Evolutionary Reinforcement Learning Framework. Figure 5 shows the proposed EvoGenRL approach for optimizing Wireless Sensor Networks (WSNs) using GAN, RL, and DE. The process starts with data collection from the WSN, including node energy levels, packet delivery ratios, and fault conditions. This data is fed into the GAN module. In the GAN module, the generator generates synthetic fault scenarios, which are checked using real-world fault data by the discriminator to provide realistic training datasets for RL. The augmented dataset is passed to the RL agent, formulated as a Markov Decision Process (MDP) with states (e.g., energy, packet delivery ratio, fault states), actions (e.g., route change, retransmit), and rewards that consider energy efficiency and fault recovery. The RL agent is an iterative reward-based framework of learning optimal policies, whereas the DE optimizer is a hyperparameter (e.g., learning rate, exploration rate, etc.) optimization that attempts to achieve convergence and improved performance. DE is used to mutate, crossover and select candidate solutions to identify the best values of hyperparameter. The model that is trained by using the RL is then implemented in the WSN environment to allow real-time fault-tolerant and energy-efficient functioning. The loop is dynamic in the way it responds to the network changes to maintain stable, scalable and optimal networks. Fig. 5. Open in a new tab Flowchart of the proposed study. Result and discussion The resulting Evolutionary Reinforcement Learning Framework was implemented in Python with the utilization of TensorFlow to perform reinforcement learning and the use of SciPy to perform the optimization of differential evolution. The structure uses GANs to generate different fault conditions, which encourages the stability of training. Application of DE to optimize hyperparameters to RL, i.e., learning rate and exploration rate, made the agents more flexible with respect to adaptation in changing environments. The RL agent real-time decision-making, based on the metrics, including the energy consumption and the time to resume fault conditions, can be included in the goal of the framework, which is the resilience of the WSN and harness energy. Using the Python libraries to conduct efficient computation, the study demonstrates the possibility of applying machine learning methods to ensure the sustainable functioning of fault-tolerant WSN, as well as provides additional opportunities to enhance adaptive networks to better. To test the utilization of the proposed protocol, an initial energy of 2 J/node was chosen in this study because it is comparable to popular WSN simulation parameters in NS-3 and previous research and has enough capacity to produce realistic energy depletion, routing behavior, and lifetime differences between protocols. Experimental setup The experimental evaluation of the proposed EvoGenRL framework was carried out using a controlled network simulation environment, standardized hardware configuration, and statistically validated performance measurements. All experiments were executed with fixed seeds for reproducibility, and each configuration was independently repeated multiple times to ensure statistical confidence. Simulation environment The experiments were performed using the following integrated simulation and learning environment: Network Simulator: NS-3 (Version 3.39). Programming Framework: Python 3.9. Deep Learning Library: TensorFlow 2.10. Optimization Library: SciPy 1.9. Dataset Processing Tools: NumPy 1.23, Pandas 1.5. The NS-3 simulator was used for packet-level wireless network simulation, while RL, GAN, and DE modules were executed externally in Python and interfaced with NS-3 during training. Hardware & machine specifications All simulations were conducted on a dedicated workstation with the following specifications. Hardware Specifications are shown in Table 3 . Table 3. Hardware specifications. Component Specification CPU Intel Core i7-11700 F (8 cores, 2.5 GHz base, 4.9 GHz boost) GPU NVIDIA RTX 3060 (12 GB GDDR6) RAM 32 GB DDR4 Storage 1 TB NVMe SSD Operating System Ubuntu 20.04 LTS CUDA Version 11.6 Open in a new tab The GPU was used for training GAN and RL networks, while NS-3 simulations ran on CPU. Simulation parameters Table 4 summarizes all fixed simulation parameters used across experiments. Table 4. Simulation parameters. Parameter Value Number of Sensor Nodes 50, 100, 150, 200 Simulation Area 500 × 500 m² Transmission Range 50 m Initial Energy per Node 2 J Packet Size 512 bytes Fault Probability 0.1, 0.2, 0.3 Simulation Time 1000 s MAC Protocol IEEE 802.15.4 Routing Baselines LEACH, DRLSS, IEE-DLSA, HOA-WSN GAN Training Epochs 1000 epochs RL Discount Factor (γ) 0.90 RL Learning Rate (α) Optimized by DE RL Exploration Rate (ε) Optimized by DE DE Population 20 DE Generations 50 DE Mutation Factor (F) 0.7 DE Crossover Rate (Cr) 0.8 Open in a new tab Number of runs & statistical validation To ensure statistical reliability, every experimental configuration was executed 30 independent times using distinct random seeds. Metrics Aggregation. For each metric, the following were computed across 30 runs: Mean. Standard deviation. 95% Confidence Interval using the 2-tailed Student t-distribution. Minimum & maximum values given in (15) 15 Where, , : Sample mean, : t -value at 95% confidence (df = 29), : Sample standard deviation. Statistical significance tests To determine whether EvoGenRL significantly outperformed baseline techniques. Table 5 shows the statistical Significance Tests. Table 5. Statistical significance tests. Test Condition Purpose Shapiro–Wilk Normality Test p > 0.05 To check distribution of paired differences Paired t-test If normal Compare EvoGenRL vs. baseline Wilcoxon Signed-Rank Test If non-normal Non-parametric alternative p-value Threshold α = 0.05 Significance decision Effect Size Cohen’s d/Rank Biserial Practical significance Open in a new tab Performance outcome Figure 6 shows energy usage patterns over time in a wireless sensor network. The curve depicts a smooth decrease in usage of energy over time, reflecting the efficacy of energy optimization mechanisms. The data points indicate the levels of energy at given steps of time with the steep initial decline leveling over time, reflecting improvements in sustainability. Fig. 6. Open in a new tab Energy consumption trends over time. Figure 7 presents the fault recovery timeline in three major steps namely fault detection, fault isolation and recovery. The fault detection takes 1.5 s and fault isolation, which is the slowest, takes 1.75 s. The quickest is fault recovery and it takes 0.8 s. The graph identifies the success of the suggested system in the process of identifying, isolating, and recovering faults within the shortest time possible. The individual color of each bar graphically reflects the difference in time between these critical stages of the system and the fact that the system has an effective process of fault management. Fig. 7. Open in a new tab Fault recovery timeline. Figure 8 represents the trend of network lifetime behaviour in active nodes depending on the time (cycles) of a WSN. Since the number of active nodes at the first 100 cycles is 100, it reduces linearly to 75 at the 50 th cycle which is the failure of nodes due to energy depletion or faults. The steady decreasing pattern means that the pace of degradation is consistent and there are no adaptive fault tolerance and energy saving strategies implemented. This pattern points at the necessity to employ efficient management techniques, such as the proposed EvoGenRL framework, to sustain node activity, enhance the network lifetime, and enhance the overall WSN reliability and performance. Fig. 8. Open in a new tab Network lifetime progression. Figure 9 named Packet Delivery Dynamics shows the Node ID versus Packet Delivery Rate with a color scale that represents the Reliability Indicator. The rate of packet delivery increases with an increase in node ID that runs between 1 and 5 and between 10 and 30. The intensity and size of color of the bubbles is indicative of increased reliability of the nodes with good delivery rates. An example is node 5 that has the highest delivery rate and reliability. This trend implies that certain nodes are more efficient and trusted in the process of forwarding packets perhaps due to their better placement, energy levels, or network state and hence the need to have adaptive routing processes. Fig. 9. Open in a new tab Packet delivery dynamics. The topology stability of a WSN is illustrated in Fig. 10 in which five nodes (1 to 5) are linked in a closed loop fashion. Every node is connected to two neighboring nodes which offer redundant paths and enhance reliability of communication. This topology is stable and symmetrical, and it cannot partition the network in the case of the failure of one node. The architecture provides resiliency and data availability. The equally spaced links, also enable equal energy consumption resulting in a long life of the network and consistent performance during dynamic or failover operation- the main goal of smart network architectures including EvoGenRL. Fig. 10. Open in a new tab Topological stability. Figure 11 demonstrates the learning of a reinforcement learning agent during 10 episodes in terms of training reward. The level of performance of the agent starts at about 15, then slowly rises to nearly 100 after episode 10. This is a good trend that indicates effective learning and convergency of policy. The consistent rise in reward is a sign that the agent is successfully altering its behavior to provide greatest long-term rewards, one of the key ideas of the reinforcement learning. This is the kind of performance that warrants the power of the applied training policy, e.g. in the example of the EvoGenRL framework of maximizing energy efficiency and fault tolerance in WSNs. Fig. 11. Open in a new tab Policy learning curve. Figure 12 indicates that three of the most vital reinforcement learning variables, which include Gamma, Learning Rate, and Epsilon are in the process of being optimized. The chart indicates that the value of Gamma, controlling the discounting effect over the future rewards, is the most significant, indicating that optimization of rewards at long term is given more attention. Learning Rate and Epsilon is on the other hand low and is indicative of conservative updates and scant exploration. Such a hyperparameter configuration is likely the result of DE optimization with EvoGenRL framework with the aim of convergent stability and policy learning resilience in Wireless Sensor Networks. The diagram identifies the importance of proper hyperparameter tweaking towards effective and robust agent activity. Fig. 12. Open in a new tab Hyperparameter evolution. Figure 13 illustrates the energy usage (in joules) across five WSN nodes. The highest consumption at approximately 3.5 J is exhibited by Node 4 followed by Nodes 2 and 5 at approximately 3.0 J and the lowest consumption is at approximately 2.3 J at Node 1. This mismatch may be attributed to variation in the work load, routing tasks or node locations. The nodes with very high energy may trigger untimely failures, thus reducing network lifetime. These observations are the reason why there has to be energy-aware policies like the ones in the EvoGenRL framework to ensure that each resource is equally utilized. Fig. 13. Open in a new tab Energy consumption per node. Figure 14 shows the percentage of faults per node of a WSN. The highest number of faults has been caused by Node 5 of 30% and Node 2 of 23.3%. The number of faults in node 1 and 3 is less, 16.7% and 10 respectively. This imbalance means that some nodes are more vulnerable- maybe with workload, location or connectivity. These trends can be determined by asset identification, which makes it possible to carry out the actions of fault resilience aimed at enhancing the stability of the network and expanding the duration of its operation, as addressed by the framework of EvoGenRL. Fig. 14. Open in a new tab Fault distribution across nodes. Frequency distribution of packet loss in a WSN is depicted in Fig. 15 . Most of the measurements were 0 to 2 packet losses, which indicates fairly consistent communication. Nevertheless, losses (4–6 packets) more rarely occurred, but require attention, and point to a non-stable situation or interference sometimes. The even distribution within lower values and the spikes of the extremes indicate variation in the network performance. This analysis is vital in identifying weaknesses areas and also in optimization of routing or retransmission strategy. Frameworks like EvoGenRL aim at reducing these losses by reducing fault tolerance and providing reliable communication. Fig. 15. Open in a new tab Data packet loss analysis. Figure 16 shows Energy Utilization Efficiency for five different Nodes (Node1 through Node5). Base Energy indicates the natural energy consumed at every node. Optimization Gain indicates extra energy gained or conserved as a result of some optimization procedure. Node2, for example, is seen to have the maximum amount of total energy, with most of it being due to optimization gain, whereas Node1 has the minimum total energy. The chart clearly indicates how optimization adds to the total energy consumption at any node. Fig. 16. Open in a new tab Energy utilization efficiency. Figure 17 shows how the number of Connected Nodes varies over Time (cycles). The y-axis is the number of connected nodes from 90 to 100, and the x-axis is time in terms of cycles from 0 to 50. The graph shows a step-by-step reduction in connected nodes. There are 100 connected nodes at time 0 initially. Thereafter, this number stabilizes for a duration, before declining to 98 at about 10 cycles, then 96 at about 20 cycles, and so on. The connectivity decreases in discrete steps progressively, which signifies the loss of nodes with time. At 50 cycles, the connectivity in terms of connected nodes has decreased to 90. Fig. 17. Open in a new tab Node connectivity over time. Transmission Delay per Node of Transmission Delays of the five Nodes (Node1 to Node5) is visualized in Fig. 18 . The box plots show the distribution of the transmission delays of a particular node. The middle delay is denoted by the horizontal line in every box. The box itself covers the interquartile (IQR), the 25 th percentile to the 75 th percentile. The whiskers run to minima and maxima values that lie in 1.5 times the IQR. The median delay of Node3 is the largest, and the range is the largest, indicating more fluctuating and mostly higher delays. The median delay is lowest in node2 and the distribution is tighter that means that there are more consistent and shorter delays in node2. The plot effectively compares the central tendency and spread of transmission delays across the different nodes. Fig. 18. Open in a new tab Transmission delay per node. Figure 19 titled Training and Testing Loss Over Epochs illustrates the performance of a machine learning model. Both lines show a decreasing trend, indicating that the model is learning and its error is reducing over time. Training loss consistently stays below testing loss, which is typical. As epochs increase, both losses continue to drop, suggesting improved model performance. The gap between the two curves indicates some degree of generalization, with the testing loss showing how well the model performs on unseen data. Fig. 19. Open in a new tab Training and testing loss over epochs. Performance metrics Energy consumption (EC) Energy Consumption refers to the total energy used by all nodes in the WSN during its operation. It includes energy for data transmission, reception, and processing. Lower energy consumption indicates a more energy-efficient network. It is shown in Eq. ( 16 ). 16 Where, : The energy required to transmit data. : The energy required to receive data. : Energy expended during data processing. There are a total of nodes. Network lifetime Network Lifetime measures the duration until the first node in the WSN exhausts its energy. Maximizing lifetime ensures longer and more reliable network operation. This metric highlights the sustainability of the network under given workloads. It is shown in Eq. ( 17 ). 17 Where, the node’s initial energy is represented by . Average Energy Consumption per Unit of Time (EC average) = . Packet delivery ratio (PDR) PDR is the ratio of data packets successfully delivered to the destination to the total number of packets sent. It evaluates the reliability of the network. It is shown in Eq. ( 18 ). 18 Where, : The quantity of packets that the destination received. : The total quantity of packets transmitted. Latency Latency is the average time taken for a data packet to travel from the source to the destination. Lower latency implies faster communication. It is shown in Eq. ( 19 ). 19 Where, : The time at which a packet arrives. : The time when a packet was transmitted. : The total quantity of packages. Throughput Throughput is the total volume of data successfully transmitted across the network in a given time frame. It indicates network capacity. It is shown in Eq. ( 20 ). 20 Where, : Total data successfully transmitted. : Total operating time. The performance metrics of the proposed method in Table 6 highlight its efficiency in WSN. EC is minimized at 2.2 J, demonstrating effective resource management and energy-efficient operations. The Network Lifetime reaches 1700 cycles, indicating prolonged operational periods before node depletion, crucial for sustainable WSN performance. A high PDR of 99.7% ensures reliable data communication, reflecting robust fault recovery and routing mechanisms. The Latency of 3.2 ms showcases swift data transmission, critical for real-time applications. Lastly, a Throughput of 350 kbps highlights the system’s capability to handle substantial data traffic effectively. These results collectively underscore the robustness and efficiency of the proposed framework in enhancing WSN performance. The visual representation is given below in Fig. 20 . Table 6. Performance metrics of the proposed study. Metric Name Highest Value Energy Consumption (EC) 2.2 J Network Lifetime 1700 cycles Packet Delivery Ratio (PDR) 99.7% Latency 3.2 ms Throughput 350 kbps Open in a new tab Fig. 20. Open in a new tab Performance metrics of the proposed study. The comparative performance analysis in Table 7 highlights the superiority of the proposed method over existing approaches. In terms of Energy Consumption (EC), the proposed method achieves a minimal value of 2.2 J, significantly outperforming methods like HybeMetaWSN (40 J) and HOA-WSN (34.4 J). The Network Lifetime is extended to an impressive 1700 cycles, far exceeding HOA-WSN’s 80.6 cycles. For Packet Delivery Ratio (PDR), the proposed method delivers an exceptional 99.7%, showcasing higher reliability compared to IEE-DLSA (93.131%) and DRLSS (53%). Furthermore, Latency is minimized to 3.2 ms, which is faster than DRLSS (4.2 ms) and HOA-WSN (6 ms). Lastly, the Throughput is boosted to 350 kbps, a remarkable improvement over IEE-DLSA (93.615 kbps). These results demonstrate the effectiveness of the proposed method in achieving energy efficiency, extended network lifetime, enhanced reliability, reduced latency, and improved throughput in Wireless Sensor Networks (WSNs). Table 7. Comparison of the proposed method with existing methods. Methods Energy Consumption (EC) Network Lifetime Packet Delivery Ratio (PDR) Latency Throughput HybeMetaWSN 19 40 J 75.6 40% 7.1 38 IEE-DLSA 22 7.644 N/A 93.131 N/A 93,615 HOA-WSN 19 34.4 80.6 46 6 45 DRLSS 19 28.4 75.6 53 4.2 52 EEDC: Tier-based RHCER 28 −31% (energy, relative) N/A + 38% (packet-drop reduction, relative) N/A N/A fuzzy neural networks 31 9.75 mJ 5250 cycles 98.5% 0.019 s (19 ms) 0.98 Mbps Proposed Method 2.2 J 1700 cycles 99.7% 3.2 ms 350 kbps Open in a new tab This comparative analysis shows that the uniqueness of the suggested framework is rather obvious. Although the traditional WSN routing plans either maximize one parameter or use a set of heuristics, the new approach will incorporate GAN-based fault scenario optimization, DE-tuned reinforcement learning, and adaptive decision optimization. The results are the lowest energy consumption of 2.2 J, maximum network lifetime of 1700 cycles and optimum PDR of 99.7, as shown by the table. These enormous latency and throughput gains further corroborate the fact that a hybrid EvoGenRL framework delivers a better but most importantly a holistic solution to the current methods. According to the simulation outcomes, the suggested EvoGenRL framework offers an efficient solution to WSNs energy efficiency, fault tolerance, and topological stability. The energy consumption curve demonstrates a much slower depletion curve than the state of the art approach, which implies that learning of the RL agent to optimize routing and node-sleep strategies in reducing unnecessary transmissions is successful. It is also supported by the use of DE-tuned hyperparameters, which enable the convergence on training to be much faster and more stable. This results in increased energy reserves in the network that explains relatively long lifecycle of the network of 1700 cycles. There is also a significant improvement in the fault-recovery performance. The different types of fault scenarios that the GAN fault scenarios subject it to throughout the training process assist it in responding fast and correctly to fault scenarios in real operations. The ability of the framework to manage failures proactively is proven by the fact that a decrease in the time spent on fault detection and isolation, as shown in the simulation results. The packet delivery ratio of 99.7 is high which indicates good recovery of node and link failures and consistent data flow even in the degrading network conditions. The topology related results include increased stability in routing path represented by the reduced number of route changes with time. This is because of the RL policy, which solely chooses all stable links in an energy-efficient manner and does not undertake unnecessary re-routing. A more stable routing topology minimizes the latency, control overhead and the workload of nodes leading to higher reliability. Moreover, the enhanced 350 kbps throughput is a positive indication of the ability of the framework to maintain efficient communication channels, even when overloaded by heavy traffic or when performance is spoiled by faults. Overall, the findings of simulations can be used to confirm the strength and flexibility of the EvoGenRL framework. The system is constantly better than the state-of-the-art approaches in all the main performance indicators through the addition of GAN-based scenario generation, DE optimization, and RL-driven decision-making. These gains indeed confirm that EvoGenRL can enable resilient, energy-aware, and scalable WSN operations that are appropriate for real-world deployments. Discussion An Evolutionary Reinforcement Learning framework, EvoGenRL, which outperforms existing frameworks in all the metrics evaluated for different values of the problem parameters, provides a complete solution to the problems facing WSNs. Guided by the design of optimal deployment frameworks for several types of RF and optical transceivers, it integrates some advanced techniques like RL, DE, and GAN to maintain a leading energy efficiency, network reliability, and communication performance balance. In comparison to HybeMetaWSN (40 J) and HOAWSN (34.4 J), the proposed framework shows a remarkably low (i.e. 2.2 J) energy consumption. The reduction is due to RL based optimization of node activity and DE adjusted hyperparameters jointly reducing unnecessary energy expenditure. Its network lifetime is extended to 1700 cycles, rather than the closest competitor HOA-WSN (80.6 cycles). The reason for this extended longevity is the dynamic energy balancing strategies enabled by the RL policies which aware the node workloads and thus avoid premature failures. Another domain where EvoGenRL stands out is in reliability with the Packet Delivery Ratio (PDR) set at 99.7%, while HOA-WSN (46%) and DRLSS (53%) are comparatively low. In order to perform this remarkable, our framework exploits GANs to learn the various fault scenarios during training such that it can adapt to the real world challenges such as node failures and packet losses. The latency is reduced to 3.2 ms, compared to competing systems such as DRLSS and HOA-WSN (4.2 ms and 6 ms respectively). RL’s efficient routing decisions that prefer low delay routes whereas preventing data degradation leads to this improvement. It boosts the throughput up to 350 kbps, an improvement by a drastic proportion than methods such as IEE DLSA (93.615 kbps) and DRLSS (52 kbps). Such capability of the framework to dynamically optimize the communication protocols to achieve the highest possible data transmission even in the presence of high traffic loads is what they assign this enhancement to. Overall, the EvoGenRL framework provides a systemic and innovative solution to constrain the limitations of existent methods revealing itself able to provide a sustainable and efficient solution for next generation WSNs. Conclusion and future works Energy efficiency, fault resilience and communication performance for Wireless Sensor Networks (WSNs) were optimized in this study using the proposed Evolutionary Reinforcement Learning (EvoGenRL) Framework. A framework is proposed, based on Reinforcement Learning (RL), Differential Evolution (DE) and Generative Adversarial Networks (GANs), that exhibited dramatic improvements over prior methods. This consumed 2.2 J of energy, extended the network lifetime to 1700 cycles, increased the packet delivery ratio (PDR) to 99.7%, lowered latency to 3.2 ms and increased throughput to 350 kbps. These findings confirm the effectiveness of EvoGenRL to face key issues in WSNs and establish a benchmark for subsequent advances. The strengths of the framework are not without limitations. Second, it may have too high computational complexity for effective deployment in real time in resource constrained environments. Second, the training process requires massive computing power (including GAN fault simulation). As future work, integrating this framework with real world IoT platforms for practical validation and leveraging it with lightweight versions of EvoGenRL for deployment in ultra-low power devices. Moreover, improving its applicability, multi objective optimization approaches to handle trade-off between energy efficiency, data integrity and latency, can also be pursued. There is an open window for further researches to proceed on the base of work with this IS presented, as well as to introduce the innovations in intelligent and sustainable management of WSN. Author contributions Lakshmi S: Conceptualization, Methodology, Data Curation, Writing the Original Draft. Aswath S: Software, Validation, Formal Analysis, Visualization. Swaminathan A: Investigation, Resources, Writing the Review & Editing. Anandakumar Haldorai: Supervision, Project Administration, Funding Acquisition. Data availability The datasets generated and/or analysed during the current study are available in kaggle repository.https://www.kaggle.com/datasets/halimedogan/wireless-sensor-network-data. Declarations Competing interests The authors declare no competing interests. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. References 1. Al Aghbari, Z., Pravija, P., Raj & Khedr, A. M. A robust fault-tolerance scheme with coverage preservation for planar topology based Wsn. Wireless Pers. Commun. 129 (3), 2011–2036 (2023). [ Google Scholar ] 2. Shyama, M., Pillai, A. S. & Anpalagan, A. Self-healing and optimal fault tolerant routing in wireless sensor networks using genetical swarm optimization. Comput. Netw. 217 , 109359 (2022). [ Google Scholar ] 3. Wang, Z. Reliability analysis of social network data transmission in wireless sensor network topology, J. Sens. , 2022( 1), 6256884 (2022). 4. Heidari, A., Amiri, Z., Jamali, M. A. J. & Jafari, N. Assessment of reliability and availability of wireless sensor networks in industrial applications by considering permanent faults. Concurrency Computation: Pract. Experience . 36 (27), e8252 (2024). [ Google Scholar ] 5. Dhanasekaran, S., Ramalingam, S., Baskaran, K. & Vivek Karthick, P. Efficient distance and connectivity based traffic density stable routing protocol for vehicular ad hoc networks. IETE J. Res. 70 (2), 1150–1166 (2024). [ Google Scholar ] 6. Wu, H., Han, X., Yang, B., Miao, Y. & Zhu, H. Fault-tolerant topology of agricultural wireless sensor networks based on a double price function. Agronomy 12 (4), 837 (2022). [ Google Scholar ] 7. Zhao, Z., Liu, C., Guang, X., Zhao, Z. & Qu, W. A Reliability-Driven topology restoration strategy for underwater wireless sensor networks in dynamic ocean environments. IEEE Internet Things Journal , (2024). 8. Tabatabaei, S. A Fault-Tolerant clustering approach for target tracking in wireless sensor networks. Wireless Pers. Commun. 137 (4), 2303–2322 (2024). [ Google Scholar ] 9. Mehdiyev, S. Assessing the Impact of Energy Consumption of Wireless Sensor Networks on the Fault Tolerance of Cyber-Physical Systems. 10. Xue, X. et al. A hybrid cross layer with harris-hawk-optimization-based efficient routing for wireless sensor networks. Symmetry 15 (2), 438 (2023). [ Google Scholar ] 11. Li, H., Liu, X., Huang, M., Chen, Q. & Chen, L. ICCO: immune clustering-based coverage optimisation algorithm in wireless sensor networks. Int. J. Internet Protoc. Technol. 15 (2), 76–93 (2022). [ Google Scholar ] 12. Devi, V., Kumar, R. & Kumar, V. Energy Efficient Transmission and Fault Tolerance in WSN-Based IoT Network using HS-EDEEC and Probability Based Approach, in International Conference on Integrated Circuits, Communication, and Computing Systems (ICIC3S) , IEEE, 2024, pp. 1–8., IEEE, 2024, pp. 1–8. (2024). 13. Xu, X., Tang, J. & Xiang, H. Data transmission reliability analysis of wireless sensor networks for social network optimization, J. Sens. , 2022( 1), 3842722 (2022). 14. Yigit, Y., Dagdeviren, O. & Challenger, M. Self-stabilizing capacitated vertex cover algorithms for internet-of-things-enabled wireless sensor networks. Sensors 22 (10), 3774 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Ramalingam, S., Dhanasekaran, S., Sinnasamy, S. S., Salau, A. O. & Alagarsamy, M. Performance enhancement of efficient clustering and routing protocol for wireless sensor networks using improved elephant herd optimization algorithm. Wireless Netw. 30 (3), 1773–1789 (2024). [ Google Scholar ] 16. Amron, K., Kusumawinahyu, W. M., Anam, S. & Mahmudy, W. F. Multi-Tier Topology Design of Wireless Sensor Networks using Multi-Objective Particle Swarm Optimization, in Proceedings of the 7th International Conference on Sustainable Information Engineering and Technology , pp. 103–110. (2022). 17. Priyadarshi, R. Exploring machine learning solutions for overcoming challenges in IoT-based wireless sensor network routing: a comprehensive review. Wireless Netw. 30 (4), 2647–2673 (2024). [ Google Scholar ] 18. Singh, A. et al. Resilient wireless sensor networks in industrial contexts via energy-efficient optimization and trust-based secure routing. Peer-to-Peer Netw. Appl. 18 (3), 1–17 (2025). [ Google Scholar ] 19. Prasad, K. V. & Periyasamy, S. Secure-energy efficient bio-inspired clustering and deep learning-based routing using blockchain for edge assisted WSN environment. IEEE Access. 11 , 145421–145440 (2023). [ Google Scholar ] 20. Hajar, M. S., Kalutarage, H. K. & Al-Kadri, M. O. 3R: A reliable multi agent reinforcement learning based routing protocol for wireless medical sensor networks. Comput. Netw. 237 , 110073 (2023). [ Google Scholar ] 21. Song, Y., Liu, Z., Li, K., He, X. & Zhu, W. Research on High-Efficiency routing protocols for HWSNs based on deep reinforcement learning. Electronics 13 (23), 4746 (2024). [ Google Scholar ] 22. Hussain, A. et al. Integrated Energy-Efficient distributed link stability algorithm for UAV networks. Computers Mater. Continua . 81 (2), 2357 (2024). [ Google Scholar ] 23. Udayaprasad, P. et al. Energy efficient optimized routing technique with distributed SDN-AI to large scale I-IoT networks. IEEE Access. 12 , 2742–2759 (2024). [ Google Scholar ] 24. Khan, S. B. et al. Artificial intelligence in next-generation networking: energy efficiency optimization in IoT networks using hybrid LEACH protocol. SN Comput. Sci. 5 (5), 546 (2024). [ Google Scholar ] 25. Thangavelu, A. & Rajendran, P. Energy-Efficient secure routing for a sustainable heterogeneous IoT network management. Sustainability 16 (11), 4756 (2024). [ Google Scholar ] 26. Kumar, S. P., Garg, S., Alabdulkreem, E. & Miled, A. B. Advanced generative adversarial network for optimizing layout of wireless sensor networks. Sci. Rep. 14 (1), 32139 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Alharbi, N. H. Enhancing graph-routing algorithm for industrial wireless sensor networks, PhD Thesis, University of Glasgow, (2024). [ DOI ] [ PMC free article ] [ PubMed ] 28. Gupta, D., Wadhwa, S., Rani, S., Khan, Z. & Boulila, W. EEDC: an energy efficient data communication scheme based on new routing approach in wireless sensor networks for future IoT applications. Sensors 23 (21), 8839 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Rao, A. K., Singh, B. K. & Nagwanshi, K. K. An optimized routing protocol for energy-efficient data transmission in agricultural environments using WSN-based IoT networks. Int. J. Adv. Technol. Eng. Explor. 11 (120), 1562 (2024). [ Google Scholar ] 30. Wireless Sensor Network Data. & Accessed June 03, 2025. [Online]. Available: https://www.kaggle.com/datasets/halimedogan/wireless-sensor-network-data 31. Gopalan, S. H., Takale, D. G., Jayaprakash, B. & Raj, V. P. An energy efficient routing protocol with fuzzy neural networks in wireless sensor network. Ain Shams Eng. J. 15 (10), 102979 (2024). [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Data Availability Statement The datasets generated and/or analysed during the current study are available in kaggle repository.https://www.kaggle.com/datasets/halimedogan/wireless-sensor-network-data. Articles from Scientific Reports are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (3.5 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top