This paper has been accepted to appear in the IEEE European Symposium on Security and Privacy (Euro S&P), 2026.
Automated Stealthy Wear-Out Attack on Digital Twins With Deep Reinforcement Learning Joshua Haworth⊠ , Aryan Pasikhani, George Pavlides, Prosanta Gope, John Clark
arXiv:2607.10830v1 [cs.CR] 12 Jul 2026
Department of Computer Science, University of Sheffield, United Kingdom {jhaworth1, aryan.pasikhani, p.gope, john.clark}@sheffield.ac.uk, [email protected]
Abstract—Digital Twins (DTs) have emerged as pivotal enablers of Industry 4.0, offering transformative capabilities such as real-time monitoring, advanced simulation, and precise control of physical assets. By bridging the physical and virtual domains, DTs facilitate seamless integration of data-driven decision-making and operational optimisation. However, this seamless interaction significantly expands the attack surface of industrial systems, creating vulnerabilities that adversaries can exploit. This paper introduces a novel and stealthy wear-out attack leveraging Deep Reinforcement Learning (DRL) to target DT-enabled infrastructures. The adversary strategically and covertly manipulates control signals, inducing increased torque on a specific joint to accelerate wear and tear while evading detection by a stateof-the-art anomaly detection system. Extensive benchmarking of reinforcement learning algorithms - including Twin Delayed Deep Deterministic Policy Gradient (TD3), Soft Actor-Critic (SAC), Proximal Policy Optimisation (PPO), and Advantage Actor-Critic (A2C) - revealed that SAC consistently outperformed its counterparts in terms of sample efficiency, stability, and overall attack effectiveness. We evaluate the proposed adversary in an industrial setting using the UR10e robotic arm. Results demonstrate the adversary’s ability to significantly elevate torque levels on the targeted joint, leading to accelerated degradation and increased maintenance costs, all while operating stealthily and avoiding detection. Our findings highlight the substantial risks posed by DRL-driven adversaries to DT-enabled environments and emphasise the critical need for robust defence mechanisms to protect critical industrial systems. Index Terms—Digital Twins, Reinforcement Learning, Adversarial Machine Learning
1. Introduction Striving to realise the Industry 4.0 vision represents one of the top priorities for manufacturers* . The use of industrial robots, cloud computing and vast amounts of sensor data collected throughout a product’s lifecycle are expected to improve production efficiency, flexibility and decision making, while at the same time reducing waste and minimising carbon emissions [1], [2]. Digital Twin (DT) technology is one of the vital enablers for the realisation of Industry 4.0 [2]. It provides a virtual replica of a physical entity, system or process, that gets updated in real-time according to past data, real-time * https://www.sap.com/products/scm/industry-4-0/ industry-4-0-strategy.html
sensor readings and a predefined physical model [2]. DTs seamlessly integrate and analyse physical and digital data recorded throughout a product’s lifecycle, providing new services which can be utilised to adjust how an operation is performed in the physical space according to direct orders from the virtual space [2]. This can improve the industrial process’ performance, enhance product designs and streamline operations, but also allow for Prognostics and Health Management (PHM) to detect, in a timely manner, potential asset faults and degradations, thus reducing maintenance costs [2]. Tao et al. [3] proposed a five-dimensional model of DT in which there is a clear separation between the elements involved. Specifically, the DT architecture is decomposed into Physical Entity (PE), Virtual Entity (VE), Services (Ss), DT Data (DD) and Connection (CN). As a faithful replica of the PE, the VE can be utilised to run simulations or tests of various scenarios without affecting the real-world counterpart. This is particularly beneficial for testing configuration changes before deploying them to critical physical systems or for performing security assessments on systems whose potential downtime is unacceptable [4]. DTs are usually hosted on the cloud, providing them with vast amounts of computational power, which allows them to perform complex calculations and analyses that would otherwise not be feasible on the, usually more resource-constraint, physical counterparts [5]. Whilst the benefits of DT are clear [2], [4], the two-way communication between the physical and cyber spaces, combined with the introduction of ”smart” services, broadens the attack surface and presents new security challenges to the system [5]. For example, an attacker that has infiltrated the system or compromised one or more elements of the DT model (e.g. VE, Ss, DD or CN) could manipulate the VE’s instructions to command the PE to perform hostile or unsafe actions. Such actions have the potential of harming assets in the physical space, producing defective products and resulting in significant financial losses for the manufacturer. In case the affected industrial system is part of critical infrastructure, such as a power grid, the consequences of malicious compromise can be of significant scale and even lead to the loss of human life [6]. Moreover, since the DT model includes an accurate replica of a physical object and its operational logic, unauthorised access to the DT’s elements can lead to the leak of confidential information regarding how processes are performed and sensitive Intellectual Property [7]. With the security risks that arise due to the use of DT technology, it is apparent that its operation should be safeguarded via thorough security assessments and the
design of appropriate defence measures. A considerable amount of work has been conducted on implementing defence mechanisms on the DT-side in order to detect and mitigate attacks that aim to exploit the existence of a DT to deceive the physical object into falling into an unsafe state (e.g. [5], [8]–[10]). However, in the existing literature, the adversary is non-intelligent, and their malignant modifications are constant, monotonous or probabilistic [5], [9], [11], [12], with no concern for intelligently adapting and adjusting their attack strategies based on the observed environment. In recent years, Artificial Intelligence (AI)-assisted attacks targeting industrial systems emerged (e.g. [13], [14]) that are not only highly effective in achieving their hostile objectives but also have high evasion rates against existing defence measures. In particular, Deep Reinforcement Learning (DRL) algorithms have demonstrated promising performance as an offensive tool against industrial environments [15]–[17], capable of synthesising highly potent and stealthy attack strategies. Their ability to learn how to conduct effective attacks with little knowledge of the victim system (a.k.a grey box) or potential attack approaches makes them a versatile tool for discovering new attack strategies [15]. Furthermore, as a data-driven method that depends on (partial) system observations, DRL agents are capable of discovering and exploiting complex state dynamics and inter-dependencies, making them a powerful technique for attacking convoluted systems [15], [17]. Also, the fact that DRL agents take strategic actions that aim to maximise the cumulative reward received enables them to perform better in the long run compared to alternative short-sighted attack methods [18]. So far, no published research investigates the possibility of utilising AI-based methods to exploit and attack DT-enabled infrastructure. With the increasing adoption of DT technology by world leaders in various fields like power grids, wastewater management, car manufacturing, aerospace engineering, and healthcare [2], safeguarding DT-enabled infrastructure is more relevant than ever before. Any compromise in their intended functionality has the potential of a wide negative impact on manufacturers and customers across the world. With the rising occurrence and sophistication of damaging attacks against industrial environments (e.g. the Stuxnet attack against Iran’s nuclear facilities [19], the BlackEnergy malware attack on Ukraine’s power grid † and the HatMan malware targeting Schneider Electric safety controllers ‡ ), it is apparent that industrial systems are a sought-after target by cyber criminals and need to be secured. By proactively identifying attack vectors, we can inform and facilitate the development of effective and robust defence measures [15].
1.1. Our Contribution In this paper, we employ the power of DRL to train an adversarial agent capable of conducting stealthy wear-out attacks against DT-enabled infrastructure. In our specific †
https://www.cisa.gov/news-events/ics-alerts/ir-alert-h-16-056-01 https://www.cisa.gov/sites/default/files/documents/ MAR-17-352-01%20HatMan%E2%80%94Safety%20System% 20Targeted%20Malware S508C.pdf ‡
use case, the DRL agent attempts to make subtle changes to the control signals issued by the VE for a physical robotic arm in order to gradually wear out a selected joint of it. This leads to the physical system’s faster deterioration, degrades operational efficiency and effectiveness, and increases the fault occurrence probability and the owner’s maintenance and replacement costs. In summary, this paper’s contributions are: •
•
•
We develop the first DRL-based adversary who can model the robotic arm’s safe zones of operation via the DT, and perform an intelligent wear-out attack on the targeted joint. Here, we consider a grey-box attacker by altering the access to environment observations the agent has. We demonstrate how our novel DRL-assisted adversary can autonomously manipulate robotic operations and induce wear-out effects by employing a stealthy low-and-slow attack strategy. The adversary effectively evades detection by an ensemble of Autoencoder-based anomaly detectors, highlighting its capability to exploit vulnerabilities over prolonged periods. We publicly share our implementation code along with the generated adversarial policies and results§ for reproducibility purposes and also to facilitate their use by the research community for the development of robust and resilient defence systems.
To the best of our knowledge, no work currently exists that investigates the possibility of a stealthy wear-out attack against physical assets in a DT-enabled setting via the use of DRL.
2. Related Work Digital Twins (DTs) are becoming integral to modern Industrial Control System (ICS), enabling tighter coupling between physical processes and digital operations. As industrial environments adopt greater automation and connectivity, DTs provide continuous synchronisation between virtual and physical entities, supporting real-time monitoring, prediction, and control [2], [3]. Their growing deployment across manufacturing, energy, and transportation has shifted DTs from auxiliary analytical tools to essential components of Industry 4.0 infrastructures [1], [2]. Attacks against DT-enabled infrastructure: As DTs become increasingly adopted within industrial settings, consideration needs to be given to new attack vectors introduced through bidirectional synchronisation. The existing literature focuses on adversaries that introduce predefined manipulations to the communication between the PE and VE. Early work by Eckhart and Ekelhart [20] explore state synchronisation vulnerabilities in CyberPhysical System (CPS), modelling an insider and Manin-the-Middle (MitM) adversaries capable of spoofing or modifying direct communication between the VE, and PE. The attack presented relies on the direct manipulation of values without incorporating stealth or situational adaptivity, demonstrating an initial focus on basic intrusion § All related materials, including datasets and code, are available for the research community here: https://anonymous.4open.science/ r/Stealthy-Wear-Out-RL-BC8B/README.md
Legend Anomaly detector Observation Zone Network Connection
Wrist 1 joint Wrist 2 joint Wrist 3 joint Tool flange
Controller
Switch
L1 Network
Router
Internet
Adv Firewall Adv L2 Network
DNN-assisted Anomaly Detector
Switch Database
Base Base joint Shoulder joint Elbow joint Physical Entity (UR10e Robotic Arm)
Virtual Entity
Figure 1: DRL-assisted adversary targeting an industrial robotic system by exploiting vulnerabilities in L1 (Operational Technology) or L2 (Information Technology) networks. scenarios within DT-enabled environments. Subsequent works by Akbarian et al [11], and Tärneberg et al [5] explore more advanced manipulation techniques with the additional focus of attempting to bypass Intrusion Detection System (IDS) mechanisms within the system. In [11] ramping and scaling attacks are introduced as an alternative to constant signal offset, such that scaling attacks rely on a scaling factor used to manipulate the original measurement value, and ramping attacks introduce a gradually increasing manipulation. Tärneberg et al [5] build on these attack methods with the introduction of burst attacks where the adversarial manipulations are applied in ON/OFF periods that determine when the attack occurs, with the aim to reduce detectability. Both of these works demonstrate how minor manipulations can effect the system behaviour while attempting to avoid security mechanisms, highlighting the potential consequences of data injection attacks against DT systems. Dietz et al [21] further formalise adversarial capabilities in a DT-enabled industrial systems, simulated via virtual machines, using the Dolev-Yao attacker model [22]. This assumes secure cryptographic primitives, full system knowledge, and possesses the ability to run infinite concurrent processes. Varghese et al [23] use the same basis outlined in [21], with consideration for a range of attacks, including network Denail of Service (DoS), command injection, and calculated (small factor scaling) and naive measurement modifications (constant / random modifications). More recently Ali et al [24] investigate DTs in an Electric Vehicle (EV) environment, where the attacker is limited to the manipulation of the voltage and phase angle. By injecting adjustments to the in-transit information vector the adversary aims to corrupt the system state while limiting the adjustments to avoid detection. A notable limitation of these False Data Injection Attacks (FDIAs) [5], [9], [11], [20], [21], [23], [24] is there reliance on manually tuned parameters. In order for these attacks to achieve any semblance of success, the attacker must explicitly calibrate the parameters used to
make manipulations for each system in order to balance impact and stealth. In contrast, an intelligent AI-based adversary could dynamically adapt its strategy in order to optimise impact and undetectability in response to the environment. This adaptability makes adversarial agents a prime candidate for overcoming the weaknesses of existing attacks against DT-enabled infrastructure. Additionally it should be noted that current work focuses on the deployment of security mechanisms as opposed to the use of advanced adversarial models to improve existing security mechanisms. (D)RL-based attacks against industrial systems: Beyond DT-enabled environments, DRL agents are being used to perform attacks against industrial and power grid systems. Mohamed and Kundur [15] employ DRL to synthesise attacks against load frequency control used by power grids. The adversarial agent aims to trigger unwanted protection mechanisms which can lead to sudden power imbalance, grid instability and subsequently, blackout. They demonstrate that an DRL agent is capable of generating attacks with little to zero prior knowledge of the victim system and is even able to adjust its attacks to damage different systems. Their agent executes FDIA (corrupting measurements and control signals) as well as load switch attacks to destabilise the target grid. Using the Deep Deterministic Policy Gradient (DDPG) DRL algorithm, due to its simplicity and support for continuous action and observation spaces, they were able to achieve near-optimal results in respect to the effectiveness of attacks, with the authors claiming that additional training episodes could lead to optimality. A number of studies demonstrated the use of (D)RL to synthesise attack strategies against power grid systems. In [25]–[27], the authors use (deep) Q-Learning to perform line-switching attacks that induce sudden changes in grid topology and cause cascading failures, leading to blackouts. Chen et al. [28] trained a Q-Learning With Nearest Sequence Memory agent that, with little knowledge of the victim system, is capable of performing stealthy FDIAs,
modifying system measurements to deceive existing controls in triggering power outages. In [29], a multi-agent DRL algorithm using DDPG is proposed that is capable of generating stealthy, coordinated FDIAs against microgrids, which were shown to be effective against a stateof-the-art detection scheme. Shereen et al. [16] proposed the use of Proximal Policy Optimization (PPO) DRL algorithm to discover efficient but subtle attack policies that inject false measurements in the readings transmitted to automatic generation control, causing it to issue inaccurate control commands to the generators that can have catastrophic consequences for the power system. In a more complex work, [17] introduce a Monte Carlo Simulation to identify the most vulnerable operation intervals of a power grid and use a DDPG agent to launch scaling attacks on the power flow sensor readings sent to the controlling automatic generation controller during those vulnerable time periods, thus destabilising the grid. To obstruct the attack’s detection they also launch coordinated GPS timestamp spoofing attacks on the phaser measurement unit data. As shown, the use of DT in environments can lead to severe cyberattacks, and (D)RL can be a potent tool for conducting attacks against industrial systems. However, to the best of our knowledge, DRL has not yet been applied to maliciously exploit the presence of DT in an infrastructure, and particularly to perform stealthy wearout attacks against the infrastructure’s physical systems.
3. Threat Modelling Adversary’s Objectives. The adversary aims to covertly manipulate the control signals transmitted from the VE to the PE in order to achieve long-term physical degradation of the targeted joint(s) of the robotic system. This can be defined by three main objectives: •
•
•
Accelerate Mechanical Wear-Out: Gradual physical degradation achieved by subtly manipulating the waypoint coordinates transmitted to the robotic arm, thereby inducing higher joint torque and mechanical stress. Sustained increases in load accelerate wear, reducing component lifespan and ultimately raising maintenance costs. Undetectability within the Network: The adversary aims to remain undetected by both the DT’s anomaly detection mechanisms (e.g., autoencoderbased detectors) and any built-in safety monitors. By applying small perturbations to waypoint commands, the attacker avoids triggering alarms while still imparting cumulative mechanical stress. Persistence in the System: Instead of inducing an immediate malfunction, the adversary employs repeated minute modifications that gradually erode system reliability. This “low and slow” approach requires continual adaptation to operational conditions, enabling maximised long-term wear-out while remaining inconspicuous.
Adversary’s Capabilities. Figure 1 depicts the attack surface targeted by our DRL-assisted adversary within a DT-enabled control system. The adversary resides on the communication path between the VE and the PE, enabling full interception, manipulation, and injection of control traffic consistent
with the Dolev–Yao threat model [22]. From this vantage point, the adversary can alter waypoint commands (control signals) as well as telemetry information exchanged between components, including joint positions, velocities, and accelerations. Key Assumption: We assume the adversary can access either L1 or L2 , but not both. Simultaneous access enables direct manipulation of the PE while evading the IDS through controller-level adjustments. Notably, if the adversary gains L1 access, modifications can be applied directly rather than through control signals sent from the VE. Our proposed attack aligns with the Exploitation and Actions on Objectives phases of the Cyber Kill Chain † , where the adversary gathers intelligence on system vulnerabilities and executes targeted disruptions to achieve its malicious objectives while maintaining operational stealth. We assume the adversary has compromised the VE–PE communication pipeline - for example, via malware leveraging publicly disclosed vulnerabilities such as CVE2024-2442 (exposing ICS protocol interfaces), CVE2024-2882 (enabling unauthorised code execution and data manipulation), or CVE-2023-5885 (facilitating persistent command-and-control). Such compromise may occur via human–machine interfaces at Level 1 or through a remote gateway associated with DT services. Additionally, we assume that the communication between the PE and VE occurs using widely used industrial systems protocols such as MODBUS-TCP, which have known and documented vulnerabilities (for example CVE-202562578- data is transferred in plaintext and CVE-202548466 - which enables attackers to send command packets remotely). Following infiltration, the adversary deploys a DRL-based policy that selectively manipulates waypoint commands transmitted from the VE to the PE. The adversary introduces small, per-axis modifications to reshape the trajectory without causing overtly invalid or unreachable poses. These modifications are formally defined through: [∆x, ∆y, ∆z] = [x, y, z] + [mx , my , mz ]
(1)
Where [x, y, z] represents the original waypoint, [mx , my , mz ] forms the modification made in each axis, and [∆x, ∆y, ∆z] is the manipulated waypoint received by the PE. The modification range was determined based on preliminary experiments, which showed that UR10e can reach waypoints with coordinates (x, y, z) in the range [-1, 1]. Since one of the main requirements of this study is that the agent’s actions be stealthy, the modification’s range was set to 10% of the arm’s reach. This way, it provides the agent with an action space that is at once flexible enough to explore damaging strategies and constrained enough to facilitate convergence to a stealthy policy. Our work assumes a grey-box threat model, in which the adversary has partial visibility into the DT system, † https://www.lockheedmartin.com/en-us/capabilities/cyber/ cyber-kill-chain.html
specifically, limited knowledge of the feature set and operational telemetry reflected in the observation space. Such information may stem from system logs, publicly available documentation (e.g., technical specifications), or prior reconnaissance. With this partial insight, the adversary can craft subtle, targeted perturbations that remain within expected operational bounds, improving their ability to evade anomaly detection while exerting strategic influence over the system’s control behaviour. Adversary’s Observation In the context of a Universal Robots UR10e ‡ system, the adversary has access to six joints’ runtime data, reflecting the arm’s standard degrees of freedom. Specifically: •
•
•
Original Waypoints: The baseline coordinates generated by the DT (or human operator) for each motion segment. Modified Waypoints: The altered waypoint coordinates as they were executed at the previous step (or “episode”) after the addition of the adversarial offsets. Joint States: Position, velocity, and acceleration for all six joints as measured at the previous step (or “episode”). These provide feedback on how past perturbations affected actual system dynamics and how close the manipulations were to triggering alarms (via the reward function).
Fig. 1 schematically illustrates how the adversary sits “in the loop”, monitoring and adjusting these data streams. In an out-of-context scenario (e.g., UR5 or a different brand of robot with a different number of joints), the attacker must adapt its policy to whichever number of degrees of freedom the new arm supports but can follow the same methodology: intercepting digital commands, applying minimal offsets, and observing sensor readouts to refine the DRL policy. Overall, through minimal real-time modifications of waypoint commands—powered by DRL training and guided by stealth feedback—the adversary accomplishes sustained, covert wear-out of a targeted joint in the UR10e system or any similarly configured DT-driven robotic platform.
•
•
• •
•
The adversary’s policy πϕ : X → A maps observations to actions and is optimized to maximize the expected cumulative reward: "∞ # X ∗ t π = arg max Eπ γ R(st , at , st+1 ) . (2) π
1)
•
S : The set of all possible system states, with st ∈ S representing the system’s underlying state at time t. A: The set of actions available to the adversary, where at ∈ A represents perturbations to waypoint coordinates. The per-axis modifying actions follow ∆x , ∆y , ∆z ∈ [−0.1, 0.1].
2)
https://www.universal-robots.com/products/ur10e/
ε ∼ N (0, σ)
(3)
Soft Target Network: To update Q-values, a soft Q-function is used to estimate expected return with entropy: Q̂(xt , at ) = r(xt , at ) + γExt+1 ∼p [Vψ̄ (xt+1 )]
3)
(4) Dual Critic: Two Q-functions are employed as the critic, such that the minimum value is used to update the value and policy: h Jπ (ϕ) = Ext ∼D DKL π ′ (·|ot )
exp(Qπold (xt ,·)) Z πold (xt )
i
(5) Soft Target Network Updates: The target network is updated using a smoothing coefficient (τ ) to improve stability: ψ̄ ← τ ψ + (1 − τ )ψ̄
(6)
3.1.1. Reward Function Formulation:. The adversarial goal is to maximise the mechanical stress (torque) on the target joint while remaining undetected by the anomaly detection system. The reward function is formulated as follows: R(xt , at , xt+1 ) =
where: • • •
‡
Observation-Based Policy: The adversary selects actions using a deterministic policy with inherent Gaussian noise for policy exploration: at = fϕ (εt ; xt ),
3.1. Formalisation of Attack
•
t=0
Algorithm: To address partial observability, we designed our adversary based on [30] to operate on observations xt rather than full states st . The algorithm incorporates the following modifications:
4)
The adversarial strategy operates in a Partially Observable Markov Decision Process (POMDP), under the assumption that the adversary may not have complete observability of the environment. For instance, certain joint readings (e.g., velocity and acceleration) may be inaccessible, reflecting real-world scenarios where the adversary has limited visibility into the system state. The POMDP is defined as the tuple (S, A, T, R, X, O, γ), where:
T: The transition probability function T (st+1 |st , at ), describing the likelihood of transitioning to state st+1 from st after taking action at . R: The reward function R(st , at , st+1 ), quantifying the adversary’s success in causing increased mechanical stress while avoiding detection. X : The observation space, where xt ∈ X represents the adversary’s partial view of the system. O: The observation probability function O(xt |st ), describing the likelihood of observing xt given the true state st . γ : The discount factor (0 ≤ γ < 1), prioritizing immediate rewards over future ones.
( τtarget · (1 − Panom ), if waypoints are reachable −1, otherwise (7)
τtarget represents the total torque applied to the targeted robotic joint. Panom is the probability of detection assigned by the anomaly detection system. A penalty of −1 is assigned if the manipulated waypoints are deemed unreachable.
3.1.2. Training Procedure. : The adversarial agent is trained using experience replay with the following steps:
used to execute the planned trajectory and move the arm’s end effector to the intended waypoint coordinates. Wear-Out Measurement To measure wear-out and mechanical stress on the targeted joint, we use torque 1) Sample mini-batches from the replay buffer. (Nm), as it is a good proxy for joint wear [31] and a 2) Compute new target Q-values using the soft target quantifiable measure we can evaluate in our simulation. network. MuJoCo’s main data structure (“mjData”), which holds 3) Update soft value network using the function: the simulation’s state, also keeps track of the torque h i 2 1 applied to each of the joints || , allowing us to easily JV (ψ) = Ext ∼ D 2 (V ψ(xt ) − Ea ∼πφ [Qθ (xt , at ) − log πφ (at |xt )]) access that information and use it for the calculation of (8) our reward (as explained in the DRL Environment). Other 4) Update the soft Q-function parameters using the metrics also exist that can be used to approximate the target value network: wear-out of robotic joints such as joint vibrations and 2 temperatures [31], however, this kind of information is JQ (θi ) = E(xt , at ) ∼ D 12 Qθi (xt , at ) − Q̂(xt , at ) not easily accessible in a MuJoCo simulation. (9) Target Joint In all of our experiments against UR10e, 5) Stochastic policy update using dual critic through the targeted joint we try to stealthily wear out is “Wrist the minimisation of KL divergence between the 3”, as this is one of the smallest joints of the arm and thus two critic networks. is designed to withstand the least torque** . As such, an 6) Update target network using smoothing. increase in the torque suffered by this joint is going to be By leveraging the POMDP framework, the adversarmore impactful than in other joints, resulting in quicker ial strategy effectively learns stealthy manipulation techmechanical wear-out. niques while maximizing wear-out impact in a highly dynamic industrial robotic environment. Additionally, during 4.2. DRL Agent: training the agent has access to manufacturer safety limits as the digital-twin model is known - at deployment the The DRL agent is responsible for modelling the agent no longer has access to safety limits. However, the robotic arm’s safe zones of an operation via the DT and agent does have access to anomaly scores due to adversary conducting intelligent wear-out attacks on the targeted access at Network L2 - exposing the Anomaly detection joint. The agent’s actions are modifications that get apsystems and data historian. plied to the original waypoint coordinates the robotic arm reaches, with the goal of introducing subtle changes to 4. Design & Implementation the arm’s trajectory and pose that induce additional strain on the targeted joint. This way, the targeted joint wears In this section, we describe the experimental setup out sooner, thus increasing the owner’s maintenance costs. considered to evaluate the proposed DRL-assisted wearHere, we evaluate four DRL algorithms, namely Soft out attack against a DT-enabled industrial robotic system. Actor-Critic (SAC) [32], Twin-Delayed Deep Deterministic Policy Gradient (TD3) [33], PPO [34] and Advantage Actor-Critic (A2C) [35], using their Stable-Baselines3†† 4.1. Simulation implementations. The said algorithms are selected due to their popularity, high performance and compatibility The experiments are conducted using a realistic with continuous action spaces such as ours. The main physics engine - MuJoCo † with an in-built simulation hyperparameters used for each algorithm can be found to the commercial specification of the industrial UR10e in Table 1. robot. MuJoCo is an advanced physics engine that enDRL Environment: A Gymnasium‡‡ DRL environables fast, accurate simulation of articulated structures ment was developed that uses the aforementioned MuJoCo interacting with their environment. It is used for research simulation to provide a space for the DRL agent to and development in areas like robotics and biomechanics. interact, learn and be evaluated in. For the experiments We use the high-quality UR10e MuJoCo model from conducted in this study, the robotic arm is configured MuJoCo Menagerie ‡ , as it represents a well-designed and to reach three different waypoints specified within the accurate model of the real robotic arm, providing us with simulation environment. All of the original waypoints are confidence that the simulated arm’s operation will be as reachable by the arm without additional strain on the realistic as possible and our results valid. Unfortunately, joints. Furthermore, the arm is configured to begin each while MuJoCo allows you to move each of the arm’s episode from a fully expanded horizontal home pose, joints under realistic physics rules, it does not provide move to each of the specified waypoints one by one, any out-of-the-box capability of commanding the arm to and finally return to its home pose. The observation and reach specific waypoints. As such, we used the Robotics action spaces were implemented to reflect the attacker’s Toolbox for Python library ¶ to plan the trajectory that observation and action capabilities as defined in Section each of the joints should follow in order to reach the designated waypoints. The model’s actuators were then t
i
|| †
https://mujoco.org/ ‡ See https://github.com/google-deepmind/mujoco menagerie for UR10e MuJoCo model used ¶ https://petercorke.github.io/robotics-toolbox-python/
https://github.com/google-deepmind/mujoco/issues/1095
** https://www.universal-robots.com/articles/ur/
robot-care-maintenance/max-joint-torques-cb3-and-e-series/ †† https://stable-baselines3.readthedocs.io/en/master/ ‡‡ https://gymnasium.farama.org/index.html
Learning rate Buffer size Batch size Tau Gamma Action noise Entropy Regularisation Coefficient Target policy noise Target noise clip Number of steps Number of epochs Generalised Advantage Estimator lambda Clip range
SAC 0.0003 1000000 256 0.005 0.99 N (0, 0.1) Auto
TD3 0.001 1000000 256 0.005 0.99 N (0, 0.1) -
PPO 0.0003 64 0.99 -
A2C 0.0007 0.99 -
-
0.2 0.5 -
128 10 0.95
5 1.0
-
-
0.2
-
TABLE 1: Algorithms’ Hyperparameters 3. The observations are standardised to have a mean of 0 and standard deviation of 1 using running statistics. Actions on are normalised in range [-1, 1] as far as the DRL algorithm is concerned, however it is important to note that the actual modifications applied to the original waypoint coordinates are still in the range [-0.1, 0.1] . The scaling of observations and actions is very important in DRL as it helps the Neural Networks used by the DRL algorithms to have more stable and efficient training. The agent’s reward function is designed to maximise torque at the targeted joint while simultaneously minimising the likelihood of detection. The reward function is mathematically defined in Equation 7. Initially, the environment checks whether the adversarially modified waypoints are reachable by the arm. If they are deemed unreachable, or if the action performed exceeds the safety limits specified by the manufacturer, the agent is penalised with a reward of −1 to discourage infeasible actions. If the waypoints are reachable, the MuJoCo simulation runs, calculating the total torque applied to the targeted joint and assessing the anomaly probability of the executed arm trajectory using an anomaly detection system that is detailed later in the paper. The integration of the anomaly detection system in the reward function guides the agent to execute stealthy attacks. Similar to observations and actions, rewards are scaled to maintain a standard deviation of 1 while preserving their original means. This ensures consistent reward magnitude, favouring effective training, without altering the reward sign, which could negatively impact the agent’s learning process.
4.3. Anomaly Detection System A comprehensive anomaly detection system capable of recognising abnormal robotic arm trajectories is used to assess the DRL agent’s stealthiness and support its training. The system consists of an ensemble of four Autoencoder-based anomaly detectors (Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), Residual Neural Network (ResNet), and Gated Recurrent Unit (GRU)), each using joint state data (positions, velocities, accelerations) at each time step to predict the probability of an anomaly. Each detector’s prediction is weighted by its test F1-score to produce an overall
anomaly probability, which is incorporated into the reward function to guide the agent toward taking undetectable malicious actions. The detectors are trained on sequences of joint states collected under normal operation (i.e., following the original waypoints) and, for time-enabled environments, the corresponding timestamps. For the training, validation, and testing of these anomaly detection systems a 76.4%/13.6%/10% split has been used over a dataset of 400,000 coordinate and joint position values collected over 16,000 noise inclusive trajectories. To emulate realistic behaviour, normal trajectories are captured with ±0.1 mm noise added to each waypoint, consistent with the UR10e specifications§§ . This allows them to learn a latent representation of normal arm behaviour. They operate by encoding normal inputs into a compressed representation and reconstructing them, similar to Kalman-filter residual analysis. Since they are trained solely on normal data with Gaussian noise, they reconstruct such data with minimal error. During inference, if a reconstruction exceeds a predefined threshold, it is flagged as anomalous. Thresholds are selected for each detector by training on normal and abnormal sequences and choosing the value that maximises F1-score. In known environments, the safety limit also serves as a rule-based detector: any action that breaches it is marked as anomalous. We provide a more in-depth discussion on IDS performance in Appendix C. Intrusion Detection Model LSTM AE LSTM CNN AE
ResNet AE
GRU AE
Architecture
F1-Score (%)
LSTM layers, Fully Connected Layers LSTM Layer, Fully Connected Layer Convolutional 1D layers, Rectified Linear Unit (ReLU) layers Convolutional 1D layers, ReLU layers, Batch Normalisation 1D layers, Max Pooling 1D layers, Residual Connections GRU Layers, Fully Connected Layers
96.2 100.0 97.4
99.8
98.3
TABLE 2: Anomaly Detectors’ Architectures and F1Scores
5. Experiments & Results In this section, we detail the experiments conducted to determine the most appropriate DRL algorithm for performing the desired attack. Additionally, we rigorously evaluate its performance across different scenarios and the effects of key hyperparameters. It is important to note that while the IDS mechanisms have been tested in a timeenabled environment (which provides resilience against replay attacks), the time difference between normal and manipulated behaviour is negligible. As a result, the experiments are conducted in the default environment, and temporal factors are not considered. §§ https://www.universal-robots.com/manuals/EN/HTML/SW5 19/ Content/prod-usr-man/complianceUR10e/H g5 sections/appendix g5/ tech spec sheet.htm
5.1. Algorithm Benchmarking In order to determine the most suitable DRL algorithm for conducting the intended attack, the algorithms mentioned in Section 4.2 are compared and contrasted based on their performance against the base case. The base case represents the scenario where the attacker has access to all relevant environmental observations (i.e., joint positions, velocities, and accelerations), targets the joint ”Wrist 3” of the robotic arm, and the agent trains for 10,000 episodes. Benchmarking the algorithms on the same scenario enables us to objectively and fairly determine which one is more suitable for the formulated problem. Each algorithm has been tested for a period of 100 episodes after training to further evaluate performance. The training and testing performance of each algorithm is plotted in Figure 2, and further benchmarking metrics can be found in Table 3. As shown in Figure 2, SAC performs the most consistently across training and testing. During training, all algorithms manage to obtain better performance than the baseline of “No Attack”, which represents the normal operation of the robotic arm. During testing, SAC, and PPO achieve significant improvements over the baseline, whereas TD3 and A2C demonstrate significantly weaker overall performance. Comparing the on-policy algorithms (PPO, and A2C): PPO demonstrates a consistent performance due to its clipping mechanism that prevents large changes to its policy; however, this causes the policy to be restrictive in its exploration. A2C presents a stronger, albeit much more unstable performance as it does not incorporate clipping, which allows it to better adapt to the gymnasium environment and the continuous action space. Out of the off-policy methods (SAC, TD3: SAC demonstrates consistent improvements over the training period while remaining stable, whereas, TD3 shows strong early improvements before converging at a suboptimal policy. Throughout training, SAC’s performance can largely be attributed to its inherent stochastic nature, as its entropy regularisation coefficient encourages exploration, introduces adaptability, and helps avoid local optima, thereby allowing convergence to more optimal policies. TD3, on the other hand, is a deterministic policy method, leading to poorer exploration capabilities - hence the instability during training. As shown in Table 3, SAC attains the largest AUC, with A2C following and TD3 and PPO trailing due to reduced learning. This reflects the clear sampleefficiency advantage of the off-policy methods, whose replay buffers enable faster and more stable cumulative learning than their on-policy counterparts† . Lastly, it should be noted that with a substantially larger number of training episodes, less sample-efficient on-policy algorithms could attain performance similar to SAC. Following the benchmarking experiments, additional environment configuration experiments have been conducted to determine the optimal values for both the agent’s action and observation spaces. The results of these experiments are presented in Section5.3 and Section 5.4. † For a detailed comparison between off-policy and on-policy DRL algorithms see [36].
Takeaways: SAC and A2C are very effective in conducting a stealthy wear-out attack in this environment. SAC’s superior sample efficiency and consistent performance during both training and testing make it the most appropriate algorithm overall for this study and is therefore used for all subsequent experiments. It should be noted, however, that PPO and A2C remain strong candidates and were only excluded due to instability observed during training.
(a)
(b)
Figure 2: Algorithm Benchmarking performance during (a) training and (b) testing
Mean Training Reward Mean Testing Reward Area Under Learning Curve
SAC 38.90
TD3 28.80
PPO 31.79
A2C 40.74
36.13
32.00
29.72
34.10
361281.49
319951.22
304346.14
340875.76
TABLE 3: Algorithms’ Benchmark Metrics
5.2. Exploration vs Exploitation The exploration–exploitation trade-off is a well-known challenge in DRL [37], where the agent must balance trying new actions (exploration) with selecting the bestknown action (exploitation). If the agent does not explore sufficiently, it may converge prematurely to a suboptimal policy, whereas excessive exploration can slow convergence and prevent the agent from stabilising an effective strategy.
SAC manages this trade-off through an entropy regularisation coefficient in its objective function [38]. This coefficient determines how strongly the agent is encouraged to maintain stochasticity in its policy. Higher values promote greater exploration, while lower values place more emphasis on exploiting learned behaviour. In this experiment, we evaluate a range of entropy coefficients: 0.1, 0.3, 0.5, 0.7, 0.9, and “auto”, to analyse their impact on performance. When “auto” is used, the entropy coefficient is learned dynamically during training. Figure 3a shows that the lowest coefficient (0.1) consistently yields the strongest training performance, followed by 0.3. These settings enable the agent to exploit effective behaviours early while maintaining enough stochasticity to avoid premature convergence. The remaining coefficients, particularly 0.7 and 0.9, perform substantially worse, demonstrating that excessive exploration prevents the policy from stabilising and leads to noticeably lower cumulative returns. The “auto” configuration performs comparably to the mid-range coefficients but does not surpass the best fixed setting. Testing results in Figure 3b reveal that, apart from the highest coefficient (0.9), all settings converge to similar evaluation performance, with 0.3 and 0.5 achieving the strongest results. This indicates that while low-entropy configurations accelerate learning, moderate exploration can yield slightly more stable behaviour at deployment. The poor performance of the 0.9 coefficient in both training and testing highlights the detrimental effect of persistent over-exploration. Overall, the results show that lower entropy coefficients provide the most favourable learning dynamics, while moderate coefficients offer competitive evaluation performance. Excessively large coefficients hinder both convergence and final reward, highlighting the importance of carefully balancing exploration and exploitation when training adversarial SAC agents. Key Takeaways: While exploration is necessary for diverse experience and avoiding local optima, excessive entropy reduces the ability to learn effectively. Therefore, it must be applied carefully to ensure effective convergence toward a nearoptimal solution. Lower fixed entropy coefficients provide the most reliable exploration–exploitation balance and are used for all subsequent experiments - in this case, 0.3 is used.
5.3. Action Space The action space determines the scale of perturbations the adversarial agent can inject into the control signal, thereby shaping both its learning dynamics and its eventual influence on the robotic policy. To assess how perturbation magnitude affects performance, we evaluate three ranges [−0.1, 0.1], [−0.01, 0.01], and [−0.001, 0.001], to span three orders of magnitude. As shown in Fig. 6a, the two larger ranges, [−0.1, 0.1] and [−0.01, 0.01], enable higher rewards during training, with the widest range yielding the fastest and strongest improvement. In contrast, the smallest range produces
(a)
(b)
Figure 3: Entropy Regularisation Coefficient comparison during (a) training and (b) testing.
almost no learning signal and remains close to the noattack baseline throughout, indicating that perturbations of this scale are too weak to meaningfully alter the victim’s trajectory during training. However, the evaluation results in Fig. 6b show that the intermediate range, [−0.01, 0.01], achieves the strongest performance at test time. While [−0.1, 0.1] still performs well, its impact is consistently weaker than that of the intermediate range, and the smallest range again fails to surpass the no-attack baseline. These findings indicate that although broad action ranges accelerate learning by enabling the agent to quickly discover impactful manipulations, moderately sized perturbations provide a more effective balance between influence and stability during deployment
Key Takeaways: Larger action spaces improve exploration during training, but excessively large adjustments reduce stealth and reduce the agent’s ability to generalise to evaluation conditions. Conversely, very restrictive ranges limit the agent’s ability to influence the system. The intermediate action range of [−0.01, 0.01] therefore offers the most effective trade-off, providing sufficient manipulative capability while remaining subtle enough to avoid detection, meaning this will be the action range for further experimentation.
Figure 4: Action Space Experimentation - Training Rewards
this conclusion: all configurations yield almost the same reward, substantially outperforming the no-attack baseline and demonstrating that the agent can execute a successful attack even with minimal information. Given this consistency, sub-experiments 3–5 restrict the observation space to a single joint. These results show only small differences between using position, velocity, or acceleration, with joint position performing marginally better. Overall, the agent performs well with very limited information, and the observation space has little impact on effectiveness. If anything, simpler observation spaces lead to slightly improved performance by reducing unnecessary complexity without diminishing the agent’s ability to disrupt the system. Key Takeaways: Although different observation spaces produce only minor variations in performance, the agent performs consistently well across all configurations, achieving rewards of approximately 39 in each case. Joint position emerges as the least influential feature, with simpler, alternating joint (3 joints of the 6) observation spaces yielding the strongest results. Therefore, to determine how well the agent performs with the least information, the remaining experiments will use single joint position values.
Figure 5: Action Space Experimentation - Testing Rewards
5.4. Observation Space & Rewards Relation The amount and type of data the adversarial agent has access to directly impact its ability to conduct the attack and its effectiveness. As such, in this set of subexperiments, we assess the adversarial agent’s effectiveness across varying knowledge spectra. For each subexperiment, the agent has access only to a subset of the original observation space (joint positions, joint velocities, and joint accelerations), following a grey-box threat model. This set of sub-experiments can also help determine which types of observations are most important for successfully conducting a stealthy wear-out attack. The observation space used for each sub-experiment is: •
•
• • •
Sub-experiment 1: Sub-experiment 4: Full joint information of half of the joints only (namely “Shoulder”, “Wrist 1” and “Wrist 3” joints) Sub-experiment 2: Sub-experiment 5: Full joint information of all joints (six joints - base joint is excluded for the Panda model) Sub-experiment 3: Joint position only Sub-experiment 4: Joint velocity only Sub-experiment 5: Joint acceleration only
Sub-experiments 1–2 were conducted first to determine how many known joints are necessary for the agent to perform an effective attack. Table 4 show that the training performance across all observation configurations is nearly identical, indicating that increasing the number of joints or adding additional features provides no clear learning advantage. The testing results in Table 4 reinforce
Full Information Partial Information Single Joint Information Single Joint Acceleration Single Joint Velocity Single Joint Position
Mean Training Reward 33.10 33.13 33.28
Mean Testing Reward 39.46 39.77 39.39
33.20
39.63
33.28 33.25
39.41 39.45
TABLE 4: Observation Space and Rewards relation comparison
5.5. Wear-Out Impact In order to facilitate the comprehension of a trained agent’s attack impact against the victim’s robotic arm, we compare it against the baseline “No Attack” case with respect to the total torque suffered by the target joint over 1000 episodes. For this and all subsequent experiments, the trained SAC agent with “auto” entropy regularisation coefficient and access to only the target joint’s position information (as far as joint states are concerned) is used, as this is the algorithm setup that achieves the best testing performance according to the previous experiments (See section 5.3 for the configuration experiment results). In the “No Attack” case, the mean torque at the target joint is 33.12 Nm, whereas with our trained agent, it increases to 51.15 Nm. This shows that the attack successfully increases the torque applied to the target joint by approximately 54.44% compared to its normal operation. Since the lifespan of a material decreases exponentially with the increase in stress [39], the reduction facattack k tor can be computed as ( σσnormal ) , with σ representing stress levels and k the material’s fatigue exponent (usually
ranges from 6 to 10 for aluminium materials [40] used in UR10e† ). The reduction factor can be calculated as (where 10 and 11 show the lower and upper bounds respectively): 6 51.15 ≈ 1.5446 ≈ 13.55 (10) Reduction Factor = 33.12 10 51.15 Reduction Factor = ≈ 1.5410 ≈ 77.00 (11) 33.12 Assuming that joint “Wrist 3” has a lifespan of 35000 hours ‡ , the lifespan of the joint if it is under constant attack by our agent is: 35000 Lnormal = ≈ 2583 (12) Reduction F actor 13.55 Lnormal 35000 Lattack = = ≈ 454.5 (13) Reduction F actor 77.0 Based on the above calculation, if the target joint operates continuously under the agent’s attack torque, its lifespan is drastically reduced to between approximately 454.4-2583 hours depending on the material fatigue. Lattack =
Takeaways: The trained agent increases the torque induced on the target joint by approximately 54.44% compared to the robotic arm’s normal operation. This significantly decreases the component’s lifespan from 35000 hours to between 454.5 and 2583 hours depending on the material fatigue limits, thus increasing the maintenance costs.
5.6. Undetectability One of the key desired characteristics of the adversarial agent is that it is able to conduct the attack in a stealthy manner, ie remain undetected by the security mechanisms in place within the network. In this experiment, we test the agent’s ability to perform manipulations to the control signal while remaining undetectable by the IDS specified in Section 4. For the purpose of this experiment, the agent is considered to be detectable if the anomaly score exceeds 50% - ie it is more likely that the value is anomalous than not. The attack experiment has been run for 1000 episodes in order to gather information from a sufficient number of interactions between the agent and the anomaly detection system. In this experiment, the adversarial agent achieved an average anomaly score of ≈ 20.63%, whereas, the baseline produces an average anomaly score of ≈ 17.91% due to standard process noise defined in the UR10e specification [41]. These results demonstrate that the perturbations introduced by the adversary are perceived as slightly more anomalous than the standard process noise. This means that the agent has learned to make minor adjustments that introduce higher mechanical stress to the targeted joint while effectively blending in with standard operating behaviour. In addition to showcasing the agent’s ability to perform the attack in a stealthy manner, the low anomaly scores produced by waypoints introduced by the agent’s † ‡
https://www.universal-robots.com/media/50880/ur10 bz.pdf https://www.universal-robots.com/media/8641/ur brochure gb.pdf
actions indicate that the agent can exploit deficiencies within the ensemble IDS deployed in the simulation environment. When this is combined with the results from the ”Wear-Out-Impact” experiments, the agent demonstrates that it can achieve two clear goals: cause stress-related wear-out while remaining undetectable within the network, making the attack particularly dangerous.
Key Takeaways: The agent is able to achieve an average anomaly score of ≈ 20.63%, producing a slightly higher anomaly probability compared to the standard operating behaviour of the robotic arm established in the baseline (≈ 17.91% due to process noise). This demonstrates that agent is able to maintain undetectability while exploiting deficiencies within the deployed IDS.
5.7. Naive Attacks Comparison
In order to evaluate the performance of the SAC adversarial agent against both the current, more naive methods, we have performed experiments against two naive methods: constant adjustment, and adding random values in the range [−0.01, 0.01]. This experiment was run for 1000 episodes in a similar manner to prior experiments, in order to collect enough interactions between the adversary and the anomaly detection systems. The first attack is a ”constant” attack, such that a constant value (in this case 0.07) is introduced to the original waypoint in order to adjust the standard behaviour, the second is a ”random” attack, where small random modifications in the range [−0.01, 0.01] are sampled from a uniform normal distribution to replicate the action space used by the agent while focusing on a stochastic approach. The performance of these attacks is compared to the SAC agent, and an established baseline in Table 5 The outcomes of this experiment demonstrate that both of the naive methods perform above the baseline by ≈ 8 points; however, both of these methods lead to a decrease in the mechanical stress applied to the targeted joint, which has the effect of increasing the lifespan of the joint as opposed to increasing stress to reduce the lifespan. When considering this result alongside the anomaly probability, it becomes evident that the naive attack methods perform very minor adjustments, so much so that the IDS considers the change in behaviour as less noisy than the standard process noise, leading to a minimal anomaly score which boosts the reward received despite failing to accomplish the attack goals. Comparing these results to the performance achieved by the SAC agent, we can see that the agent drastically outperforms both of the naive models, both with respect to the average reward received and the average torque applied to the targeted joint.
Key Takeaways: The adversarial agent demonstrates a considerable improvement over the implementation of the naive attack methods. Its strategic modification of the waypoints leads to a notable increase in stress applied at the target joint, while maintaining stealth within the network, and succeeds in both attack goals, whereas the naive methods only succeed at remaining undetectable over a 1000 episode period.
Normal Operation Mean Reward Mean Target Joint Torque (Nm) Mean Anomaly Probability (%)
Random Attack
SAC
27.20
Constant Attack (0.07) 32.28
33.13
40.60
33.12
32.28
33.13
51.15
17.91
1.31e-05
1.48e-05
20.63
TABLE 5: Naive Attacks vs SAC Performance Comparison
5.8. Out of Context One of the key advantages of DRL agents is their ability to operate in unforeseen environments and adapt to new situations. In this experiment, we evaluate the adversarial agent’s generalisability to unseen robotic arms on two new models: first, the UR5e (made by the same manufacturer as the UR10e) to assess performance in a semi-similar environment, and second, the Franka Panda, which introduces a substantially different configuration. The Franka Panda also differs in its joint structure, meaning that the target joint is “Joint 7” rather than “Wrist 3” in the UR models¶¶ . Additionally, the tolerance used in the inverse-kinematics solver inflates torque values for the Panda model. Although increasing this tolerance mitigates the effect, it also prevents the arm from moving, as the permissible error becomes larger than the adversarial perturbations. For consistency, we retain the original tolerance from the UR experiments and focus primarily on the agent’s behavioural trends in the Panda environment rather than the absolute torque values. Two sub-experiments are used to measure adaptability: the first deploys a “fresh” SAC agent in each new environment, and the second applies a pre-trained agent with a 2500-episode transfer-learning period (a quarter of the original training duration). For the Panda model, the agent is pre-trained on both UR10e and UR5e. Table 6 summarises these results. In the semi-known UR5e environment, both SAC agents increase the torque at the target joint relative to normal operation (8.07 Nm). The fresh agent produces a modest increase (8.92 Nm) with a small rise in anomaly probability, whereas the transfer-learning agent produces a substantially larger ¶¶ Franka Panda’s “Joint 7” is the arm’s final joint, analogous to “Wrist 3” in UR10e.
torque (17.66 Nm) but at a markedly higher anomaly probability (40.34%). This suggests that while the agent accounts for anomaly likelihood when optimising reward, it ultimately prioritises force application when it can leverage prior knowledge. Depending on the adversarial goal, this may be beneficial, though increasing the weight of the anomaly term could encourage stealthier behaviour. Overall, the UR5e results show that the agent can adapt effectively to a similar environment and still achieve its primary attack objectives. For the Panda model, the fresh SAC agent again increases the mean torque relative to normal operation, but the transfer-learning configuration does not lead to further improvement, instead failing to outperform the fresh baseline. This is likely due to the significant differences between the UR and Panda kinematic structures: while the UR10e and UR5e share characteristics that allow for beneficial transfer, the Panda model diverges substantially, limiting generalisation and in some cases hindering adaptation. Finally, because the pre-trained agent relies only on the target joint’s positional information, it avoids any incompatibility in observation-space dimensionality despite the change in degrees of freedom. This requirement for only the target joint’s state allows the agent to operate across multiple robotic platforms with differing joint counts. Key Takeaways: The pre-trained SAC agent adapts effectively to environments that share structural similarities, producing greater torque increases than a freshly trained agent under. Its use of only the target joint’s position enables straightforward transfer across robotic arms with different degrees of freedom. However, in dissimilar environments such prior knowledge becomes detrimental, as behaviours learned on UR models fail to translate effectively and ultimately limit the agent’s performance.
UR5e Results Normal Operation Mean Reward Mean Target Joint Torque (Nm) Mean Anomaly Probability (%)
Fresh SAC
6.68 8.07
9.18 8.92
Transfer Learning SAC 10.53 17.66
17.20
20.15
40.34
Franka Panda Results Normal Fresh Operation SAC Mean Reward Mean Target Joint Torque (Nm) Mean Anomaly Probability (%)
448.49 1118.96
451.66 1126.81
59.91
59.91
Transfer Learning SAC
59.91
TABLE 6: Normal Operation vs Fresh SAC vs Transfer Learning SAC Performance on UR5e and Franka Panda Environments
6. Discussion Our study shows that a DRL-assisted adversary can effectively execute a stealthy wear-out attack on a DTenabled industrial robotic system by subtly modifying the control signals transmitted from the VE to the PE. Unlike conventional attacks on DT infrastructures [5], [9], [11], [12], our approach leverages DRL to develop an intelligent adversary that adapts to its environment in real time. This enables the adversary to execute persistent, long-term attacks while remaining undetected by existing security mechanisms, eliminating the need for manual tuning to balance impact and stealth. Our adversarial agent follows a strategic “low & slow” attack methodology, carefully manipulating control signals to gradually increase the torque on the targeted joint by approximately 54% — while maintaining a low anomaly profile. This ensures the attack’s persistence by avoiding immediate detection and accelerating mechanical wear subtly over an extended period.
6.1. DRL Algorithm and Agent Performance A key factor in the success of this stealthy adversarial approach is the choice of an appropriate DRL algorithm. Our benchmarking demonstrated that SAC, overall, yields the most effective and sample-efficient adversarial policy for stealthy wear-out attacks in our experimental setup. It also showed that off-policy DRL algorithms (TD3 and SAC), generally demonstrate superior sample efficiency compared to on-policy DRL algorithms (PPO and A2C). This advantage of off-policy methods is attributed to their ability to store and reuse past experiences via replay buffers, which accelerates convergence and enhances learning efficiency. However, despite TD3’s high sample efficiency, its deterministic nature hindered its convergence to an optimal policy, whereas SAC’s stochastic approach encouraged the exploration of more beneficial policies. A useful insight gained from this research is the role of the entropy regularisation coefficient in balancing the exploration-exploitation trade-off for SAC algorithm. Constrained or automatically adjustable entropy regularisation coefficients help the agent gather sufficient diverse experiences to escape local optima without excessively slowing convergence. The automatically adjustable entropy coefficient appeared to be the most suitable for balancing exploration and exploitation in this problem, yielding the highest performance. Furthermore, it is noteworthy that the trained DRL agent can conduct a highly effective attack with minimal knowledge (e.g., only the position data of the targeted joint). This suggests that even the slightest amount of information can be sufficient for an intelligent agent to cause significant damage to the target system. Additionally, the DRL-assisted adversary’s strategic actions were shown to lead to a considerably more effective attack compared to naive constant or random actions. Another notable advantage of the adversarial agent trained is that it can discover and exploit weaknesses in the anomaly detection defence system used by the DT. Specifically, it exploits deficiencies in the anomaly detection system, allowing it to construct damaging waypoints
that are misclassified as less anomalous than the original waypoints. Moreover, the trained agent can adapt to new security mechanisms with which it has no prior experience. The pre-trained DRL agent was also shown to be generalisable and easily adaptable to new environments, regardless of the victim arm’s degrees of freedom, and to effectively damage a completely different robotic arm with only a limited number of fine-tuning episodes.
6.2. Task Complexity and Testbed The agent is trained and evaluated on navigation between three predefined waypoints, this is simpler than standard industrial operations. We argue, however, that the attack surface our adversary exploits is invariant to task complexity: ICS operation fundamentally consists of state-to-state transitions mediated by waypoint or set point commands, and our adversary manipulates commands within this framework. The control-loop primitives exercised in our setup—inverse kinematics resolution, trajectory planning, per-axis waypoint dispatch—are identical to those used in richer tasks such as multi-stage assembly or pick-and-place Extension to more complex tasks is reliant on agent training to handle the required operations. Our transfer-learning results in §5.8 support this: a pretrained SAC adversary adapts to a structurally similar arm (UR5e) within a quarter of the original training time, while structurally dissimilar platforms require larger retraining phases. We expect similar behaviour across more varied trajectories, which will require proportionally more episodes to converge on an optimal attack, but the underlying mechanics remain unchanged. Training is performed entirely offline against a replicated DT model, and the resulting policy adds no network traffic at deployment, suggesting the main barrier for mounting this attack is access to the VE-to-PE communication rather than computational power. We use MuJoCo as it supports extensibility across robotic platforms and reproducibility of the adversarial dynamics studied. Simulation allows other researchers to replicate our results, benchmark alternative defences against the same adversary, and adapt the methodology to new platforms. Several aspects of real-world deployment lie outside the scope of the present simulation. Extended operation under sustained adversarial torque would engage wear mechanisms beyond the stealthy wear-out employed, including more severe mechanical degradation, bearing fatigue, and specific failure modes, each of which would add to a more precise characterisation of component lifespan. Additionally a real-world deployment adds environmental considerations such as humidity, environmental and physical temperature, and power supply fluctuations - which may affect agent performance due to action trajectories, potentially leading to abnormal mechanical states.
6.3. Defence Mechanisms The main contribution of this paper is offensive: a novel DRL-driven wear-out attack, and the public release of the associated code, datasets, and trained policies. By disclosing this material, we aim to enable defenders to harden existing IDS mechanisms against the class of stealthy, adaptive adversaries we describe. Beyond direct
hardening, however, our framework is well-suited to a more defensive strategy: using the offensive policy as the training partner for a learning-based defender. An extension of our framework is to treat the defender as an agent embedded in the DT pipeline, observing joint telemetry and selecting defensive actions. The defender’s reward captures the objective of detecting manipulated trajectories while minimising false positives on legitimate operation. The defender can be trained against the offensive policies generated by our framework rather than pre-recorded attacks. This exposes the agent to precise low-and-slow manipulations that the reconstruction-based detectors fail to detect. Iterating adversary and defender training in a adversarial loop, would yield defenders more robust to unseen strategies than detectors trained on static datasets. Purely data-driven detectors are at a structural disadvantage against adversaries optimised to lie on the manifold of plausible trajectories. Physics-informed detectors, scoring deviations between observed joint dynamics and predictions, combined with cross-layer defences that jointly analyse telemetry and control signals, can additionally flag inconsistencies between the waypoints issued by the VE and the resulting joint behaviour, which is the discrepancy our attack introduces. A full evaluation of these directions is beyond the scope of this paper, whose contribution is the attack itself. We note that the most robust deployments are likely to combine several approaches and we view our publicly released attack as a method on which this defensive research can be trained to protect against.
7. Conclusions
6.4. Future Development
Despite the promising results achieved in this study, several areas warrant further exploration. Future work should investigate the use of a diverse anomaly detection system that combines rule-based and advanced machine learning techniques to enhance robustness against adversarial attacks. Examining the impact of the attack on non-targeted joints would provide deeper insights into potential collateral mechanical stress. Given the widespread adoption of DT technology in industrial domains, ensuring the security and resilience of these systems against sophisticated adversarial threats is of paramount importance. By demonstrating the capabilities of DRL-driven adversarial strategies, this research aims to raise awareness and support the development of robust defensive measures that can effectively counter such intelligent, stealthy attacks. Additionally, this work aims to illustrate the threat of DRL-assisted attacks against DT-enabled systems and emphasises the need to train existing defence mechanisms using real network traffic and attack-driven logs within the DT pipeline. This also lays the foundation for future work that leverages these insights to strengthen and adapt current defence strategies, helping reduce the threat posed by DRL-assisted adversaries. We include a discussion on ethics and dual-use in Appendix B.
This work presents a novel and covert wear-out attack using Deep Reinforcement Learning (DRL) to damage Digital Twin (DT)-enabled infrastructures. The adversarial agent strategically introduces perturbations to the control signals received by the physical robot, causing 54% more mechanical stress in the targeted joint while evading detection by an ensemble of Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN) and Residual Neural Network (ResNet) Autoencoder anomaly detectors. The agent’s ”low & slow” approach allows it to maintain a higher torque at the target joint for long periods, thereby stealthily accelerating its degradation and increasing maintenance costs. The adversarial agent shows great effectiveness in a grey-box setting with minimal information and rapid adaptability to unforeseen environments, indicating how dangerous it can be. We hope that the proposed adversarial agent can be used by researchers to facilitate the design and development of robust defence measures that can prevent such covert, intelligent attacks.
References [1]
M. Ghobakhloo, “Industry 4.0, digitization, and opportunities for sustainability,” Journal of Cleaner Production, vol. 252, p. 119869, 2020. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S0959652619347390
[2]
F. Tao, H. Zhang, A. Liu, and A. Y. C. Nee, “Digital twin in industry: State-of-the-art,” IEEE Transactions on Industrial Informatics, vol. 15, no. 4, pp. 2405–2415, 2019.
[3]
F. Tao, M. Zhang, Y. Liu, and A. Nee, “Digital twin driven prognostics and health management for complex equipment,” CIRP Annals, vol. 67, no. 1, pp. 169–172, 2018. [Online]. Available: https: //www.sciencedirect.com/science/article/pii/S0007850618300799
[4]
B. Sousa, M. Arieiro, V. Pereira, J. Correia, N. Lourenço, and T. Cruz, “Elegant: Security of critical infrastructures with digital twins,” IEEE Access, vol. 9, pp. 107 574–107 588, 2021.
[5]
W. Tärneberg, P. Skarin, C. Gehrmann, and M. Kihl, “Prototyping intrusion detection in an industrial cloud-native digital twin,” in 2021 22nd IEEE International Conference on Industrial Technology (ICIT), vol. 1, 2021, pp. 749–755.
[6]
G. Yadav and K. Paul, “Assessment of scada system vulnerabilities,” in 2019 24th IEEE International Conference on Emerging Technologies and Factory Automation (ETFA), 2019, pp. 1737– 1744.
[7]
I. A. Fernandez, S. Neupane, T. Chakraborty, S. Mitra, S. Mittal, N. Pillai, J. Chen, and S. Rahimi, “A survey on privacy attacks against digital twin systems in ai-robotics,” 2024. [Online]. Available: https://arxiv.org/abs/2406.18812
[8]
E. C. Balta, M. Pease, J. Moyne, K. Barton, and D. M. Tilbury, “Digital twin-based cyber-attack detection framework for cyberphysical manufacturing systems,” IEEE Transactions on Automation Science and Engineering, vol. 21, no. 2, pp. 1695–1712, 2024.
[9]
F. Akbarian, W. Tärneberg, E. Fitzgerald, and M. Kihl, “A security framework in digital twins for cloud-based industrial control systems: Intrusion detection and mitigation,” in 2021 26th IEEE International Conference on Emerging Technologies and Factory Automation (ETFA ), 2021, pp. 01–08.
[10] M. Ali, G. Kaddoum, W.-T. Li, C. Yuen, M. Tariq, and H. V. Poor, “A smart digital twin enabled security framework for vehicle-togrid cyber-physical systems,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 5258–5271, 2023. [11] F. Akbarian, E. Fitzgerald, and M. Kihl, “Intrusion detection in digital twins for industrial control systems,” in 2020 International Conference on Software, Telecommunications and Computer Networks (SoftCOM), 2020, pp. 1–6.
[12] E. C. Balta, M. Pease, J. Moyne, K. Barton, and D. M. Tilbury, “Digital twin-based cyber-attack detection framework for cyberphysical manufacturing systems,” IEEE Transactions on Automation Science and Engineering, vol. 21, no. 2, pp. 1695–1712, 2024.
[28] Y. Chen, S. Huang, F. Liu, Z. Wang, and X. Sun, “Evaluation of reinforcement learning-based false data injection attack to automatic voltage control,” IEEE Transactions on Smart Grid, vol. 10, no. 2, pp. 2158–2169, 2019.
[13] K. Chung, Z. T. Kalbarczyk, and R. K. Iyer, “Availability attacks on computing systems through alteration of environmental control: smart malware approach,” in Proceedings of the 10th ACM/IEEE International Conference on Cyber-Physical Systems, ser. ICCPS ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 1–12. [Online]. Available: https://doi.org/10.1145/3302509.3311041
[29] A. J. Abianeh, Y. Wan, F. Ferdowsi, N. Mijatovic, and T. Dragičević, “Vulnerability identification and remediation of fdi attacks in islanded dc microgrids using multiagent reinforcement learning,” IEEE Transactions on Power Electronics, vol. 37, no. 6, pp. 6359–6370, 2022.
[14] J. Chen, X. Gao, R. Deng, Y. He, C. Fang, and P. Cheng, “Generating adversarial examples against machine learning-based intrusion detector in industrial control systems,” IEEE Transactions on Dependable and Secure Computing, vol. 19, no. 3, pp. 1810– 1825, 2022. [15] A. S. Mohamed and D. Kundur, “On the use of reinforcement learning for attacking and defending load frequency control,” IEEE Transactions on Smart Grid, vol. 15, no. 3, p. 3262–3277, May 2024. [Online]. Available: http://dx.doi.org/10.1109/TSG. 2023.3343100 [16] E. Shereen, K. Kazari, and G. Dán, “A reinforcement learning approach to undetectable attacks against automatic generation control,” IEEE Transactions on Smart Grid, vol. 15, no. 1, pp. 959– 972, 2024. [17] S. Maiti, A. Balabhaskara, S. Adhikary, I. Koley, and S. Dey, “Targeted attack synthesis for smart grid vulnerability analysis,” in Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 2576–2590. [Online]. Available: https://doi.org/10.1145/3576915.3623155 [18] H. Li, X. Sun, and Z. Zheng, “Learning to attack federated learning: A model-based reinforcement learning attack framework,” Advances in Neural Information Processing Systems, vol. 35, pp. 35 007–35 020, 2022. [19] R. Langner, “Stuxnet: Dissecting a cyberwarfare weapon,” IEEE Security & Privacy, vol. 9, no. 3, pp. 49–51, 2011. [20] M. Eckhart and A. Ekelhart, “A specification-based state replication approach for digital twins,” in Proceedings of the 2018 Workshop on Cyber-Physical Systems Security and PrivaCy, ser. CPS-SPC ’18. New York, NY, USA: Association for Computing Machinery, 2018, p. 36–47. [Online]. Available: https://doi.org/10.1145/3264888.3264892 [21] M. Dietz, L. Hageman, C. von Hornung, and G. Pernul, “Employing digital twins for security-by-design system testing,” in Proceedings of the 2022 ACM Workshop on Secure and Trustworthy Cyber-Physical Systems, 2022, pp. 97–106. [22] D. Dolev and A. Yao, “On the security of public key protocols,” IEEE Transactions on Information Theory, vol. 29, no. 2, pp. 198– 208, 1983. [23] S. A. Varghese, A. D. Ghadim, A. Balador, Z. Alimadadi, and P. Papadimitratos, “Digital twin-based intrusion detection for industrial control systems,” in 2022 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops). IEEE, 2022, pp. 611–617.
[30] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning. Pmlr, 2018, pp. 1861–1870. [31] T. Kot, Z. Bobovský, A. Vysocký, V. Krys, J. Šafařı́k, and R. Ružarovský, “Method for robot manipulator joint wear reduction by finding the optimal robot placement in a robotic cell,” Applied Sciences, vol. 11, no. 12, 2021. [Online]. Available: https://www.mdpi.com/2076-3417/11/12/5398 [32] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” 2018. [Online]. Available: https://arxiv.org/abs/1801.01290 [33] S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in International conference on machine learning. PMLR, 2018, pp. 1587–1596. [34] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017. [Online]. Available: https://arxiv.org/abs/1707.06347 [35] V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” 2016. [Online]. Available: https://arxiv.org/abs/1602.01783 [36] R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction. MIT press, 2018. [37] H. Wang, T. Zariphopoulou, and X. Zhou, “Exploration versus exploitation in reinforcement learning: a stochastic control approach,” 2019. [Online]. Available: https://arxiv.org/abs/1812. 01552 [38] OpenAI Spinning Up, “Soft Actor Critic,” https://spinningup. openai.com/en/latest/algorithms/sac.html, 2018, accessed: 202503-21. [39] S. Japp, “Fatigue of structures and materials,” Netherlands: Kluwer academic publisher, 2014. [40] R. Stephens, A. Fatemi, R. Stephens, and H. Fuchs, Metal Fatigue in Engineering, ser. A Wiley-Interscience publication. Wiley, 2000. [Online]. Available: https://books.google.co.uk/books?id= B2aAPVa1TloC [41] [Online]. Available: https://www.universal-robots.com/blog/ determine-roi-for-your-palletizing-cobot-application/
Appendix A. Acknowledgements
[24] M. Ali, G. Kaddoum, W.-T. Li, C. Yuen, M. Tariq, and H. V. Poor, “A smart digital twin enabled security framework for vehicle-togrid cyber-physical systems,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 5258–5271, 2023.
This work was supported in part by the Engineering and Physical Sciences Research Council (EPSRC) under Award EP/V039156/1.
[25] J. Yan, H. He, X. Zhong, and Y. Tang, “Q-learning-based vulnerability analysis of smart grid against sequential topology attacks,” IEEE Transactions on Information Forensics and Security, vol. 12, no. 1, pp. 200–210, 2017.
Appendix B. Ethical Discussion
[26] Z. Ni, S. Paul, X. Zhong, and Q. Wei, “A reinforcement learning approach for sequential decision-making process of attacks in smart grid,” in 2017 IEEE Symposium Series on Computational Intelligence (SSCI), 2017, pp. 1–8. [27] Z. Wang, H. He, Z. Wan, and Y. Sun, “Coordinated topology attacks in smart grid using deep reinforcement learning,” IEEE Transactions on Industrial Informatics, vol. 17, no. 2, pp. 1407– 1415, 2021.
This work introduces an AI-assisted methodology for generating stealthy wear-out attacks against digital twin–enabled systems. While the primary objective is to advance defence mechanisms, we acknowledge the inherent dual-use risks associated with this framework and its potential for misuse.
The methods described in this paper could be repurposed by adversaries to conduct wear-out attacks against digital twin–enabled systems, increasing operational costs for targeted systems. Specifically, AI-driven optimisation of adjustments to operational commands, while maintaining stealth, amplifies these risks. In recognition of these risks, we disclose all code and the datasets used to train both the anomaly detection models and the reinforcement learning environment. This enables security teams to develop defensive mechanisms against the proposed attack, improving overall system security. Additionally, we release our models to support the development of more advanced agents, encouraging continued research into RL-assisted adversaries and their role in strengthening system security. In conclusion, we acknowledge the dual-use nature of the proposed methodology but contend that disclosing the model and data supports the security community in identifying emerging threats to digital twin systems and fosters deeper discussion on defending against AI-assisted attacks.
(a) Training Rewards
Appendix C. IDS Performance Results In this section, we present the full performance results of the IDS used within the MuJoCo simulation environment that the experiments were performed in. Table 7 reports the performance of each detector averaged across the three robotic arm models used in this study (UR5e, UR10e, and Franka Panda), and Table 8 reports performance in the time-enabled environment, in which timestamps are appended to joint-state sequences to provide resilience against replay-style attacks.
C.1. Basic Environment Table 7 presents the results of the ResNet autoencoder, CNN autoencoder, and standard LSTM models each achieve perfect scores across all five reported metrics (accuracy, F1, precision, recall, AUC), correctly classifying every sequence in the held-out evaluation set. Two interpretations are consistent with these results. The first is that the deviation between normal and adversarially modified trajectories is sufficiently large relative to the ±0.1 mm Gaussian process noise applied during normal operation that the differences between decision boundaries is negligible for the presented architectures. The second is that the detectors have over-fit to the specific structure of the training trajectories; this would be expected to manifest as degraded performance under distribution shifts, and we observe partial evidence of this in the time-enabled environment. The two explanations are not mutually exclusive: ResNet and CNN architectures tend to fit narrow distributions tightly, and the fact that all three top-performing models converge on identical perfect scores is consistent with learning closely related decision boundaries on the same problem. The GRU and LSTM autoencoders trail the top three. The GRU autoencoder achieves 96.97% accuracy, F1 of 0.9841, precision of 0.9697, recall of 1.0000, and AUC of 0.9267. The difference between precision and recall (a
(b) Testing Rewards
Figure 6: Action Space Experimentation.
recall of 1.0000 with precision below 0.97) indicates that the model captures malicious sequences in the evaluation set but produces a small number of false positives on normal trajectories. For an IDS deployed in an industrial setting where false alarms incur operational cost, this is the more tolerable failure mode of the two, missed detections would enable the attacks studied in this paper, whereas false alarms can be confirmed by operators. The corresponding AUC of 0.9267 indicates that the GRU autoencoder’s reconstruction-error distribution still separates the two classes well across thresholds, but with less margin than the top three models. The LSTM autoencoder is the weakest detector in the basic environment, reporting 93.33% accuracy, F1 of 0.9649, precision of 0.9394, recall of 0.9933, and AUC of 0.7813. The gap between recall (0.9933) and AUC (0.7813) is substantially larger than the other models which indicates that while the LSTM autoencoder catches most adversarial sequences the underlying score distribution is poorly calibrated: small threshold shifts would produce large changes in the precision/recall trade-off. Second, the precision of 0.9394 is the lowest of any model in either table, meaning the LSTM autoencoder produces the highest false-positive rate at its chosen operating point. Together, these observations suggest the LSTM autoencoder learns a less stable reconstruction of normal behaviour than other models.
TABLE 7: Performance metrics (average over robotic arm models). Model (Avg)
Metric
Accuracy (%) F1-Score Precision Recall AUC
GRU AE
LSTM AE
ResNet AE
CNN AE
LSTM
96.97 0.9841 0.9697 1.0000 0.9267
93.33 0.9649 0.9394 0.9933 0.7813
100.00 1.0000 1.0000 1.0000 1.0000
100.00 1.0000 1.0000 1.0000 1.0000
100.00 1.0000 1.0000 1.0000 1.0000
TABLE 8: Performance metrics (average over robotic arm models with time consideration). Model (Avg)
Metric
Accuracy (%) F1-Score Precision Recall AUC
GRU AE
LSTM AE
ResNet AE
CNN AE
LSTM
96.36 0.9808 0.9632 1.0000 0.9125
92.36 0.9594 0.9366 0.9853 0.8174
99.88 0.9993 0.9987 1.0000 0.9996
98.55 0.9921 0.9845 1.0000 0.9676
100.00 1.0000 1.0000 1.0000 1.0000
C.2. Time-Enabled Environment Adding timestamps to the inputs produces small but consistent degradations in four of the five models (seen in Table 8). The CNN autoencoder drops from 100% to 98.55% accuracy, the ResNet autoencoder drops from 100% to 99.88% the GRU autoencoder drops from 96.97% to 96.36% , and the LSTM autoencoder drops from 93.33% to 92.36%. The corresponding F1, precision, and AUC metrics shift in the same direction for each affected model, with AUC providing the most sensitive signal: the CNN autoencoder’s AUC falls from 1.0000 to 0.9676, and the ResNet autoencoder’s AUC falls from 1.0000 to 0.9996. The relative performance of the ResNet autoencoder’s AUC suggests the residual connections provide a stabilising effect on the additional input dimension, allowing the detector to incorporate timestamps without disrupting its learned representation of joint behaviours. The standard LSTM model is the only detector that maintains performance across both environments, reporting 100% accuracy, F1, precision, recall, and AUC in both Table 7 and Table 8. This is a notable result: while the autoencoders treat timestamps as an additional reconstruction the LSTM uses timestamps directly as a classification feature, gaining grounding without reducing reconstruction cost.