arXiv:2609.35644v1 [cs.NI] 28 Sep 2026
Energy-Driven Evaluation of Network Digital Twinning Applied to mmWave Beam Management João Borges 1 , Bruno Castro 1 , Kleber Cardoso Andrey Silva 3 , and Aldebaro Klautau 1 1
2
,
LASSE - 5G and IoT Research Group, Federal University of Pará (UFPA), Belém 66075-750, Brazil 2 Universidade Federal de Goiás, Goiânia 74001-970, Brazil 3 Ericsson Research, Rodovia Engenheiro Ermênio de Oliveira Penteado, Indaiatuba 13337-300, Brazil
Abstract Network Digital Twins (NDTs) are important enablers of 6G and future networks. However, there is a lack of studies regarding practical aspects, such as the impact of simultaneously changing twinning rate, fidelity, and other NDT operational parameters. For instance, works often consider the impact of operational parameters in isolation or with physical twin (PTwin) implementations relying on simulations. Therefore, the main contribution of this work is to provide an in-depth study on the performance impacts of both twinning rate and fidelity, using PTwins that rely on measurements obtained from hardware. We also investigate the optimization of these two operational parameters to minimize energy consumption, exploring a Bayesian Optimization (BO) method and what-if analysis. In this work, the NDT models an indoor propagation environment to optimize beam management. The virtual twin (VTwin) was implemented with the Sionna ray tracing (RT) simulator and the PTwin with our in-house setup composed of customized Wi-Fi radios with 32 antenna elements operating at 60 GHz. The study reveals important aspects of wireless channel NDTs, suggesting that fidelity levels vary throughout the experiment and with the what-if difficulty. Moreover, the twinning rate can be adjusted using sample-efficient methods, such as BO, even at lower fidelity levels.
Keywords: 6G, digital twin, millimeter wave, physical layer, ray tracing. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.
1
1
Introduction
The so-called Digital Twin (DT) systems are currently envisioned as key enablers of the sixth generation of cellular networks (6G) [1, 2]. Among its main features are the capacity of performing actions such as real-time system management, as well as a process known as what-if analysis, where a hypothetical scenario is first tested in a virtual replica before deployment in the real system [3, 4]. These features prospect the existence of a safe test ground for network rollouts, expansions, and automations, so to support them, the DT systems exploit heavy integration with sophisticated simulations and Artificial Intelligence (AI) [5]. However, as the reliance on these technologies grows, so do the concerns with resource consumption [5] [6]. For use cases such as Multiple Input Multiple Output (MIMO) beam management, this is further enforced due to a current lack of works focused on measuring and optimizing the resources employed to enable Network Digital Twin (NDT) operation. This work explores the performance impact of varying the twinning rate and fidelity, which are two of the most characteristic DT parameters and part of the DT definition according to the Digital Twin Consortium [7], with the goal of saving energy resources without compromising the system performance. Some related works focus on developing state-of-the-art methods to increase fidelity and maintain synchronization, e.g. [8], which is out of the scope of this work. The purpose of this paper is to evaluate the effects of simultaneously varying NDT parameters and checking their accuracy vs energy consumption performance. The problem is posed as a sequential discrete-continuous decision problem with episodic feedback, where we check the effects on the system performance when using different twinning rate policies present in the literature and their respective configurations, under two different fidelity levels, first using an ideal replica and then an imperfect one. For the twinning rate policies choice and configuration we use two approaches, a grid-search method, using a limited set of configurations, and a Bayesian Optimization (BO) method, for a more sample-efficient parameter space exploration. Besides these, we also consider a “max-rate” and a “noretwinning” policies for baseline purposes, the first which performs twinning on every step and the second that only performs twinning in the first step and do not execute it again. For the performance evaluation, we use a composite metric which considers both twinning error and energy consumption. Tests are carried out on an in-house NDT system, where the Physical Twin (PTwin) is composed of customized Wi-Fi radios with 32 antenna elements operating at 60 GHz and the Virtual Twin (VTwin) by the NVIDIA Sionna ray tracing (RT) simulator [9]. The main contributions of this paper are: • the investigation of the simultaneous effect of two main NDT operational parameters on the system performance; • the exploration of a Bayesian method for sample-efficient NDT twinning method and parameter selection;
2
• the introduction of a standards-based NDT architecture for the beam selection, described in Section 4; and • a hardware implemented PTwin, described in Section 5. The remaining content of this paper is structured as follows: Section 2 navigates the DT/NDT terminology and core concepts necessary for the paper, as well as the related work in DT operation, with emphasis on twinning. Then, Section 3 addresses the mathematical definition of the energy-driven evaluation problem for the NDT twinning and Multiple Input Single Output (MISO) beam selection. Next, on Section 4 we explain the NDT architecture as well as the what-if process pipeline we follow. In Section 5, the twinning rate and fidelity experiments, as well as the NDT setup considered in this work are described, and then discussed. Finally, Section 6 concludes the paper, summarizing the contributions and proposing future works.
2
Background and Related Work Table 1: Related work Twin rate
Twin fidelity
PTwin uses physical
evaluation
evaluation
measurements
Muñoz et al., 2022 [10]
✗
✓
✓
✗
DT fidelity evaluation in CPSs
Tan et al., 2022 [11]
✓
✗
✗
✗
Optimizing DT synchronization in manufacturing Network infrastructure for DT synchronization
Work
NDT
Use case
Cakir et al., 2023 [12]
✓
✗
✗
✓
Zheng et al., 2023 [13]
✓
✗
✗
✓
Data synchronization in vehicular networks
Tan et al., 2024 [14]
✓
✗
✗
✗
Optimizing DT synchronization in manufacturing
Muñoz et al., 2024 [15]
✗
✓
✓
✗
DT fidelity evaluation in CPSs
Salehi et al., 2024 [16]
✗
✓
✓
✓
MIMO beam selection in V2X scenarios
Huang et al., 2025 [17]
✗
✓
✗
✓
MIMO beam selection in V2X scenarios
Our work
✓
✓
✓
✓
MISO beam selection in indoor scenarios
In this section, we provide a brief overview of core concepts useful for the literature analysis, such as the difference between DTs and NDTs, as well as an explanation about the twinning rate and fidelity, which are the DT parameters explored in this investigation. Finally, at the end of the section, we provide insights into the current related works studying the individual effects of these concepts over DTs and NDTs systems and list this work main differences.
2.1
Digital Twins and Network Digital Twins
The term DT has a widely accepted origin as part of the work of Michael Grieves and John Vickers [18, 19]. According to their early description, the digital twin is a virtual representation of a physical target, containing real information and consisting of three elements: the physical target itself, its virtual representation, and the bi-directional data flow between them. DTs are able to perform tasks such as system state monitoring, logging and, using what-if scenarios it can
3
preemptively test changes on the system e.g. upgrades and downgrades, as well as run feature tests or check its behavior in alternative scenarios, all these in the virtual replica, without risking the real system [4]. Besides DT, the more specific term NDT is used when it comes to the communication networks. Here we adopt the Third Generation Partnership Project (3GPP) definition of NDTs [20], which is also shared by the European Telecommunications Standards Institute (ETSI) Zero-touch Network and Service Management (ZSM) [21], and the Internet Engineering Task Force (IETF) [22], where an NDT is a DT in which the replicated element is the network itself, or just part of it. From this definition, we can infer that DTs of end-to-end (E2E) networks, single domains, or even just the communications channel, or the user device, are all NDTs. There is also some disambiguation to be made regarding the naming convention. For instance, the Open RAN (O-RAN) [23] and also the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) [24] both refer to NDTs by the term Digital Twin Network (DTN). Alternatively, some works use DTN to define a network of many single DTs [25].
2.2
Types of Twinning Rates
The twinning rate is defined in [19] as the frequency that the virtual replica updates its states based on the twinned target. In other words, it is the rate in which the synchronization between VTwin and PTwin occurs. The authors in [26] provide 11 synchronization patterns for general DTs, which are named event-driven, time-driven, hybrid, adaptive, multi-level, event-chain, data-driven, context-aware, multi-modal, asynchronous, and manual. From these, we consider three patterns, regarded as within the scope of the setup and experiments used in this work, and provide practical implementations of them. They are: the event-driven, time-driven, and adaptive. The first pattern, event-driven is further divided into two variants, the threshold-based and the event-based ones, meant to cover metric monitoring events and also more intermittent triggers, respectively. For their practical implementation, the first variant, named threshold-based, defines a given numeric Key Performance Indicator (KPI) to monitor, which in this paper is the system Received Signal Strength (RSS), and it monitors whether it falls beyond a given dBm value. The second variant, identified as event-based, defines a list of N events, not linked to the numeric KPI monitoring, where the twinning will occur. From now on, we refer to these two variants (threshold-based and event-based ) as their own synchronization patterns for easier identification of the twinning method used in the experiments. Besides these, the next pattern, i.e., the time-driven, employs as trigger of the twinning process a fixed interval of time. Finally, the last one, i.e., the adaptive synchronization pattern, uses both an interval of time and a threshold, similar to the time-driven and threshold-based modes. The difference is that the time interval is variable, and also, twinnings triggered by trespassing the threshold take precedence over those triggered by time. More specifically, in this work, a given initial interval and threshold are defined and followed for 4
a given amount of time. Once this time is finished, the interval is incremented, and the process continues as long as the monitored metric stays within the threshold. However, if the threshold is crossed, the twinning is executed, and the interval is sent back to the initial value.
2.3
Twin Fidelity Performance Metric
The term fidelity appears throughout DT literature [27–31], with [19] defining the fidelity of a VTwin as the abstraction level, the set of parameters used in the replica, and its actual accuracy in representing the PTwin. In this work, we follow this definition, however, when considering different fidelities in our experiments, we only vary the actual accuracy of representing the target, and do not focus on the abstraction level and set of parameters used in the representation. Some works evaluate digital twin fidelities by comparing the sequence of system states [10, 15]. We perform a similar evaluation, but on a higher level, by comparing the results from the top-K chosen beams. This is due to the lack of lower-level information availability in the physical measurements of the dataset. Regarding the accuracy metric used for this evaluation, we adopt the “precision@k”, which is defined in [32]: precision@k =
|Iv ∩ Ip | , k
(1)
where Iv is the set of best indices suggested by the VTwin and Ip is the set of best indices from PTwin measurements. As mentioned in [32], the term precision references the common Machine Learning (ML) performance measure with the same name, which is the ratio of true positives to the total number of predicted positives, both true and false ones. For the precision@k, we consider the true positives as |Iv ∩ Ip |, which represents the VTwin correctly guessing the beam indices that compose the Top-K, and the total amount of predicted positives as k, given that the beam indices outside the Ip are predicted negatives.
2.4
Related Work
The investigation of DT systems is extensive due to its pervasiveness, being present on several industries, such as robotics, aerospace, manufacturing, transportation, agriculture, and healthcare, to name a few [33]. However, despite the large number of works, the majority of these are focused on investigating the benefits the DT brings to the system where it is installed, but do not consider the impacts of different deployment parameters on the use case performance, nor the DT’s cost of operation. Those that do so, deal with them separately, for instance, with a focus either on the twinning rate or the virtual replica fidelity, but not on two or more parameters simultaneously. The number of works further reduces when it comes to the specific case of NDTs and when the implementation of the PTwin in hardware is considered, as some works target only the VTwin, without any physical measurements. Also, in the current literature, 5
Network application
7 3 PTwin (physical network)
1 Data collection
What-if intelligence (functional model)
2
8 Unified Optimization 6 5 Emulation data repository VTwin (basic model) 4 Unified data models Network Digital Twin
Recommendation system (network optimization) Proposal system Recommended actions
9
Actuator agent Control
10
Figure 1: Overview of the system architecture, with elements and interactions inspired by the ones from the ITU-T. The elements and/or names in both italic and blue represent those that are not present and/or are named differently in the ITU-T architecture. there is a frequent assumption of the ideal case for the VTwin. But a virtual replica can rarely achieve 100% perfect mimicking, and this non-ideal condition affects the search for an optimal twinning rate, increasing the relevance of considering both the twinning rate and fidelity during the investigations. In Table 1, we list the works that most closely align with ours when it comes to the presence of the DT parameters studied, as well as the presence of an actual hardware implementation of the PTwin, and also with the indication of it being an NDT and their target use case. In [10] and [15], Muñoz et al. investigate methods of evaluating the fidelity level of DT of Cyber-Physical Systems (CPSs). These works focus on comparing traces of robotic systems’ operations to assess the real and virtual counterparts’ alignment quality. In [11] and [14], Tan et al. explore the search for an optimal twin synchronization for general DTs using mathematical formulations, but with a fixed DT fidelity and without a hardware implementation. In [12], Cakir et al. investigate how the communication system configuration, alongside the desired twinning rate, affect both, the synchronization performance between PTwin and VTwin, determining the maximum achievable twinning rate, and also the use case performance metric, but without addressing variations in twinning fidelity. In [13], Zheng et al. focus on the synchronization of vehicle users (VUEs) and the DT in an Internet of Vehicle (IoV) setting by using a deep learning (DL)-based network delay prediction approach, while also not considering variations on fidelity or the usage of measurements from real hardware as the PTwin. In [16], Salehi et al. leverage the FLASH dataset [34] as the PTwin and explore the usage of multiple VTwins with different fidelity levels to balance beam selection performance gains and the what-if analysis latencies associated with each of them, but do not explore the twinning rate. In [17], Huang et al. leverage knowledge about the scenario to optimize the tradeoff between system complexity and its performance. This is done by varying the number of antennas and, therefore,
6
the VTwin fidelity, and performing a scenario cognition-empowered beam selection that balances performance and system complexity, without relying solely on maximum fidelity. They also, however, do not explore the variation of a second parameter, namely, the twinning rate. Finally, to the best of the authors’ knowledge, there is currently no prior work investigating the energy-versus-performance impact of varying multiple DT parameters (i.e., twinning rate and fidelity) while considering a DT for communications network, in other words, an NDT, involving both, a hardware implemented PTwin and a VTwin, applied for the millimeter wave (mmWave) beam selection problem.
3
Problem formulation
In this section, we provide the mathematical definition for both the energydriven twinning evaluation and the MISO beam selection problem. We start with basic terminology definitions and proceed to explain the system evolution with the formulas of the monitored performance and energy metrics, then we explain the optimization function and why we can use a BO approach as a solution.
3.1
Basic definitions
The NDT system management considered in this work is described as a dynamic decision making problem where the NDT system must choose between a twinning method m ∈ M from a set of M possible twinning methods, and then its respective parameters θm ∈ Θm , from a set Θm of parameters, to reduce both, the twinning error and the system energy consumption. This choice occurs at the start of an episode e from a total of E episodes, where e = 1, ..., E, each occurring over a series of steps s, with s = 1, ..., S, considering S describing the total amount of steps needed for the completion of an episode, so that for all episodes we have a total number of steps T , where T = E × S. This series of choices is summarized in the tuple a, which is an action that must be taken at the start of each episode e, so that we have ae = (me , θme ).
3.2
System evolution
The system evolves in episodes, each ending with a given cost C, that considers both, the cumulative twinning performance P and the cumulative energy consumption J. The performance P is the cumulative absolute error between d while J is the sum of the energy the measured RSS and its estimation (RSS), consumption W in Watt-hour of performing the twinning process followed by what-if analysis, both over the entire episode e, and subject to the chosen action ae . Finally, the performance P and the cumulative energy consumption J of
7
the episode e using the action ae are respectively defined as Pe (ae ) =
S X
d s| ; |RSSs − RSS
∀ e ∈ E,
(2)
s=1
and Je (ae ) =
S X
Ws ;
∀ e ∈ E.
(3)
s=1
Due to their different value ranges, these two metrics have their samples normalized using min-max scaling, bringing their values to the [0, 1] interval, ˜ Finally, we and with their normalized versions receiving the notation P̃ and J. can say that the cost C at the end of the episode e, using the method me and parameters θme is Ce (me , θme ) = Ce (ae ) = wp × P̃e (ae ) + wj × J˜e (ae ),
(4)
where wp and wj are complementary scalar weights within the range [0, 1]. These were added to adjust the focus on one of the two metrics, depending on the system optimization priority, namely the twinning error and the energy expenditure, respectively. For this work we use the values wp = wj = 0.5.
3.3
Optimization function and solution approach
Considering the described dynamic decision-making problem, we have the following equation as the target â(m̂, θ̂m ) = argmax E[Ce (me , θme )] ∀ e ∈ E,
(5)
m,θm
which represents the search for the optimal action â, described in terms of the optimal twinning method m̂ and its respective optimal parameters θ̂m . These values enable the system to achieve the optimal cost by minimizing the expected twinning error and the energy expenditure of each episode. Due to the nature of this problem, the method chosen for this search leverages a BO method, more specifically, the Tree-structured Parzen Estimator (TPE) [35], because of its capabilities of accommodating categorical variables such as the ones from the discrete twinning method choice, and also dealing with hierarchical sequential decision problem, with discrete-continuous decision and episodic feedback, such as the choice of twinning rate method, and then the search for its parameters.
3.4
The MISO beam selection process
In this work, we consider the MISO beam selection problem to be described by a scenario with a transmitter (Tx) and receiver (Rx) Uniform Planar Arrays (UPAs), with precoder represented by the vector f, which is used alongside the
8
PTwin
Unified data repository
Recommendation system
What-if intelligence
Actuator agent
VTwin
Loop [twinning rate] 1: Data collection 2: What-if analysis starts 3: Requests updated Rx position 4: Sends updated Rx position Update states 5: Requests what-if analysis (emulation) 6: Sends the simulation results (optimization)
Performs ray-tracing simulation
Calculates the best beam index 7: Stores the ray-tracing simulation and what-if analysis results 8: Sends what-if results Generates recommendation 9: Sends recommended actions to actuator agent 10: Implements the recommended action on the PTwin (control)
Figure 2: Sequence diagram of a what-if scenario generation using the proposed ITU-T inspired NDT system architecture. channel matrix H to obtain the receiving signal y for a given beam pair index i. In this scenario, the receiving antenna uses only one element, therefore, the receiving signal y, given the index i, is calculated by yi = Hfi ,
(6)
and the optimal beam choice î is given by î = arg max |yi |.
(7)
i∈1,...,I
4
NDT System
In this section, we present the NDT architecture considered for the experiments and explore the what-if scenario generation. The architecture is inspired by the one from Y.3090 [36], defined by the ITU-T and shown in Fig. 1, where the names in black correspond to terminology and elements directly from the ITU-T architecture, while the ones in blue represent alternative names or added elements. Also, we explain how they interact with each other during a what-if scenario implementation via a sequence diagram presented in Fig. 2. This what-if scenario considers the beam selection problem described in Subsection 3.4. As illustrated in Fig. 1 and Fig. 2, we have the PTwin, also named physical network in the system architecture, which is composed by a scenario with a fixed and a mobile UPA, that can be either the Tx or the Rx, depending on the simulation settings. The position of this mobile UPA is captured at a
9
Rx
Tx (a) Top view of the physical area covered in this work for the in-house NDT setup, obtained by a camera mounted on the roof.
(b) Top view of the room virtual model for the in-house NDT setup, generated using the open-source software Blender.
Figure 3: Real and virtual versions of the room used in the in-house NDT setup experiments. (a) Physical area: a laboratory room with the main elements being a whiteboard, side and center tables, and a robot carrying the MikroTik setup, which obtained the RSSIs; (b) virtual model: generated using the open-source software Blender. The figure also highlights the placement of the Tx (red), the Rx (green), and the path followed by the robot (black line). The red arrow indicates the orientation of the Tx, while the Rx always points to the direction of the robot movement around the center table. given frequency, defined by the twinning rate parameter, and logged at an element named unified data repository in the NDT system, which in this case is a database. This is the step 1, denoted as data collection. Once reaching the unified data repository, the Rx position data is considered to be at disposal of the NDT. This data collection step is a loop and keeps repeating until a what-if analysis is initiated by the recommendation system in step 2. Once started, the what-if intelligence requests updated Rx position information from the unified data repository during step 3, that are then transmitted to the VTwin at step 4, which updates its internal states with the newest Rx position. The NDT is composed by a unified data repository and a unified data model, which in turn is composed of VTwin and what-if intelligence, also known as basic model and functional model, respectively. The VTwin is a virtual representation of the PTwin, described with a given fidelity, while the what-if intelligence is the element capable of extracting useful insights from this representation by simulating/emulating proposed actions, e.g., to optimize the system. This interplay between the VTwin and the what-if intelligence is described in both, the Y.3090 and our architecture, in terms of emulation and optimization, identified by the step 5 and step 6. In step 5, the what-if intelligence requests an analysis to check the best beam index (emulation), for which the VTwin performs a RT procedure and in step 6, it sends the RT data back to the what-if intelligence, that calculates the best beam index according to the received data (optimization). Then, the RT data, alongside the what-if results, are propagated to the
10
unified data repository for logging purposes, and only the what-if results to the recommendation system, in steps 7 and 8, respectively. In step 9, the recommendation system, via the proposal system, generates the recommendation, consisting of the best index for the Tx, and sends to the actuator agent. Finally, in step 10, the actuator agent implements it in the PTwin (control).
5
Evaluation
In this section we explain the twin fidelity and twinning rate analysis experiments as well as the experimental setup used for them. At the end, we provide the results and discussion regarding the interplay between twinning rate and fidelity in regards to the adopted energy-aware performance metric.
5.1
Analysis on the Twin Fidelity and the Twinning Rate
For the experiments of the twin fidelity impact on the system performance, we execute the beam selection procedure on the VTwin and compare with the ground-truth results. The experiment consists of evaluating the precision@k throughout the path covered by the mobile UPA, as shown in Fig. 3(b). The experiment is composed of 47 episodes, each with 58 steps, where for every step we have a top-K for both, PTwin and VTwin, so that we can calculate the precision@k of all the steps for every episode. Concerning the twinning rate, we are evaluating the variation on the cumulative cost, integrating twinning error and energy expenditure as defined in Eq. 4. In the experiments, we consider a receiver moving in the indoor scenario shown in Fig. 3(a). The NDT system has at its disposal a model-driven VTwin based on ray tracing simulation using NVIDIA Sionna [37] for channel modeling. The tests are carried out on a single test episode, while the other episodes are used for parameter search using the TPE. We follow the process described in Fig. 2, where the system initially stays in a loop, feeding the unified data repository with up-to-date receiver positions, as described in step 1. However, differently from the default behavior shown there, once the experiment starts, each twinning is immediately followed by the trigger of a what-if analysis. This what-if consists of obtaining the beam index with higher RSS from a codebook of 64 possible Tx beams. The Rx adopts a single receiving beam, with a directional pattern informed in Table 2. The Rx always points to the direction of the robot movement around the center table. The process then follows through with steps 2 to 10. For this experiment, we first perform an initial twinning at the first step and then, for the retwinning policy, we contrast the use of the twinning rates mentioned in Section 2.2, i.e., 1) time-driven, 2) threshold-based, 3) adaptive, and 4) event-based. More specifically, the time-driven method consists of defining a fixed interval for the twinning to occur, which for this experiment must assume the value of an integer within the range [2, 58]. The threshold-based
11
Table 2: Simulation Parameters Adopted for the VTwins Simulation parameters Sionna version
1.1.0
Material database
ITU
Transmitter
UPA 6x6
Receiver
Single element
Frequency/bandwidth
60 GHz
Antenna pattern (Tx/Rx)
TR.38901
Antenna polarization
Vertical
Max. num. of ray interactions
6 6
Max. num. of paths per Tx
10 (to allow virtually all)
Num. of rays launched per Tx
106
Synthetic array
True
Allow LOS paths
Enabled
Specular reflection
Enabled
Diffuse reflection
Enabled
Refraction
Enabled
monitors if the RSS achieved at the latest step falls below a given float number, and if it does, a twinning is triggered. The adaptive uses both an interval and a threshold, with the difference that the interval is variable instead of fixed. It works by following the initial interval two times, and if the RSS stays within the threshold during this period, then the interval is incremented by one. Otherwise, if the RSS falls below the threshold, then a twinning is immediately triggered and the interval goes back to the initial value. This behavior reduces the energy consumption of executing a twinning followed by a what-if analysis if the performance is not degraded enough to call for a new twinning. Finally, the event-based twinning rate works by defining a set composed of N events, which are integer numbers within the range [1, 58], representing the steps for twinning to occur. The first four methods are initially implemented with parameters chosen using a rudimentary grid-based hyperparameter search. After that, for a sampleefficient search within the constrained dataset, a TPE approach is used to find both, a performative method and its respective hyperparameter(s). The grid search is an exhaustive parameter sweep over a finite set of parameters that compose the grid [38]. The parameters for it are predefined and chosen empirically
12
K=16
Precision@K (%)
Precision@K (%)
100 95 90 85 80 75 70 65 60 55 50 45 40 35 30 25 20 15 10 5 0
K=32
58 55 52 49 46 43 40 37 34 31 28 25 22 19 16 13 10 7 4 1
58 55 52 49 46 43 40 37 34 31 28 25 22 19 16 13 10 7 4 1 Step
(a) Fidelity boxplots assuming K = 16.
100 95 90 85 80 75 70 65 60 55 50 45 40 35 30 25 20 15 10 5 0 Step
(b) Fidelity boxplots assuming K = 32.
Figure 4: Boxplots of the fidelity at each step considering the entire dataset. The dots represent outliers, while the lower/upper whiskers are, respectively, the smallest/largest values besides the outliers. Using the “precision@k” metric, the fidelity varies throughout the experiment evolution and with the chosen K. Lower K values suggest a more stringent what-if analysis, trying to pinpoint the best index, while higher K values widen the search to a bigger subset. as described in the following: • time-driven – the intervals can be a value from the set {2, 4, 8, 10}, representing the steps. For instance, if 2 is chosen, the twinning is executed every 2 steps; • threshold-based – the threshold can assume one of the following values in dBm {−65, −60, −55, −50}; • adaptive – needs to define two parameters: – a threshold (in dBm) from the set {−65, −60, −55, −50} and – the initial interval (in seconds) from the set {1, 2, 4, 8}; For the event-driven, aiming for better coverage of the search space we adopt varying set sizes (N = 1, 2, 3), with 5 candidates for each. Also, the choice of values to be swept by the grid-search is based on a Stratified Mixed-Strength Design, a methodology inspired by Mixed-Strength Covering Arrays (MCAs) [39–41] from the area of combinatorial design and experimental optimization, used to uniformly partition the events within the [1, 58] value range. Therefore, the sets are the following: • Single-Event Sets (N = 1): {6}, {18}, {30}, {42}, {54}. • Two-Event Sets (N = 2): {3, 35}, {15, 47}, {22, 58}, {10, 28}, {40, 51}. 13
• Three-Event Sets (N = 3): {2, 29, 57}, {12, 33, 49}, {7, 20, 45}, {19, 38, 53}, {5, 25, 41}. Now, the TPE is used to search for both, a twinning rate type and its respective parameters, however, differently from the grid-search, the TPE can handle a much larger search space. It begins with a categorical search space composed of time, threshold, adaptive or event. After that, the search follows as described below: • time-driven – the intervals can be an integer value from the range [1 : 58] (steps); • threshold-based – the threshold can assume any float value within the range [−80 : 0] (dBm); • adaptive – needs to define two parameters: – a float threshold from the range [−80 : 0] (in dBm) and – an integer initial interval from the range [1 : 58] (in steps); • event-based – parameters choice is again, a set of size N = 1, 2, 3, where each element can assume integer values from the interval [1 : 58], which represents the experiment “episode” step. In summary, for this experiment, the goal is to maximize the use case performance metric while at the same time reducing the number of times we trigger a twinning process request to the NDT. In our case, this is immediately followed by a what-if scenario calculation, therefore also saving energy resources as reflected in the performance metric defined in Eq. 4.
5.2
In-house experimental setup
This work considers the indoor environment illustrated in Fig. 3(a). The data acquisition relies on a low-cost experimental platform designed for AI-based beam management research. The setup consists of two MikroTik wAP 60G radios serving as the Tx and Rx. They operate in the 60 GHz mmWave band with 6 × 6 UPA antennas, having the four corners disabled and so, considering only 32 active elements. The Tx is positioned next to a wall, by the side of the room, while the Rx is mounted on the top of a mobile robot, that follows a circular path around the center table, with their placement shown in Fig. 3(b). The MikroTik radios are commercial off-the-shelf devices, customized with an OpenWRT operating system and a modified firmware for the IEEE 802.11ad baseband controller, enabling the implementation of custom beamforming codebooks and the real-time extraction of the Received Signal Strength Indicator (RSSI) [42]. In previous works, we developed a synchronized client-server system that manages the entire data collection process, capturing time-aligned multimodal data that combines wireless metrics (e.g., the RSSI) with images from RGB 14
and depth cameras [42, 43]. To enable mobility, the Rx radio is mounted on a robotic platform programmed to autonomously follow a predefined path in the environment. This integrated system allowed for the creation of robust datasets that correlate wireless channel quality with visual information in dynamic scenarios. Further details on the platform architecture and its validation can be found in [42, 43]. Also, in this work, we leveraged the digital model of the room developed in [32], where the experimental setup is located, shown in Fig. 3(b). But, while in [32] we used the proprietary software Remcom Wireless InSite [44], here we use the open-source NVIDIA Sionna as the VTwin. Next, we focus on the description of the experiments for monitoring the impact of varying these two parameters on the system performance.
5.3
Results
Time-driven (interval=10) Threshold-based (th=-60) Adaptive (th=-65,int=4) Event-driven (positions=[14,46]) TPE (Event-driven: positions=[14, 38]) Max-rate No-retwinning
Sample Index
(a) Using the ideal VTwin.
17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 0
Time-driven (interval=10) Threshold-based (th=-60) Adaptive (th=-65,int=4) Event-driven (positions=[14,46]) TPE (Event-driven: positions=[14, 38]) Max-rate No-retwinning
58 55 52 49 46 43 40 37 34 31 28 25 22 19 16 13 10 7 4 1
58 55 52 49 46 43 40 37 34 31 28 25 22 19 16 13 10 7 4 1
17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 0
Cumulative Weighted Cost
Cumulative Weighted Cost
Results obtained from the in-house experimental setup for the VTwin fidelity assessment are illustrated in Fig. 4, where in Fig 4(a) and Fig 4(b) we show the fidelity levels achieved for precision@k = 16 and precision@k = 32, respectively. One can notice that the fidelity levels vary throughout the experiment evolution and is also heavily influenced by the chosen k. This result shows that a high fidelity can be achieved partially, throughout the experiments and also when the task at hand involves less accuracy. For instance, by not using K = 1 and instead increasing it, we are changing the what-if task from suggesting the single best beam index to just narrowing down the suggested optimal beam index to a K-sized subset of the original codebook.
Sample Index
(b) Using the Sionna VTwin.
Figure 5: Cumulative weighted cost of both versions of the VTwin, the “ideal” one, that yields perfect fidelity, and the one using Sionna RT simulator, with varying fidelity. In Fig. 5(a) and Fig. 5(b), we present the variation of the cumulative cost
15
defined in Section 3.2 when using different types of twinning rates and fidelities. In Fig. 5(a) we explore the case where the fidelity is ideal, i.e. the VTwin is always at the highest fidelity possible and we focus only on variations of the twinning rate, while in Fig. 5(b) we explore the case where the fidelity varies throughout the experiment by using the Sionna-based VTwin. In both figures we compare the twinning methods described in 2.2, as well as the max-rate, no-retwinning and the TPE approaches. For the twinning methods, we first execute a preliminary step, which is a grid search for time-driven, thresholdbased, adaptive, and event-driven. From this search, we choose only the most performative parameters from each twinning method to be compared with the remaining approaches. Next, the max-rate and the no-retwinning are brought into comparison to provide the baseline for both, a system that is constantly twinned, meaning that for every step it executes a new twinning, and for a system that only performs twinning on the first step and then do not perform it again. Lastly, the TPE approach comes with a more sophisticated search through the parameter space. Notice that, the less costly and therefore best method in Fig. 5(a), TPE, displays robustness against performance degradations in Fig. 5(b), when the VTwin presents the non-ideal conditions shown in Fig. 4.
5.4
Discussion
From the results we found that, even when calibrated, VTwin fidelity levels are usually dependent on the experiment time evolution and the stringency of whatif conclusions. We also checked that a correctly chosen twinning rate method and respective parameters yield relevant gains in performance. To provide insights into solving the latter, we proposed the usage of the TPE bayesian optimization method for a sample-efficient search within the available candidates, which gave positive results. We also observed that, the correct choice of twinning rate method and parameters can greatly influence energy savings and high-performance maintenance. The TPE method, investigated as candidate solution for the twinning rate method and respective parameters selection, provided the best results for the integrated energy vs performance metric introduced in Section 5 when compared with methods and parameters chosen using the grid-search method with empirical value ranges, when used in both ideal and non-ideal fidelity conditions, which was the case when using Sionna as the VTwin. Regarding the non-ideal fidelity observed on the Sionna-based VTwin, two possible causes are: simplified assumptions regarding the radiation patterns adopted in the simulated environment antennas, and also, simplifications of the room geometry itself, which for instance, do not capture the entire complexity of the objects present in the room.
16
6
Conclusion
In this work, we investigated the tradeoff between VTwin state update rate and virtual to real accuracy, namely the twinning rate and fidelity, respectively. By establishing an integrated comparison metric that allowed for energy-aware performance comparison, we evaluated different twinning rates present in the literature at different fidelity stages. We also investigated the usage of a BO approach, more specifically TPE, to provide a sample-efficient NDT parameter optimization alternative. To achieve our results, we used data involving both real measurements acting as the PTwin and its respective model-based VTwin, leveraging a simulated environment and RT channel modelling. The results showed that the fidelity levels are dependent on both the experiment time evolution and the objective of the what-if. In other words, the fidelity will vary throughout the experiment, and it will be more or less prominent if the goal of the what-if is to narrow down the set of options for the suggested action, or to provide a single optimal suggestion, respectively. The findings represent a contribution to the investigations regarding the NDT deployment payoff and resource-aware operation, which tends to be of critical importance as the technology adoption grows, given the projected massive pervasiveness of 6G networks. Future work is to be focused on expanding this investigation to more datasets and more sophisticated setups, also exploring the viability of performing the same type of experiments for other network domains or even E2E networks.
Acknowledgements This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nı́vel Superior - Brasil (CAPES) – Finance Code 001; in part by the Conselho Nacional de Desenvolvimento Cientı́fico e Tecnológico (CNPq); in part by the Brasil 6G project (01245.020548/2021-07), supported by Rede Nacional de Ensino e Pesquisa (RNP) and Ministério da Ciência, Tecnologia e Inovacão (MCTI); in part by the Innovation Center, Ericsson Telecomunicações Ltda., Brazil; in part by the OpenRAN Brazil - Phase 2 project (MCTI grant Nº A01245.014203/2021-14); and in part by the Project Smart 5G Core And MUltiRAn Integration (SAMURAI) [MCTI/Comitê Gestor da Internet no Brasil (CGI.br)/Fundacão de Amparo à Pesquisa do Estado de São Paulo (FAPESP)] under Grant 2020/05127-2
References [1] D.-H. Tran et al., “Network Digital Twin for 6G and Beyond: An End-toEnd View Across Multi-Domain Network Ecosystems,” IEEE Open Journal of the Communications Society, vol. 6, pp. 6866–6911, 2025.
17
[2] L. U. Khan, W. Saad, D. Niyato, Z. Han, and C. S. Hong, “Digital-twinenabled 6G: Vision, architectural trends, and future directions,” IEEE Communications Magazine, vol. 60, no. 1, pp. 74–80, 2022. [3] 3rd Generation Partnership Project, “Study on management and orchestration,” 3GPP, Technical Report TR 32.801-01, April 2026. [Online]. Available: https://www.3gpp.org/DynaReport/32801-01.htm [4] IBM Topics, “What Is a Digital Twin?” https://www.ibm.com/think/ topics/digital-twin, 2024, accessed: 2026-05-17. [5] M. Sheraz, T. C. Chuah, Y. L. Lee, M. M. Alam, A. Al-Habashna, and Z. Han, “A Comprehensive Survey on Revolutionizing Connectivity Through Artificial Intelligence-Enabled Digital Twin Network in 6G,” IEEE Access, vol. 12, pp. 49 184–49 215, 2024. [6] W. Liu, Y. Fu, Z. Shi, and H. Wang, “When Digital Twin Meets 6G: Concepts, Obstacles, and Research Prospects,” IEEE Communications Magazine, vol. 63, no. 3, pp. 16–22, 2025. [7] Digital Twin Consortium, “The Definition of a Digital Twin,” https://www.digitaltwinconsortium.org/initiatives/ the-definition-of-a-digital-twin/, 2026, accessed: May 12, 2026. [8] G. Khalaf, M. Itani, and S. Sharafeddine, “A UAV-Aided Digital Twin Framework for IoT Networks With High Accuracy and Synchronization,” IEEE Transactions on Network and Service Management, vol. 23, pp. 3013– 3025, 2026. [9] J. Hoydis et al., “Sionna RT: Differentiable Ray Tracing for Radio Propagation Modeling,” in 2023 IEEE Globecom Workshops (GC Wkshps), 2023, pp. 317–321. [10] P. Muñoz, M. Wimmer, J. Troya, and A. Vallecillo, “Using trace alignments for measuring the similarity between a physical and its digital twin,” in Proceedings of the 25th International Conference on Model Driven Engineering Languages and Systems: Companion Proceedings, ser. MODELS ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 503–510. [Online]. Available: https://doi.org/10.1145/3550356.3563135 [11] B. Tan and A. Matta, “Optimizing Digital Twin Synchronization in a Finite Horizon,” in 2022 Winter Simulation Conference (WSC), 2022, pp. 2924– 2935. [12] L. V. Cakir, S. Al-Shareeda, S. F. Oktug, M. Özdem, M. Broadbent, and B. Canberk, “How to synchronize Digital Twins? A Communication Performance Analysis,” in 2023 IEEE 28th International Workshop on Computer Aided Modeling and Design of Communication Links and Networks (CAMAD), 2023, pp. 123–127. 18
[13] J. Zheng et al., “Data Synchronization in Vehicular Digital Twin Network: A Game Theoretic Approach,” IEEE Transactions on Wireless Communications, vol. 22, no. 11, pp. 7635–7647, 2023. [14] B. Tan and A. Matta, “The digital twin synchronization problem: Framework, formulations, and analysis,” IISE Transactions, vol. 56, no. 6, pp. 652–665, 2024. [Online]. Available: https://doi.org/10.1080/24725854. 2023.2253869 [15] P. Muñoz, M. Wimmer, J. Troya, and A. Vallecillo, “Measuring the Fidelity of a Physical and a Digital Twin Using Trace Alignments,” IEEE Transactions on Software Engineering, vol. 50, no. 12, pp. 3122–3145, 2024. [16] B. Salehi et al., “Multiverse at the Edge: Interacting Real World and Digital Twins for Wireless Beamforming,” IEEE/ACM Transactions on Networking, vol. 32, no. 4, pp. 3092–3110, 2024. [17] Y. Huang and Y. Zhao, “Beam management for millimeter-wave mobile communications based on digital twin-enabled scenario cognition,” Scientific Reports, vol. 15, no. 1, p. 13802, 2025. [18] M. Grieves, “Digital twin: manufacturing excellence through virtual factory replication,” White paper, vol. 1, no. 2014, pp. 1–7, 2014. [19] D. Jones, C. Snider, A. Nassehi, J. Yon, and B. Hicks, “Characterising the Digital Twin: A systematic literature review,” CIRP Journal of Manufacturing Science and Technology, vol. 29, pp. 36–52, 2020. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S1755581720300110 [20] 3rd Generation Partnership Project (3GPP), “Study on management aspect of Network Digital Twin,” 3GPP, Tech. Rep. TR 28.915, 2024. [Online]. Available: https://www.3gpp.org/dynareport/28915.htm [21] ETSI, “ETSI GR ZSM 015 V1.1.1: Zero-touch network and Service Management (ZSM); Network Digital Twin,” 2024. [22] IETF, “Network Digital Twin: Concepts and Reference Architecture,” 2025. [23] ORAN, “Research Report on Digital Twin RAN Use Cases,” 2024. [24] ITU, “Y.3090 : Digital twin network - Requirements and architecture,” 2022. [25] Y. Wu, K. Zhang, and Y. Zhang, “Digital Twin Networks: A Survey,” IEEE Internet of Things Journal, vol. 8, no. 18, pp. 13 789–13 804, 2021. [26] W. Alghamdi and E. Albassam, “Synchronization Patterns for Digital Twin Systems,” Journal of Applied Data Sciences, vol. 5, no. 3, pp. 1026–1037, 2024. 19
[27] M. Zhang, Y. Zuo, and F. Tao, “Equipment energy consumption management in digital twin shop-floor: A framework and potential applications,” in 2018 IEEE 15th International Conference on Networking, Sensing and Control (ICNSC), 2018, pp. 1–5. [28] W. Luo, T. Hu, C. Zhang, and Y. Wei, “Digital twin for CNC machine tool: modeling and using strategy,” Journal of Ambient Intelligence and Humanized Computing, vol. 10, no. 3, pp. 1129–1140, 2019. [29] C. Zhuang, J. Liu, and H. Xiong, “Digital twin-based smart production management and control framework for the complex product assembly shop-floor,” The international journal of advanced manufacturing technology, vol. 96, no. 1, pp. 1149–1163, 2018. [30] Y. Zheng, S. Yang, and H. Cheng, “An application framework of digital twin and its case study,” Journal of ambient intelligence and humanized computing, vol. 10, no. 3, pp. 1141–1153, 2019. [31] J. Guo, N. Zhao, L. Sun, and S. Zhang, “Modular based flexible digital twin for factory design,” Journal of Ambient Intelligence and Humanized Computing, vol. 10, no. 3, pp. 1189–1200, 2019. [32] G. Silva et al., “Digital Twin of a Beam Selection Procedure in an Indoor Scenario,” in XLIII Brazilian Symposium on Telecommunications and Signal Processing - SBrT 2025, 2025. [Online]. Available: https: //doi.org/10.14209/sbrt.2025.1571157127 [33] M. Bokhtiar Al Zami, S. Shaon, V. Khanh Quy, and D. C. Nguyen, “Digital Twin in Industries: A Comprehensive Survey,” IEEE Access, vol. 13, pp. 47 291–47 336, 2025. [34] B. Salehi, J. Gu, D. Roy, and K. Chowdhury, “FLASH: Federated Learning for Automated Selection of High-band mmWave Sectors,” in IEEE INFOCOM 2022 - IEEE Conference on Computer Communications, 2022, pp. 1719–1728. [35] J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl, “Algorithms for Hyper-Parameter Optimization,” in Advances in Neural Information Processing Systems, J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Weinberger, Eds., vol. 24. Curran Associates, Inc., 2011. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/ 2011/file/86e8f7ab32cfd12577bc2619bc635690-Paper.pdf [36] International Telecommunication Union, “Y.3090: Digital twin network Requirements and architecture,” International Telecommunication Union (ITU), Geneva, Switzerland, Recommendation Y.3090, 2022, approved in 2022-02-13. [Online]. Available: https://www.itu.int/rec/T-REC-Y. 3090-202202-I/en
20
[37] J. Hoydis et al., “Sionna: An Open-Source Library for Next-Generation Physical Layer Research,” arXiv preprint, Mar. 2022. [38] A. Géron, Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow. ” O’Reilly Media, Inc.”, 2022. [39] C. J. Colbourn, “Combinatorial aspects of covering arrays,” Le Matematiche, vol. 59, no. 1–2, pp. 125–167, 2004. [40] H. Avila-George, J. Torres-Jimenez, V. Hernández, and L. GonzalezHernandez, “Simulated annealing for constructing mixed covering arrays,” Journal of Applied Research and Technology, vol. 10, no. 4, pp. 658–671, 2012. [41] L. M. Moore, M. D. McKay, and K. S. Campbell, “Combined array experiment design,” Reliability Engineering & System Safety, vol. 91, no. 10, pp. 1281–1289, 2006. [42] V. Conceição et al., “Low-Cost Experimental Platform for AI-Based Beam Management Using WiFi Radio,” in XLIII Brazilian Symposium on Telecommunications and Signal Processing - SBrT 2025, 2025. [Online]. Available: https://doi.org/10.14209/sbrt.2025.1571157315 [43] V. Conceição, “Samurai mmWaves testbed,” https://github.com/Manky0/ samurai-servers, 2025, accessed: 20 August 2025. [44] “Wireless EM Propagation Software - Wireless InSite - Remcom,” https://www.remcom.com/wireless-insite-em-propagation-software, accessed: 2023-11-02.
21