ConceptioArchivearXiv CS
arXiv CSopen access

DPDS: A DPDK-Based Packet Delayer and Spacer

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributedsystemsprotocols
networking, internet, protocols, distributed systems

This work has been submitted to the IEEE Open Journal of the Communications Society for possible publication. Received XX Month, XXXX; revised XX Month, XXXX; accepted XX Month, XXXX; Date of publication XX Month, XXXX; date of current version 15 June, 2026. Digital Object Identifier 10.1109/TBA

DPDS: A DPDK-Based Packet Delayer and Spacer ETIENNE ZINK , FABIAN IHLE , AND MICHAEL MENTH

(Senior Member, IEEE)

Chair of Communication Networks, University of Tübingen, 72076 Tübingen, Germany

arXiv:2606.17716v1 [cs.NI] 16 Jun 2026

CORRESPONDING AUTHOR: M. Menth (e-mail: [email protected]). The authors acknowledge funding in part by the Deutsche Forschungsgemeinschaft (DFG) under grant 503231190, and in part by the Open Access Publishing Fund of the University of Tübingen. Furthermore, the authors acknowledge the use of Claude Opus 4.8 to assist in developing scripts for experiment automation and data visualization presented in this paper.

ABSTRACT In this paper we tackle the problem of adding varying delay to packets for link emulation.

Naive approaches either add more delay than desired or cause packet reordering, both of which are undesirable. We develop adaptive delay correlation, which adds positively correlated delays to packets efficiently. It takes a mean delay and standard deviation (jitter) as input, as well as a half-life period to control the delay dynamics. We investigate the accuracy and dynamics of the resulting packet delays with and without bandwidth limitation. As a result we give a recommendation for the configuration of the halflife period. We implement adaptive delay correlation in a DPDK-based packet delayer and spacer (DPDS), investigate its performance on hardware, and compare it with the widely used link emulator NetEm and the recently developed DPDK-based emulator MoonEm. DPDS outperforms both of them with a zero-loss throughput of 95 Gbit/s for constant delay and, with spacing enabled, 85 Gbit/s for varying delay with 3 ms jitter. Further, DPDS supports packet reordering with zero-loss throughputs of 73 Gbit/s and 58 Gbit/s for constant and varying delay, respectively, as well as policing and two packet loss models. INDEX TERMS Software-Defined Networking, Network Link Emulation, Varying Packet Delays, Band-

width Limitation, Adaptive Delay Correlation, DPDK.

I. Introduction

Network link emulators (NLEs) are often needed for performance studies of network protocols. They typically delay packets according to a specified mean delay and standard deviation (jitter), limit bandwidth using a spacer, police packets exceeding a traffic contract, and drop packets according to some loss model. As network speeds keep increasing, NLEs that support high traffic rates are also needed. While fast NLEs add only constant packet delays, some studies point out that packet delay over networks typically varies [1], [2]. However, there is currently no high-performance NLE that supports varying delay appropriately. Adding varying delay is harder than it appears. If the added delay for a packet is substantially longer than the delay for its successor, the successor packet will face additional delay, leading to achieved mean delays that are longer than desired. Packet reordering according to transmission times after adding delay solves the problem, but packet reordering

is detrimental to many transport or application protocols, such as TCP [3]–[5], QUIC [6], and RDMA [7], [8]. An alternative to reordering is smoothing consecutive delays with an exponential moving average (EMA) so that the resulting packet delays are correlated. However, the configuration of the EMA’s internal weight is left to the user and the weight does not adapt to the current traffic rate. In this paper we develop an efficient algorithm that adds varying delay to packets using the EMA approach. The algorithm is configured with a desired mean delay and jitter, and a half-life period to control the delay dynamics. Moreover, the algorithm continuously measures the current packet rate and adapts the EMA’s internal weight according to the configuration parameters and the measured packet rate. We further investigate the performance of uncorrelated delays, adaptive delay correlation, and packet reordering using simulation with and without bandwidth limitations. The results give insights into the accuracy of the different

This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ VOLUME ,

1

ZINK et al.: DPDS: A DPDK-Based Packet Delayer and Spacer

Table 1. Delay capabilities of existing emulators and DPDS. Configurable, Fixed, and Adaptive indicate whether the correlation weight is userselectable, fixed by the implementation, or recomputed from the current rate.

Constant delay

Varying delay

Reordering

Correlation

In-kernel Dummynet [14] NetEm [15]

✓ ✓

× ✓

× Configurable

eBPF + QDisc [18] 6GDetCom Emulator [19] Rattan [20] TheaterQ [21]

✓ ✓ ✓ ✓

× ✓ × ✓

× × × ×

DEMU [16] MoonGen LTE [22]

✓ ✓

× ×

× ×

MoonEm [17] SmartNet [23] DPDS

✓ ✓ ✓

× × ✓

× Fixed Adaptive

Kernel bypass

approaches and lead to a recommendation for the configuration of the half-life period. Building on these insights, we implement adaptive delay correlation in a DPDK-based NLE, called DPDK-based packet delayer and spacer (DPDS). We evaluate its zero-loss throughput (ZLT) and maximum supported rate and compare them with those of the widely used NLE NetEm and the recently developed DPDK-based emulator MoonEm. We show that the desired delay is well achieved under ZLT, and that DPDS also works as intended under varying traffic rates, which is challenging as its internal operation depends on measured traffic rates. The remainder of this paper is structured as follows: Section II reviews related work. Section III introduces adaptive delay correlation. Section IV evaluates adaptive delay correlation in simulation. Section V describes the emulator DPDS. Section VI evaluates the DPDS prototype by experimentation. Section VII concludes the paper. II. Related Work

Nussbaum and Richard [9] distinguish two concepts of network emulators: virtual network emulators (VNEs) and network link emulators (NLEs). VNEs emulate a whole network. Examples of VNEs are Mininet [10], its extension Containernet 2.0 [11], and Kollaps [12]. In contrast, NLEs apply link characteristics to packets as they are transmitted or received on an interface. The emulated conditions include packet delay, limited bandwidth, and packet loss [13]. Examples of NLEs are Dummynet [14], NetEm [15], DEMU [16], and MoonEm [17]. In this paper, we focus on NLEs. We therefore use the terms NLE and emulator interchangeably. 2

We classify NLEs into two categories: emulators implemented in the Linux kernel and emulators implemented with kernel bypass frameworks. Table 1 presents an overview of NLEs of both classes. It further distinguishes whether the emulators support constant delays, i.e., a jitter/standard deviation of zero, or varying delays, i.e., a jitter greater than zero. To accurately emulate varying delays, emulators implement either packet reordering (sorting packets according to their transmission times) or delay correlation. Although reordering can accurately emulate varying delays, it impairs in-order delivery. A. In-Kernel Emulators

Dummynet and NetEm are the most commonly used emulators in academic research [22] and have been thoroughly evaluated in several studies [9], [13], [22], [24]. Rizzo [14] introduced Dummynet in 1997, and Carbone and Rizzo [25] enhanced it further in 2010. Dummynet emulates bandwidth limitation and constant delays on a single workstation. NetEm was introduced by Hemminger [15] in 2005 to be more extensible than Dummynet. It is implemented as a queueing discipline (QDisc) in Linux’s traffic control, and can be combined with other QDiscs, like the token bucket filter (TBF) [26] to implement bandwidth limitation. NetEm supports delay emulation with reordering and correlated delays. It implements delay correlation through an EMA with a configurable correlation weight. The correlation weight determines how similar consecutive delays are. However, according to Hemminger [15], NetEm’s correlated delays are inaccurate. Later, additional in-kernel emulators were developed to overcome the limitations of Dummynet and NetEm. Becker et al. [18] implemented a constant delay emulator based on eBPF and a Linux QDisc to increase scalability in largescale virtual edge testbeds. Haug et al. [19] introduced the 6GDetCom Emulator, a custom Linux QDisc, which emulates varying delays of 5G TSN bridges with packet reordering at rates of up to 1 Gbit/s [27]. Wang et al. [20] developed Rattan, a Rust-based constant delay emulator. However, Rattan has not yet been thoroughly evaluated. Recently, Ottens et al. [21] introduced TheaterQ, a custom Linux QDisc, to emulate varying delay with reordering [28]. Of these four emulators, only two, i.e., 6GDetCom Emulator and TheaterQ, support varying delays with reordering. Neither supports correlated delays. B. Kernel Bypass-Based Emulators

Emulators based on kernel bypass frameworks were developed to surpass the accuracy and throughput of existing inkernel emulators [16], [17]. Kernel bypass frameworks, like DPDK [29] and Snabb [30], leverage direct memory access to completely bypass the Linux kernel networking stack, and pass packets from the network interface card (NIC) directly to the user space application [31]. This reduces the overhead of the generalized Linux networking stack. VOLUME ,

Three DPDK-based NLEs have been proposed: DEMU [16], MoonGen LTE [22], and MoonEm [17]. DEMU, developed by Aketa et al. [16], was the first to accurately emulate short (sub-millisecond) constant delays. It was later enhanced by Sasaki et al. [32] to emulate packet loss and by Puakalong et al. [33] to emulate bandwidth limitations. Stratmann et al. [22] and Lachnit et al. [17] developed two emulators based on the MoonGen [34] traffic generator. Stratmann et al. [22] developed an emulation of various constant delays and bandwidths for the uplink and downlink of an LTE testbed. Lachnit et al. [17] introduced MoonEm, an extension to MoonGen which uses NIC hardware timestamping to accurately emulate constant delays. None of these three emulators support varying delays. In contrast, Vogt et al. [23] demonstrated SmartNet, a DPDK-based VNE for SmartNIC-aware network emulation. SmartNet primarily focuses on emulating a virtual network, but can also emulate various link characteristics, such as varying delays, inside this virtual network. It supports correlated delays through an EMA with a fixed correlation weight [35]. III. Adaptive Correlated Packet Delays

In the following we develop a method to add correlated delay values to consecutive packets. It greatly diminishes the drawbacks of independent delays, which are a larger mean delay and lower jitter than desired. The method is based on the EMA that generates delay values Ci and averages them to delay values Di , which are used for delaying the packets. We first derive appropriate parameters for C to meet a desired mean and jitter for D. This approach depends on a correlation weight α, which controls its dynamics. Then, we make the technique adaptive to different packet rates R by setting α such that a desired half-life period th is met for generated delay values Ci . Finally, we make it adaptive to changing packet rates R by measuring the rate on the fly. A. Exponentially Weighted Moving Average for Correlated Packet Delays

We successively generate delays Ci and average them with a correlation weight parameter α by Di = α · Di−1 + (1 − α) · Ci

for i ≥ 1 and D0 = C0 . This is an EMA, background information is available in [36]. The generating delays Ci are independent and identically distributed random variables, C for short. Therefore, the averaged delays Di have the same mean as the generated delays Ci . The variance of Di can be computed as X V AR[Di ] = V AR[αi · C0 + (1 − α) · αj · Ci−j ] 0≤j<i

= α2·i · V AR[C0 ] X + ((1 − α) · αj )2 · V AR[Ci−j ]. 0≤j<i

VOLUME ,

In the limit for i → ∞ and for 0 < α < 1, we obtain σ 2 = V AR[D] = lim V AR[Di ] = i→∞

1−α · V AR[C]. 1+α

Thus, to obtain a desired mean µ and jitter σ for correlated delays Di , the generating delay’s mean should be set to µ and its variance to V AR[C] =

1+α 2 ·σ . 1−α

While some NLEs offer the use of the EMA for correlated delays [15], [23], none of them accounts for the derived relation between generating and averaged delays. B. Adapting the Dynamics to Different Rates

The packet rate R governs the dynamics of correlated delays, i.e., the speed at which the time series Di takes significantly different values over time. We first assume a constant packet rate R. A generating delay Ci contributes to Di+k with weight (1 − α) · αk , i.e., Ci influences Di+k only αk times as much as it influences Di . The half-life distance kh is the number of averaging steps after which this influence is at most halved, i.e.,  kh = min k ∈ N, k > 0 : αk ≤ 12 . We want to configure a half-life period th , i.e., the time after which the impact of Ci on the averaged delays is halved, which is th = kRh , so that kh = th · R. This yields p α = th ·R 1/2. Thus, we propose to set the correlation weight α depending on a configured half-life period th of the generating delays C and the packet rate R. C. Adapting the Dynamics to Changing Rates

As the packet rate may be unknown and change over time, it must be measured during emulation. We measure it with R = N T , where N is the number of packets observed in an interval of duration T . The intervals need to be successive and nonoverlapping. Whenever an interval closes, R and the weight α are recomputed and applied to all subsequent packets until the next interval closes. To yield accurate results, the interval should be short compared to the time scale on which the rate changes. To keep the estimate well-defined, we count arriving packets and close the current interval with the first packet that arrives after a minimum duration Tmin has elapsed. This packet simultaneously opens the next interval. Since each interval is thus both opened and closed by a packet arrival, at least one packet is always counted, i.e., N ≥ 1, and the measured rate R = N T with T ≥ Tmin is strictly positive by construction. Therefore, no special handling of zero rates is required. We use a default of Tmin = 10 ms in our evaluations. 3

ZINK et al.: DPDS: A DPDK-Based Packet Delayer and Spacer

90

60

30

0 0.1

Jitter

B. Impact of Half-Life Period

We study the impact of the half-life period on delay dynamics and the accuracy of the desired mean and jitter. 4

10

100

0.1 ms

0.3 ms

1 ms

3 ms

(a) Mean delay accuracy. 0

-20

-40

-60

0.1

1

10

100

Rate (Gbit/s) Jitter

0.1 ms

0.3 ms

1 ms

3 ms

(b) Jitter accuracy. Figure 1. Simulated accuracy of uncorrelated, non-reordered varying delays for a desired mean delay of 10 ms across various rates and jitters.

A. Problem of Uncorrelated, Non-Reordered Varying Delays

μ+σ

Delay

The simplest approach to emulate varying packet delays is their generation without any correlation or reordering. However, such uncorrelated delays are inaccurate: even though the delays are generated independently, they still influence each other. If the added delay for a packet is substantially longer than the delay for its successor, the successor packet will face additional delay. This additional queuing delay increases the achieved mean delay and reduces the achieved jitter. Figure 1 shows the relative deviations for uncorrelated delays. The deviation is given in percent, i.e., relative to the desired 10 ms delay or the different jitters. Figure 1a shows that the achieved mean delay is always larger than desired. In contrast, Figure 1b shows that the achieved jitter is always smaller than desired. The largest deviations exceed 110% for the achieved mean delay and reach −70% for the achieved jitter. Both the achieved mean delay and jitter deviate substantially even for a small jitter (0.1 ms) and traffic rate (100 Mbit/s). The deviations worsen for increasing jitter and traffic rate. Uncorrelated delays are therefore unsuitable for emulating varying delays even for small jitter and traffic rate.

1 Rate (Gbit/s)

Jitter deviation (%)

We evaluate the adaptive delay correlation using a discrete event simulation (DES) to avoid potential effects from a realworld implementation. The DES is available on GitHub [37]. In the simulation, constant bit-rate (CBR) traffic with 1518byte packets is delayed by a delay D derived with adaptive delay correlation and a normal distribution. Unless stated otherwise, each experiment takes 10 min and is repeated ten times to compute confidence intervals with a confidence level of 95%. Most confidence intervals are so small that they collapse to a point and are barely visible. We first describe the problem of naively implementing uncorrelated, non-reordered varying delays. Afterward, we study the impact of the half-life period on delay dynamics and accuracy for a mean delay of 10 ms. Then we generalize the findings to other delay values. We compare correlated packet delays with packet reordering. First, we consider unlimited bandwidth for the experiments, i.e., packets can be sent without queuing delay. Second, we take the effect of limited bandwidth on accuracy into account. Both consider correlated packet delays and packet reordering. Finally, we investigate the influence of different underlying distributions on adaptive delay correlation’s accuracy.

Mean delay deviation (%)

IV. Simulation of Varying Delays

μ

μ−σ

t1

t2

t3

t4

t5

Time Delay

Passage time

Figure 2. Definition of passage times.

1) Impact on Delay Dynamics

We study how fast a time series of correlated delays changes over time. To that end, we quantify it by measuring the passage times in these time series. The concept of passage times is illustrated in Figure 2. A passage time is the duration between the time series exceeding an upper threshold and subsequently falling below a lower threshold, and vice versa. In our experiments we set the thresholds to µ − σ and µ + σ . Figure 3 shows that the passage time depends only on the half-life period but not on the traffic rate, which was a design goal. The fact that the passage times are also independent of the desired jitter is caused by the thresholds, which scale with the jitter in our experiments. The result means that the amplitude of the oscillations of the correlated delays increases linearly with the desired jitter, but their frequency remains the same. Thus, the dynamics of correlated delay can be well controlled by the half-life period th . VOLUME ,

Mean delay deviation (ms)

Passage time (ms)

300

100

30 0.1

1

10

0.5 0.4 0.3 0.2 0.1 0.0

100

0.1

1

10

Rate (Gbit/s) Jitter th

100

Rate (Gbit/s)

0.1 ms

0.3 ms

1 ms

7 ms

21 ms

70 ms

3 ms

Jitter

0.1 ms

0.3 ms

Mean delay

3 ms

1 ms

10 ms

3 ms 30 ms

(a) Mean delay accuracy.

Figure 3. Simulated passage times across various half-life periods, rates, and jitters at a desired mean delay of 10 ms. Jitter deviation (ms)

Mean delay deviation (%)

0.0

8 6 4

-0.1

-0.2

-0.3

-0.4

2

0.1

1

10

100

Rate (Gbit/s) 0 0.1

1

10

100

Rate (Gbit/s) Jitter

0.1 ms

Mean delay

0.1 ms

0.3 ms

1 ms

7 ms

21 ms

70 ms

th

Jitter

0.3 ms 3 ms

1 ms

10 ms

3 ms 30 ms

3 ms

(b) Jitter accuracy. Figure 5. Simulated accuracy of adaptive delay correlation (th = 21 ms) across various desired mean delays, rates, and jitters.

(a) Mean delay accuracy.

Jitter deviation (%)

1 0 -1 -2 -3 0.1

1

10

100

Rate (Gbit/s) Jitter th

0.1 ms

0.3 ms

1 ms

7 ms

21 ms

70 ms

3 ms

(b) Jitter accuracy. Figure 4. Simulated accuracy of adaptive delay correlation across various half-life periods, rates, and jitters at a desired mean delay of 10 ms.

passage times sufficiently short for any values of desired jitter. Figure 4b illustrates the deviations of the achieved jitter. The deviations tend to be negative, i.e., achieved jitter is slightly smaller than desired. This is because short delays are likely to be increased by larger delays of preceding packets, leading to smaller jitter than intended. Overall, the observed jitter deviations are small, mostly around zero. However, when the desired jitter is large and the half-life period short, deviations of up to −3% can occur. This problem is avoided with our chosen default value of th = 21 ms. C. Accuracy Independent of Mean Delay

2) Impact on Accuracy

As outlined in Section IV-A, accuracy of the achieved packet delay is a major problem when varying packet delay is desired. Delay correlation significantly mitigates this problem but cannot avoid it. We investigate the influence of the halflife period and desired jitter on the accuracy of the achieved delay’s mean and jitter. Figure 4a shows the relative deviations of the achieved mean delay, which are all positive, i.e., the achieved delay is only increased compared to the desired mean delay. We observe that the deviation is largely independent of the traffic rate. Overall, the deviations are very small, mostly less than 1%, but a short half-life period (7 ms) and large jitter (3 ms) can lead to deviations of 10% (1 ms). While a half-life period of 70 ms minimizes potential inaccuracies, it leads to slower dynamics (see Figure 3). Thus, we recommend a half-life period of th = 21 ms as it leads to good accuracy and keeps VOLUME ,

The accuracy of adaptive delay correlation is essentially independent of the desired mean delay. To demonstrate this, we compare mean delays of 3 ms, 10 ms, and 30 ms. Figure 5 shows the absolute deviations (in ms) of the mean delay and jitter for different combinations of the current rate and jitter. The deviations are largely independent of the desired mean delay and traffic rate. However, the deviation increases with the desired jitter. Thus, the results confirm that adaptive correlation with a half-life period of th = 21 ms achieves the desired mean delay and jitter. There is one exception: when the desired jitter approaches the desired mean (3 ms each). This problem arises because the normal distribution then frequently generates negative delay values, which cannot be realized in practice. This increases the achieved mean delay and decreases the achieved delay jitter compared to the desired values. More complex distributions for delay generation mitigate this problem (see Section IV-F). 5

ZINK et al.: DPDS: A DPDK-Based Packet Delayer and Spacer

Mean delay deviation (%)

Swaps per packet

10 000 1 000 100 10 0

1.5

1.0

0.5

0.0 0.1

1

10

100

0.1

1

10

Rate (Gbit/s) 0.1 ms

Configuration

0.3 ms Reordered

3 ms

Jitter

Correlated

Figure 6. Number of packet swaps required to sort the packets in ascending sequence number order. Simulation results for reordering packets according to their transmission time and adaptive delay correlation (th = 21 ms) at a desired mean delay of 10 ms across various rates and jitters.

D. Comparison with Packet Reordering

Reordering packets according to their transmission time is NetEm’s default configuration. With packet reordering, packets are sent at intended transmission times and do not have to wait for preceding packets to be sent. Therefore, the desired mean delay and jitter can be exactly met at the expense of out-of-order (OOO) delivery. Reordering introduces OOO delivery, which negatively impacts the performance of protocols such as TCP [3]–[5], QUIC [6], and RDMA [7], [8]. The severity of this problem depends on how much reordering occurs. Figure 6 shows the number of swaps per packet required to order them into the original sequence. This is commonly referred to as the Kendall tau or bubble-sort distance (per packet). Adaptive delay correlation produces no reordering by design, whereas the number of swaps per packet for reordering increases with both the current rate and the desired jitter. We therefore conclude that reordering should not be the default strategy for varying delay emulation at high traffic rates and jitters. E. Impact of Limited Bandwidth

Limited bandwidth implies positive transmission times and queuing delays for packets, both of which were neglected in the previous experiments. With limited bandwidth, the achieved packet delay may be increased for both packet reordering and correlated packet delay. We quantify both in the following. We consider different utilizations u ∈ {0, 0.5, 0.8, 0.99}. For u = 0, we reuse the previous values with unlimited bandwidth for comparison. For u > 0, the delayed packets are spaced by a virtual queue with unlimited buffer whose processing rate is set to R/u, where R is the traffic rate.

1) Accuracy of Reordering with Limited Bandwidth

Figure 7 compiles the relative deviations for mean delay and jitter with packet reordering for different jitters, utilizations, and traffic rates. We observe that positive mean delay deviations are clearly visible for 100 Mbit/s and 1 Gbit/s, but hardly beyond (see Figure 7a). These deviations clearly increase with utilization. When the packets of a CBR stream 6

100

Rate (Gbit/s) 1 ms

0.1 ms

Utilization

0%

0.3 ms

1 ms

3 ms

50%

80%

99%

(a) Mean delay accuracy.

Jitter deviation (%)

Jitter

6

4

2

0 0.1

1

10

100

Rate (Gbit/s) Jitter

0.1 ms

Utilization

0%

0.3 ms

1 ms

3 ms

50%

80%

99%

(b) Jitter accuracy. Figure 7. Simulated accuracy of reordering under limited bandwidth across various utilizations, rates, and jitters for a desired mean delay of 10 ms.

receive varying delays, the traffic becomes more bursty. With increasing utilization, the transmission duration of packets increases, which leads to increasing queuing delay for some traffic, causing positive deviations from the desired mean delay. The deviations of the mean delay also increase with the desired jitter because larger jitter leads to larger bursts. In contrast, the deviations of the delay jitter decrease with the desired jitter (see Figure 7b). We observe this because many small delays are extended by queuing, which reduces the delay variation. Further, with packet reordering, the achieved mean delay and jitter may be larger than desired, but only for small traffic rates. For large traffic rates, the reordered traffic stream is not bursty enough to add substantial queuing delay to the traffic so that the desired mean delay and jitter are again almost perfectly met even after queuing.

2) Accuracy of Correlated Delays with Limited Bandwidth

Figure 8 shows the relative deviations for mean delay and jitter with correlated packet delays. The deviations are significantly larger than with packet reordering and do not vanish with increasing traffic rates. Apart from that, they also increase with utilization and desired jitter. The reason for the increased mean delay deviation (see Figure 8a) is the increased queuing delay. With packet delaying and reordering, a CBR traffic stream is turned into a relatively smooth traffic stream which is fed into the virtual queue so that relatively little queuing delay occurs. With correlated packet delays and without reordering, significant packet bursts may occur as differently delayed packets need VOLUME ,

100 000

20 Passage time (ms)

Mean delay deviation (%)

25

15 10 5

10 000

1 000

0 0.1

1

10

100

100

Rate (Gbit/s) Jitter

0.1 ms

Utilization

0%

0.3 ms

1 ms

3 ms

50%

80%

99%

0.1

1

10

100

Rate (Gbit/s) Normal

Distribution

(a) Mean delay accuracy.

Uniform

Gamma

Log-normal

Figure 9. Simulated passage times across various distributions and rates, at a desired mean delay of 10 ms and jitter of 3 ms.

20 3 15 10 5 0 0.1

1

10

100

Rate (Gbit/s) Jitter

0.1 ms

Utilization

0%

0.3 ms

1 ms

3 ms

50%

80%

99%

Mean delay deviation (%)

Jitter deviation (%)

25

2

1

0 0.1 Distribution

We study the impact of the underlying delay distribution on the dynamics, accuracy, and sample generation rate of correlated delays with a desired jitter of 3 ms. The distributions we investigate are the normal, uniform, gamma, and log-normal distribution.

100

Normal

Uniform

Gamma

Log-normal

(a) Mean delay accuracy.

Figure 8. Simulated accuracy of adaptive delay correlation (th = 21 ms) under limited bandwidth across various utilizations, rates, and jitters for a desired mean delay of 10 ms.

F. Impact of Delay Distribution

10 Rate (Gbit/s)

(b) Jitter accuracy.

100 Jitter deviation (%)

to wait for each other, leading to substantially more queuing delay at the virtual queue. This effect increases with queue utilization. The jitter is mostly slightly decreased (see Figure 8b) because small delays are increased by queuing. Only for very high utilization of u = 99% and large jitter (3 ms), the achieved jitter is increased due to excessive queuing delays. In summary, the accuracy of correlated packet delay can be significantly impaired by limited bandwidth, but only if the desired jitter is very large and the utilization is very high. For all other parameter settings, desired mean delay and jitter can be well met with correlated packet delays.

1

50

0

0.1

1

10

100

Rate (Gbit/s) Distribution

Normal

Uniform

Gamma

Log-normal

(b) Jitter accuracy. Figure 10. Simulated accuracy of adaptive delay correlation (th = 21 ms) under various delay distributions and rates, for a desired mean delay of 10 ms and jitter of 3 ms.

observe this because the log-normal distribution has a particularly long tail, especially after the variance scaling applied by adaptive delay correlation. The very high delays that frequently occur add significant queuing delay to successive packets, and smaller delays cannot be realized in practice.

2) Impact on Accuracy 1) Impact on Delay Dynamics

Figure 9 compiles the measured passage times of different distributions. As shown in Section IV-B1, the passage time depends only on the configured half-life period and not on the traffic rate. The passage times of the normal and uniform distributions are similar, while the gamma distribution has a slightly higher one. One exception is the log-normal distribution, whose passage time increases with the traffic rate. For 100 Gbit/s, no passage time is measured during 10 min of simulation. We VOLUME ,

The shape of the underlying delay distribution impacts the accuracy of the achieved mean delay and jitter. Figure 10 compiles the relative deviations of the mean delay and jitter for the different distributions. The mean delay deviations of all four distributions are within the same range, while the jitter deviation of the log-normal distribution is much larger than those of the other distributions. Figure 10a shows that the normal and uniform distributions both achieve a slightly larger mean delay. We observe this because both frequently generate negative delays that cannot be realized in practice (see Section IV-C). In contrast, 7

Sample generation rate (million/s)

ZINK et al.: DPDS: A DPDK-Based Packet Delayer and Spacer

DPDS

200 Receiver

150

Receive packets

Emulator 1 Delayer

2 Limiter

Transmitter 3

Dropper

4 Buffer

5 6

Transmit packets

100

50

Figure 12. General architecture of DPDS, consisting of receiver, emulator, and transmitter components.

0 Normal Distribution

Uniform Gamma Distribution Normal

Uniform

Gamma

Log-normal Log-normal

Figure 11. Simulated sampling rate across various distributions at a desired mean delay of 10 ms and jitter of 3 ms.

the gamma distribution nearly perfectly emulates the desired mean delay with less than 0.1% deviation. The normal, uniform, and gamma distributions achieve a jitter deviation of less than 1% (see Figure 10b). Therefore, the gamma distribution further increases the accuracy of adaptive delay correlation. We do not recommend using the log-normal distribution for adaptive delay correlation. Its realized mean delay is more accurate than those of the normal and uniform distributions (see Figure 10a). However, its realized jitter deviation significantly exceeds those of the other distributions. Further, its large passage times (see Section IV-F1) result in a realized mean delay and jitter that are not converged after 10 min. Therefore, the long tail of the log-normal distribution makes it unsuitable for adaptive delay correlation.

3) Impact on Sample Generation Rate

Generating each delay sample takes some time, which was neglected in the previous experiments. To evaluate this generation time, we generate one billion delay samples and measure the time it takes. The samples are generated on the same server used for the prototype evaluation (see Section VI-A3). Figure 11 compiles the average sample generation rate in million per second. The normal distribution’s generation rate serves as a baseline since most NLEs use normally distributed delays. Uniformly distributed samples are easier to generate, resulting in a generation rate that is 1.5× that of normally distributed ones. In contrast, generating samples from a gamma or log-normal distribution is more complex than from a normal distribution, halving the generation rates. Despite its accuracy advantage, the gamma distribution is not chosen as the default, because this advantage does not justify the resulting generation overhead.

A. General Architecture

DPDS is organized as a three-stage pipeline consisting of a receiver, emulator, and transmitter stage (see Figure 12). The receiver and transmitter stages implement the packet reception and transmission, respectively. The emulator stage implements all supported emulation features, such as packet delay, limited bandwidth, and packet loss. When packets enter DPDS’s emulator stage, they are processed in six steps (see Figure 12): 1 The delays with or without adaptive delay correlation are calculated (see Section V-B3). 2 The packets are either policed (see Section V-C1) or spaced (see Section V-C2) to enforce a bandwidth limit. 3 Packets are dropped according to the configured packet loss model (see Section V-D). 4 The packets’ transmission times are calculated by adding the delay and the spacing offset to the current time. 5 The packets and their respective transmission times are stored in a buffer. 6 Packets are sent once the current time exceeds their transmission time. DPDS can emulate link characteristics on a single port or between two separate ports of a host. It uses software timestamps, eliminating the need for special NIC features. The accuracy of software timestamps is evaluated in Section VIC. To enable high emulation rates, DPDS processes packets in batches of up to 16, which increases throughput at the cost of reduced delay accuracy and increased burstiness. Additionally, the receiver, emulator, and transmitter (Figure 12) can run as a single thread or as up to three threads on separate CPU cores. B. Packet Delays

DPDS implements packet delays by using a parameterized distribution combined with either reordering or adaptive delay correlation. In the following, we describe which delay distributions DPDS supports. Then, we explain how DPDS implements reordering and adaptive delay correlation.

V. The DPDK-Based Packet Delayer and Spacer (DPDS)

We developed the DPDK-based packet delayer and spacer (DPDS) for link emulation with a kernel bypass approach. DPDS emulates constant and varying delays, bandwidth limitation, and packet loss. It applies those characteristics to a single link regardless of the number of flows. DPDS is open source on GitHub [38]. 8

1) Supported Delay Distributions

DPDS supports emulating both constant and varying delays. Constant delays are implemented using a deterministic distribution. Varying delays are implemented using one of the following distributions: uniform, normal, log-normal, and VOLUME ,

gamma. These distributions are parameterized using a mean delay and jitter. The uniform and normal distributions are supported to match those of existing NLEs. The gamma distribution is supported since Mukherjee [39] showed that Internet delays follow a shifted gamma distribution.

2) Reordering

DPDS supports packet reordering by sorting packets by their transmission time. This is implemented using a binary heap for the packet buffer, which automatically sorts the packets as they are pushed into the data structure. While this enables accurate emulation of varying delays, it also introduces significant sorting overhead. This overhead makes reordering infeasible for high rates (see Section VI-B). Furthermore, it introduces OOO delivery, making it unsuitable for protocols like TCP, QUIC, and RDMA. If reordering is disabled, the packet buffer is implemented as a simple first in, first out queue.

bucket fill state represents the number of packets currently scheduled for transmission rather than the available tokens, i.e., the bandwidth budget. Each packet arrives at the spacer with a transmission time determined by the delayer (see Section V-B). The spacer then adds an additional offset to this time so that the resulting transmission time adheres to the configured rate. This offset is derived from the current fill state and the configured rate. To prevent unbounded delays, the fill state is subject to a configurable upper limit. Packets that would exceed this limit are dropped. D. Packet Loss

DPDS currently supports two different loss models: independent packet losses and the Gilbert-Elliott model [42]– [44]. For independent losses, each packet has an independent probability p of being lost. In contrast, the Gilbert-Elliott model uses a two-state Markov chain, with state-dependent loss probabilities, to emulate correlated packet losses. VI. Evaluation of the DPDS Prototype

3) Adaptive Delay Correlation

DPDS implements adaptive delay correlation as described in Section III. To this end, it correlates successive delays through an EMA whose weight is recomputed periodically from a short-window rate estimate. Unlike reordering, this preserves packet order and avoids sorting overhead, enabling higher forwarding rates (see Section VI-B). Adaptive delay correlation is parameterized by the mean delay µ and jitter σ of the underlying distribution, as well as the half-life period th . The resulting accuracy of DPDS is evaluated under both constant (Section VI-C) and changing (Section VI-D) rates.

We experimentally evaluate DPDS along two dimensions. First, we measure DPDS’s maximum forwarding rate and ZLT and compare them with NetEm and MoonEm. Those serve as baselines for in-kernel and kernel bypass-based emulators, respectively. Second, we evaluate DPDS’s accuracy and compare the results with our simulation predictions, including changing traffic rates. A. Methodology for Measuring Emulator Accuracy

In the following, we describe our experimental setup to facilitate reproducibility, consisting of the traffic generator P4TG, the measurement setup, and the testbed configuration.

C. Bandwidth Limitation

DPDS implements bandwidth limitation by either policing or spacing packets. With policing, packets that exceed the configured rate are dropped. In contrast, spacing delays packets to meet the configured rate.

1) Policing

Policing is implemented using the token bucket algorithm [40]. In this algorithm, each bit corresponds to a token. The bucket fill state increases at the configured token rate in proportion to the time elapsed since the last increase. However, the fill state cannot exceed a configured threshold. When a packet is processed and there are enough tokens available, the fill state decreases according to the packet’s length in bits (tokens). If a packet’s length exceeds the fill state, the packet is dropped.

2) Spacing

Spacing is implemented using the leaky bucket algorithm [41]. In contrast to the token bucket algorithm, the VOLUME ,

1) The Traffic Generator P4TG

P4TG [45] is a P4-based, hardware-accelerated traffic generator implemented on the Intel Tofino™ switching ASIC. It supports aggregate generation rates of up to 4 Tbit/s [46]. It performs round trip time (RTT) measurements with nanosecond granularity directly in the data plane, enabling precise calculation of the mean, standard deviation, and percentiles without sampling bias [47]. P4TG also supports automated ZLT determination [48]. The ZLT is the maximum rate at which a device under test forwards packets without loss, as defined in RFC 2544 [49]. P4TG’s source code is available on GitHub [50].

2) Measurement Setup

Figure 13 depicts the general measurement setup which can be used to evaluate an emulator’s accuracy. P4TG is directly connected to the emulator under test with a single link. In the first step 1 , P4TG measures the ZLT and maximum forwarding rate of the emulators for the evaluated configuration. In contrast to the ZLT, which is the highest 9

ZINK et al.: DPDS: A DPDK-Based Packet Delayer and Spacer

P4TG

1

Rmax RZLT

0

p

0

t0

RTT histogram

3

tmax

NIC

DPDS.

2

p

Link

tloss tp

tmin

Table 2. Measured maximum forwarding rates of NetEm, MoonEm, and

Emulator

0

Mean delay

Bandwidth limitation

Jitter 0 ms

3 ms

NetEm NetEm + TBF

5.1 Gbit/s 4.7 Gbit/s

4.7 Gbit/s 4.3 Gbit/s

MoonEm DPDS (reordered) DPDS (corr.) DPDS (corr. + spaced)

46.0 Gbit/s 78.5 Gbit/s 95.5 Gbit/s 94.8 Gbit/s

× 75.8 Gbit/s 93.6 Gbit/s 94.9 Gbit/s

Figure 13. Setup for measuring an emulator’s accuracy.

rate at which no packets are lost, the maximum forwarding rate is the highest output rate an emulator sustains under load. In the second step 2 , P4TG generates traffic and forwards it to the emulator under test. Based on the rates of the first step, we ensure not to overload the emulator. The emulator applies the desired link characteristics, such as packet delay and bandwidth limitation. After applying these characteristics, the emulator forwards the traffic back to P4TG. In the third step 3 , P4TG uses the RTT histogram to measure the emulated mean delay and jitter. This histogram is then exported via P4TG’s REST API for further use.

baseline for in-kernel Linux QDisc emulators, and MoonEm serves as a baseline for kernel bypass-based emulators. Other emulators in each class (e.g., 6GDetCom Emulator, TheaterQ, DEMU, SmartNet) may differ in performance due to architectural choices such as SmartNIC offloading. However, benchmarking them is beyond the scope of this evaluation.

1) Forwarding Rate

In the following, we compare the measured maximum forwarding rates of DPDS with those of NetEm and MoonEm. Table 2 presents the measured maximum forwarding rates for jitters of 0 ms and 3 ms.

3) Testbed Specifications and Configuration

a: NetEm

We evaluate DPDS using the measurement setup described above. P4TG v2.7.1 runs on an Intel(R) Tofino 1 and generates the traffic. DPDS runs on a bare-metal server with an eight-core Intel(R) Xeon(R) Gold 6134 CPU @ 3.20 GHz, four 32 GB DIMM DDR4 RAM modules, and one NUMA node. The server is equipped with one dual-port (100 Gbit/s per port) ConnectX-5 MCX516A-CCAT NIC with PCIe Gen4 and the MLNX OFED 23.10 driver. For efficient packet processing, the server uses 64 × 1 GB huge pages, isolated CPU cores, and a PCIe maximum read request size of 1024 B. Its operating system is Ubuntu 24.04.4 LTS with kernel 6.8.0-100-generic, and we use DPDK v25.11 and Rust v1.95. We configure all emulators under test with a mean delay of 10 ms and a normal and deterministic distribution for varying and constant delay, respectively. For DPDS, when spacing is active, we set the spacing rate to 100 Gbit/s, matching the maximum supported rate of the NIC. DPDS and NetEm run on a single, isolated CPU core, whereas MoonEm requires two CPU cores by design. We run each experiment with 1518-byte packets for 4 min and repeat it ten times to compute confidence intervals with a confidence level of 95%.

NetEm’s maximum forwarding rate is 5.1 Gbit/s and 4.7 Gbit/s for 0 ms and 3 ms jitter, respectively. The lower rate at 3 ms jitter is due to the additional sorting overhead that NetEm incurs for varying delays. A further decrease occurs when NetEm is combined with TBF at a spacing rate of 10 Gbit/s: the rates drop to 4.7 Gbit/s and 4.3 Gbit/s for 0 ms and 3 ms jitter, respectively. This decrease is due to the additional per-packet processing overhead from calculating the spacing offset and applying it.

B. Maximum Supported Rates

We compare the maximum forwarding rate and ZLT of DPDS with those of NetEm and MoonEm. NetEm serves as a 10

b: MoonEm

Since MoonEm is a constant-delay emulator, we evaluate it only for a jitter of 0 ms. We also configure MoonEm without bandwidth limitation or packet loss. MoonEm’s authors report NIC-dependent forwarding rates (Intel E810 vs. NVIDIA ConnectX-5). Our measurement therefore serves as a baseline for DPDK-based emulators on the ConnectX-5 NIC. We measured a maximum forwarding rate of MoonEm of 46 Gbit/s, approximately 9 times NetEm’s rate. This shows that the kernel bypass approach outperforms the in-kernel approach of NetEm. c: DPDS

The maximum forwarding rate of DPDS using reordering is 78.5 Gbit/s and 75.8 Gbit/s for 0 ms and 3 ms jitter, respectively. By contrast, adaptive delay correlation reaches 95.5 Gbit/s and 93.6 Gbit/s for the same jitters, which is VOLUME ,

Jitter 0 ms

3 ms

NetEm NetEm + TBF MoonEm DPDS (reordered) DPDS (corr.)

5.0 Gbit/s 4.6 Gbit/s 44.8 Gbit/s 73.5 Gbit/s 95.0 Gbit/s

4.5 Gbit/s 4.2 Gbit/s × 58.8 Gbit/s 60.8 Gbit/s

DPDS (corr. + spaced)

94.4 Gbit/s

85.2 Gbit/s

Mean delay deviation (%)

Table 3. Measured zero-loss throughput of NetEm, MoonEm, and DPDS.

30

20

10

0 60

70

80

90

Rate (Gbit/s) Jitter

0 ms

Configuration

0.1 ms

Reordered

0.3 ms

1 ms

Correlated

3 ms

Correlated + spaced

(a) Mean delay accuracy.

2) Zero-Loss Throughput

First, we compare the measured ZLT of NetEm, MoonEm, and DPDS in different configurations. Second, we explain why spacing increases the ZLT of adaptive delay correlation. For NetEm and MoonEm, we observe a transient phase with packet loss at the beginning of each measurement. The reported ZLT therefore refers to the steady-state throughput after this phase without further packet loss. DPDS does not exhibit such a transient phase. a: Comparison of DPDS with NetEm and MoonEm

Table 3 shows that DPDS achieves substantially higher ZLTs than NetEm and MoonEm across all jitter configurations. For a jitter of 0 ms, DPDS with adaptive delay correlation reaches a ZLT of 95.0 Gbit/s, 19 times that of NetEm and over 2 times that of MoonEm. At a jitter of 3 ms, DPDS with adaptive delay correlation reaches 60.8 Gbit/s, still over 13 times that of NetEm. MoonEm does not support varying delays and is therefore omitted at this jitter. DPDS with reordering achieves 73.5 Gbit/s and 58.8 Gbit/s for jitters of 0 ms and 3 ms, respectively, also clearly exceeding NetEm and MoonEm. Adding TBF to NetEm decreases its ZLT slightly due to the additional processing overhead. However, VOLUME ,

Jitter deviation (%)

800

approximately 21.6% and 23.5% higher. The difference is due to the sorting overhead of reordering, which is avoided by adaptive delay correlation. For a jitter of 3 ms, spacing slightly raises the maximum forwarding rate for adaptive delay correlation (from 93.6 Gbit/s to 94.9 Gbit/s). This is because spacing prevents the NIC from being overloaded (see Section VI-B2). Thus, correlated varying delays are more efficient for high-rate emulation than reordering. DPDS with packet reordering implements the same strategy as NetEm. However, DPDS’s maximum forwarding rate using reordering is 15.4 times that of NetEm and approximately 1.7 times that of MoonEm. As with MoonEm, we attribute this to the different implementation approaches: inkernel versus kernel bypass. Furthermore, DPDS’s maximum forwarding rate with adaptive delay correlation is 18.7 times that of NetEm and over twice that of MoonEm. Overall, DPDS outperforms both baselines through a combination of its kernel bypass implementation and adaptive delay correlation, a reordering-free emulation strategy.

600 400 200 0 60

70

80

90

Rate (Gbit/s) Jitter Configuration

0.1 ms Reordered

0.3 ms

1 ms

Correlated

3 ms

Correlated + spaced

(b) Jitter accuracy. Figure 14. Measured accuracy of DPDS for reordered and correlated (th = 21 ms, with and without spacing) delays at a desired mean delay of 10 ms across various rates and jitters.

at NetEm’s low rates, the burst-smoothing benefit of spacing observed for DPDS does not materialize. b: Effect of Spacing on Adaptive Delay Correlation’s Zero-Loss Throughput

For constant delays (jitter of 0 ms), spacing slightly decreases DPDS’s (corr.) ZLT from 95.0 Gbit/s to 94.4 Gbit/s due to the additional processing overhead. For varying delays (jitter of 3 ms), however, spacing increases the ZLT from 60.8 Gbit/s to 85.2 Gbit/s, an improvement of approximately 40%. This practically confirms an effect already observed in the spacing simulation (see Section IV-E2): adaptive delay correlation inherently produces packet bursts when successive delays are similar. At low rates these bursts fit within the NIC’s buffer, but at high rates they exceed its size and cause packet loss. Applying spacing before the NIC receives the packets from the emulator smooths these bursts so they fit within the available buffer, preventing such overload. Therefore, despite the small overhead in the constant-delay case, spacing substantially increases DPDS’s (corr.) ZLT for varying delays. C. Accuracy of Delay Strategies

We evaluate the accuracy of the DPDS prototype for reordering and adaptive delay correlation at rates not exceeding their respective maximum forwarding rates. Figure 14 shows the accuracy of the DPDS prototype for different combinations of emulation configuration, rate, and jitter. The measured accuracies confirm the simulation results from Section IVD. 11

ZINK et al.: DPDS: A DPDK-Based Packet Delayer and Spacer

1) Accuracy with Reordering

Table 4. Measured accuracy of DPDS’s adaptive delay correlation (th =

Figure 14a shows the mean delay deviations for rates below the maximum forwarding rate with DPDS and reordering. The deviations of reordering are greater than those observed in the simulation. This is because DPDS processes up to 16 packets at a time, resulting in packet bursts. Additionally, the NIC enforces a minimum inter-arrival time (IAT) between successive packets. Together, these effects prevent DPDS with reordering from realizing the fine-grained delays observed in simulation. For jitters of 0.1 ms and 0.3 ms at 70 Gbit/s, the jitter deviations exceed those of the simulations (see Figure 14b). This is due to DPDS’s batch processing and the minimum IAT enforced by the NIC, both of which become more pronounced as the traffic rate approaches the ZLT of DPDS with reordering. At small desired jitters, even modest absolute deviations translate into large relative deviations, amplifying the visible error.

21 ms) at a CBR and sinusoidal traffic pattern at a desired mean delay of 10 ms across various jitters.

Jitter

CBR

Sine

0.1 ms

Mean delay deviation (0.33 ± 0.00) % (0.29 ± 0.00) %

0.3 ms 1 ms 3 ms

(0.52 ± 0.02) % (1.17 ± 0.07) % (2.97 ± 0.44) %

0.1 ms 0.3 ms 1 ms

(0.23 ± 0.13) % (−0.45 ± 0.25) % (−0.49 ± 0.29) %

(0.64 ± 0.20) % (−0.19 ± 0.22) % (−0.20 ± 0.37) %

3 ms

(−0.55 ± 0.55) %

(−0.88 ± 0.98) %

(0.47 ± 0.01) % (1.10 ± 0.13) % (2.74 ± 0.49) % Jitter deviation

each rate remains representative for a meaningful number of packets.

2) Accuracy with Adaptive Delay Correlation

Figure 14 shows that for correlation, the mean delay and jitter deviations match those observed in the simulation for limited bandwidth (see Section IV-E2). As with reordering, the deviations are slightly increased due to DPDS’s batch processing. In some configurations, such as a jitter of 0.1 ms, adaptive delay correlation is even more accurate than reordering, suggesting that it is less sensitive to batch processing. The deviations further increase as the traffic rate approaches the adaptive delay correlation’s ZLT. Figure 14 shows that DPDS’s additional spacing has a negligible effect on accuracy, since the NIC already spaces the packets to its bandwidth. The only exception occurs for a jitter of 3 ms at 90 Gbit/s. Although DPDS spaces the packets, the packets still arrive at the NIC in bursts due to the maximum burst size of 16 packets. The NIC therefore has to smooth these bursts itself, which adds queuing delay and increases the realized mean delay. D. Adaptation to a Changing Rate

We show that adaptive delay correlation maintains accurate delay emulation even under changing traffic rates.

1) Configuration of a Changing Rate

To evaluate DPDS under a changing rate, we configure P4TG to generate a sinusoidal traffic pattern. The sinusoidal shape continuously exposes the emulator to both increasing and decreasing rates, which stresses the rate estimation in both directions. The pattern has a maximum rate of 50 Gbit/s and a period of 30 s. Over a measurement duration of 4 min, the traffic therefore goes through eight full periods, with the rate changing by approximately 3.3 Gbit/s2 on average between its minimum and maximum. This rate of change is fast enough to challenge the rate estimation but slow enough that 12

2) Accuracy at a Changing Rate

Table 4 lists the mean delay and jitter deviations at the end of each experiment. Further, it includes the deviations for CBR traffic at 50 Gbit/s as a reference. Overall, the deviations for both traffic patterns are within the same range and grow with increasing jitter. For the mean delay, the sinusoidal pattern consistently deviates less than CBR, while neither pattern is consistently better than the other for the jitter. Note that the sinusoidal pattern has a lower average packet rate than CBR at 50 Gbit/s, putting less load on the emulator. The similar deviations therefore primarily indicate that adaptive delay correlation’s rate estimation works as intended. This confirms that adaptive delay correlation with a rate-dependent EMA is suitable for emulating varying delays at changing traffic rates. VII. Conclusions

In this paper, we tackled the problem of adding varying delay to packets for link emulation. The challenge is that varying delay may add more delay to packets than intended when preceding packets are extensively delayed. Packet reordering solves this problem, but is detrimental to many applications and should be avoided. We proposed adaptive delay correlation to largely mitigate this problem. It essentially generates independent delay values and smooths them with an EMA. The algorithm takes a desired mean delay and standard deviation (jitter) as input, as well as a half-life period to control delay dynamics over time. Short-term rate measurements adapt the EMA’s weight to different traffic rates. We investigated adaptive delay correlation by simulation and showed that a half-life period of 21 ms achieves mean delay and jitter as desired while avoiding overly slow delay dynamics. Bandwidth limitation deteriorates the accuracy of VOLUME ,

mean delay and jitter at high utilization. These problems are avoided with packet reordering, but the reordering effect is substantial and undesired. We further developed a DPDK-based packet delayer and spacer (DPDS), which is a kernel bypass approach for link emulation. It implements adaptive delay correlation, packet reordering, spacing, policing, and two models for packet loss. Our performance evaluation showed that DPDS achieves a ZLT of 95 Gbit/s for constant delay and, with spacing enabled, 85 Gbit/s for varying delay with a jitter of 3 ms while meeting the desired delay. With these values, DPDS clearly outperforms the widely used NLE NetEm and the recently developed DPDK-based emulator MoonEm. With packet reordering, DPDS achieves ZLTs of only 73 Gbit/s for constant and 58 Gbit/s for varying delay, which still exceed those of NetEm and MoonEm. As adaptive delay correlation relies on continuous rate measurement, we also showed that DPDS works as desired under changing rates. DPDS is open source and available on GitHub [38]. List of Abbreviations

CBR constant bit-rate DES discrete event simulation DPDS DPDK-based packet delayer and spacer EMA exponential moving average IAT inter-arrival time NIC network interface card NLE network link emulator OOO out-of-order QDisc queueing discipline RTT round trip time TBF token bucket filter VNE virtual network emulator ZLT zero-loss throughput References [1] V. Paxson, “End-To-End Internet Packet Dynamics,” IEEE/ACM Transactions on Networking, vol. 7, pp. 277–292, June 1999. [2] T. Høiland-Jørgensen, B. Ahlgren, P. Hurtig, and A. Brunstrom, “Measuring Latency Variation in the Internet,” in ACM Conference on emerging Networking EXperiments and Technologies (CoNEXT), pp. 473–480, Dec. 2016. [3] E. Blanton and M. Allman, “On Making TCP More Robust to Packet Reordering,” ACM SIGCOMM Computer Communication Review, vol. 32, pp. 20–30, Jan. 2002. [4] M. Zhang, B. Karp, S. Floyd, and L. Peterson, “Improving TCP’s Performance Under Reordering with DSACK,” ICSI Technical Report, Aug. 2002. [5] M. Laor and L. Gendel, “The Effect of Packet Reordering in a Backbone Link on Application Throughput,” IEEE Network Magazine, vol. 16, pp. 28–36, Oct. 2002. [6] A. M. Kakhki, S. Jero, D. Choffnes, C. Nita-Rotaru, and A. Mislove, “Taking a Long Look at QUIC: An Approach for Rigorous Evaluation of Rapidly Evolving Transport Protocols,” in ACM Internet Measurements Conference (IMC), pp. 290–303, Nov. 2017. [7] R. Mittal, A. Shpiner, A. Panda, E. Zahavi, A. Krishnamurthy, S. Ratnasamy, and S. Shenker, “Revisiting Network Support for RDMA,” in ACM SIGCOMM, pp. 313–326, Aug. 2018. [8] C. H. Song, X. Z. Khooi, R. Joshi, I. Choi, J. Li, and M. C. Chan, “Network Load Balancing with in-Network Reordering Support for RDMA,” in ACM SIGCOMM, pp. 816–831, Sept. 2023.

VOLUME ,

[9] L. Nussbaum and O. Richard, “A Comparative Study of Network Link Emulators,” in Spring Simulation Multiconference, 2009. [10] N. Handigol, B. Heller, V. Jeyakumar, B. Lantz, and N. McKeown, “Reproducible Network Experiments Using Container-Based Emulation,” in ACM Conference on emerging Networking EXperiments and Technologies (CoNEXT), pp. 253–264, 2012. [11] M. Peuster, J. Kampmeyer, and H. Karl, “Containernet 2.0: A Rapid Prototyping Platform for Hybrid Service Function Chains,” in IEEE Conference on Network Softwarization (NetSoft), pp. 335–337, June 2018. [12] P. Gouveia, J. a. Neves, C. Segarra, L. Liechti, S. Issa, V. Schiavoni, and M. Matos, “Kollaps: Decentralized and Dynamic Topology Emulation,” in European Conference on Computer Systems, Apr. 2020. [13] J. Gomez, E. F. Kfoury, J. Crichigno, and G. Srivastava, “A Survey on Network Simulators, Emulators, and Testbeds Used for Research and Education,” Computer Networks, vol. 237, Dec. 2023. [14] L. Rizzo, “Dummynet: A Simple Approach to the Evaluation of Network Protocols,” ACM SIGCOMM Computer Communication Review, vol. 27, pp. 31–41, Jan. 1997. [15] S. Hemminger, “Network Emulation with NetEm,” Linux Conf Au, Apr. 2005. [16] S. Aketa, T. Hirofuchi, and R. Takano, “DEMU: A DPDK-Based Network Latency Emulator,” in IEEE Workshop on Local & Metropolitan Area Networks (LANMAN), June 2017. [17] S. Lachnit, S. Gallenmüller, E. Hauser, F. Wiedner, K. Holzinger, H. Stubbe, T. Senftl, and G. Carle, “MoonEm – High-Precision Path Property Emulation Using DPDK,” in ACM Conference on emerging Networking EXperiments and Technologies (CoNEXT), Dec. 2025. [18] S. Becker, T. Pfandzelter, N. Japke, D. Bermbach, and O. Kao, “Network Emulation in Large-Scale Virtual Edge Testbeds: A Note of Caution and the Way Forward,” in IEEE International Conference on Cloud Engineering, Sept. 2022. [19] L. Haug, F. Dürr, S. Egger, L. Grohmann, J. Gross, G. P. Sharma, and J. Sachs, “Simulating and Emulating the Characteristic Packet Delay of Logical 5g TSN Bridges,” in KuVS Workshop on Network Softwarization (KuVS NetSoft), Apr. 2025. [20] M. Wang, Y. Shen, B. Wang, H. Tong, Y. Xie, Y. Gao, Y. Liu, L. Chen, M. Xu, and J. Wu, “Rattan: An Extensible and Scalable Modular Internet Path Emulator,” July 2025. arXiv Preprint. [21] M. Ottens, K.-S. Hielscher, and R. German, “TheaterQ: A Qdisc for Dynamic Network Emulation,” Oct. 2025. arXiv Preprint. [22] L. Stratmann, B. Walker, and V. A. Vu, “Realistic Emulation of LTE with MoonGen and DPDK,” in International Workshop on Wireless Network Testbeds, Experimental Evaluation and Characterization (WiNTECH), pp. 87–94, Sept. 2020. [23] F. G. Vogt, C. Rothenberg, V. H. S. Lopes, M. C. Luizelli, F. Rodriguez, C. Papagianni, and G. Pongrácz, “SmartNet: Bridging Performance and Realism in Network Emulation with SmartNICs,” in ACM SIGCOMM Posters and Demos, pp. 184–186, 2025. [24] R. Lübke, P. Buschel, D. Schuster, and A. Schill, “Measuring Accuracy and Performance of Network Emulators,” in IEEE International Black Sea Conference on Communications and Networking (BlackSeaCom), pp. 63–65, May 2014. [25] M. Carbone and L. Rizzo, “Dummynet Revisited,” ACM SIGCOMM Computer Communication Review, vol. 40, pp. 12–20, Apr. 2010. [26] A. N. Kuznetsov and B. Hubert, “TBF Man Page.” https://www.man7 .org/linux/man-pages/man8/tc-tbf.8.html, Dec. 2001. [27] L. Haug, “Simulating and Emulating the Characteristic Packet Delay of Logical 5G TSN Bridges (slides).” https://uni-tuebingen.de/fakultaete n/mathematisch-naturwissenschaftliche-fakultaet/fachbereiche/inform atik/lehrstuehle/kommunikationsnetze/kuvs-fg-netsoft/2025/program/, Apr. 2025. visited on 2026-01-29. [28] M. Ottens, “GitHub: TheaterQ.” https://github.com/cs7org/TheaterQ, Oct. 2025. visited on 2026-01-29. [29] The Linux Foundation, “DPDK.” https://www.dpdk.org/, Oct. 2025. visited on 2026-01-29. [30] M. Paolino, N. Nikolaev, J. Fanguede, and D. Raho, “SnabbSwitch User Space Virtual Switch Benchmark and Performance Optimization for NFV,” in IEEE Conference on Network Function Virtualization and Software-Defined Networking (NFV-SDN), pp. 86–92, Nov. 2015. [31] P. Emmerich, M. Pudelko, S. Bauer, S. Huber, T. Zwickl, and G. Carle, “User Space Network Drivers,” in ACM/IEEE Symposium on Architectures for Networking and Communications Systems (ANCS), Sept. 2019.

13

ZINK et al.: DPDS: A DPDK-Based Packet Delayer and Spacer

[32] K. Sasaki, T. Hirofuchi, S. Yamaguchi, and R. Takano, “An Accurate Packet Loss Emulation on a DPDK-Based Network Emulator,” in Asian Internet Engineering Conference (AINTEC), Aug. 2019. [33] C. Puakalong, R. Takano, V. Visoottiviseth, A. Khurat, and W. Sawangphol, “A Network Bandwidth Limitation with the DEMU Network Emulator,” in IEEE Symposium on Computer Applications and Industrial Electronics, pp. 151–154, Apr. 2020. [34] P. Emmerich, S. Gallenmüller, D. Raumer, F. Wohlfart, and G. Carle, “MoonGen: A Scriptable High-Speed Packet Generator,” in ACM Internet Measurements Conference (IMC), pp. 275–287, Oct. 2015. [35] V. H. S. Lopes and F. G. Vogt, “GitHub: SmartNet.” https://github.c om/intrig-unicamp/SmartNet, Sept. 2025. visited on 2026-01-29. [36] M. Menth and F. Hauser, “On Moving Averages, Histograms and Time-DependentRates for Online Measurement,” in ACM/SPEC International Conference on Performance Engineering (ICPE), pp. 103– 114, Apr. 2017. [37] E. Zink, “GitHub: DPDS-Simulation.” https://github.com/uni-tue-kn/ DPDS-Simulation, Mar. 2026. visited on 2026-06-15. [38] E. Zink, “GitHub: DPDS.” https://github.com/uni-tue-kn/DPDS, Mar. 2026. visited on 2026-06-15. [39] A. Mukherjee, “On the Dynamics and Significance of Low Frequency Components of Internet Load,” Internetworking: Research and Experience, vol. 5, pp. 163–205, Dec. 1994. [40] J. Heinanen and R. Guerin, “RFC2697: A Two Rate Three Color Marker,” Sept. 1999. [41] J. S. Turner, “New Directions in Communications (or Which Way to the Information Age?),” IEEE Communications Magazine, vol. 24, pp. 8–15, Oct. 1986. [42] E. N. Gilbert, “Capacity of a Burst-Noise Channel,” Bell Systems Technical Journal, vol. 39, pp. 1253–1265, Sept. 1960. [43] E. O. Elliott, “Estimates of Error Rates for Codes on Burst-Noise Channels,” Bell Systems Technical Journal, vol. 42, pp. 1977–1997, Sept. 1963. [44] G. Hasslinger and O. Hohlfeld, “The Gilbert-Elliott Model for Packet Loss in Real Time Services on the Internet,” in GI/ITG Conference on Measuring, Modelling and Evaluation of Computer and Communication Systems (MMB), June 2008. [45] S. Lindner, M. Häberle, and M. Menth, “P4TG: 1 Tb/s Traffic Generation for Ethernet/IP Networks,” IEEE Access, vol. 11, pp. 17525– 17535, Feb. 2023. [46] F. Ihle, E. Zink, S. Lindner, and M. Menth, “Enhancements to P4TG: Protocols, Performance, and Automation,” in KuVS Workshop on Network Softwarization (KuVS NetSoft), Apr. 2025. [47] F. Ihle, E. Zink, and M. Menth, “Enhancements to P4TG: HistogramBased RTT Monitoring in the Data Plane,” in Workshop on Resilient Networks and Systems (ReNeSys), pp. 26–29, Sept. 2025. [48] F. Ihle, E. Zink, S. Lindner, and M. Menth, “High-Speed Generation of Periodic Traffic Patterns on P4TG for DDoS and Burst-Load Evaluation,” in IEEE Conference on Network Softwarization (NetSoft), July 2026. [49] S. Bradner and J. McQuaid, “Benchmarking Methodology for Network Interconnect Devices,” Mar. 1999. [50] S. Lindner, F. Ihle, and E. Zink, “GitHub: P4TG.” https://github.com /uni-tue-kn/P4TG, Nov. 2025. visited on 2026-01-29.

Fabian Ihle received his bachelor’s (2021) and master’s degrees (2023) in computer science from the University of Tübingen. Afterward, he joined the communication networks research group of Prof. Dr. habil. Michael Menth as a Ph.D. student. His research interests include software-defined networking, P4-based data plane programming, resilience, and Time-Sensitive Networking (TSN).

Michael Menth (Senior Member, IEEE) is a professor at the Department of Computer Science at the University of Tübingen/Germany and chairholder of Communication Networks since 2010. He studied, worked, and obtained diploma (1998), PhD (2004), and habilitation (2010) degrees at the universities of Austin/Texas, Ulm/Germany, and Würzburg/Germany. His special interests are performance analysis and optimization of communication networks, resilience and routing issues, as well as resource and congestion management. His recent research focus is on network softwarization, in particular P4-based data plane programming, Time-Sensitive Networking (TSN), Internet of Things, and Internet protocols. Dr. Menth contributes to standardization bodies, notably to the IETF.

Etienne Zink received his bachelor’s degree (2022) from the Corporate State University BadenWürttemberg and his master’s degree (2024) from the University of Tübingen, both in computer science. Afterward, he joined the communication networks research group of Prof. Dr. habil. Michael Menth as a Ph.D. student. His research interests include software-defined networking, network function virtualization, resilience, and network emulation.

14

VOLUME ,

Record · ID 282749 · SHA-256 2702e0ac83b9bf9b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.