arXiv:2609.19823v1 [cs.NI] 17 Sep 2026
PyStream: Enhancing Video Streaming Evaluation Samuel Radler∗
Leon Prüller∗
Emanuele Artioli
Alpen-Adria-Universität Klagenfurt, Christian Doppler Laboratory ATHENA Klagenfurt, Austria
Alpen-Adria-Universität Klagenfurt, Christian Doppler Laboratory ATHENA Klagenfurt, Austria
Alpen-Adria-Universität Klagenfurt, Christian Doppler Laboratory ATHENA Klagenfurt, Austria
Farzad Tashtarian
Christian Timmerer
Alpen-Adria-Universität Klagenfurt, Christian Doppler Laboratory ATHENA Klagenfurt, Austria
Alpen-Adria-Universität Klagenfurt, Christian Doppler Laboratory ATHENA Klagenfurt, Austria
Abstract As streaming services become more commonplace, analyzing their behavior effectively under different network conditions is crucial. This is normally quite expensive, requiring multiple players with different bandwidth configurations to be emulated by a powerful local machine or a cloud environment. Furthermore, emulating a realistic network behavior or guaranteeing adherence to a real network trace is challenging. This paper presents PyStream, a simple yet powerful way to emulate a video streaming network, allowing multiple simultaneous tests to run locally. By leveraging a network of Docker containers, many of the implementation challenges are abstracted away, keeping the resulting system easily manageable and upgradeable. We demonstrate how PyStream not only reduces the requirements for testing a video streaming system but also improves the accuracy of the emulations with respect to the current state-of-the-art. On average, PyStream reduces the error between the original network trace and the bandwidth emulated by video players by a factor of 2-3 compared to Wondershaper, a common network traffic shaper in many video streaming evaluation environments. Moreover, PyStream decreases the cost of running experiments compared to existing cloud-based video streaming evaluation environments such as CAdViSE.
CCS Concepts • Networks → Network simulations; Network performance modeling; Network experimentation; Network performance analysis.
Keywords Video Streaming Evaluation, HTTP Traffic Shaper, Evaluation Cost ACM Reference Format: Samuel Radler, Leon Prüller, Emanuele Artioli, Farzad Tashtarian, and Christian Timmerer. 2024. PyStream: Enhancing Video Streaming Evaluation. In ∗ Both authors contributed equally to this research.
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). MMSys ’24, Bari, Italy © 2024 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-0412-3/24/04 https://doi.org/10.1145/3625468.3652194
ACM Multimedia Systems Conference 2024 (MMSys ’24), April 15–18, 2024, Bari, Italy. ACM, New York, NY, USA, 7 pages. https://doi.org/10.1145/3625 468.3652194
1
Introduction
Video streaming is by far the most data-intensive application of the Internet, accounting for approximately 65% of the global network traffic [1]. Therefore, it is crucial to develop tools for bandwidth management, as any 1% increase in video streaming efficiency will save almost 1 billion TBs of Internet traffic [2]. Adaptive video streaming is the current state-of-the-art multimedia delivery mechanism designed to optimize the viewer’s experience by dynamically adjusting video quality based on real-time network conditions and device capabilities [3]. At its core, this technology leverages HTTP-based protocols and employs a segmented representation approach [4]. Video content is encoded into discrete segments, each representing a fraction of the entire video stream. Leveraging adaptive bitrate algorithms (ABRs), video players dynamically select each segment’s appropriate representation (bitrate and resolution), targeting crisp, seamless playback. By constantly monitoring network metrics such as available bandwidth and latency, ABRs autonomously switch between representations, striking a balance between video quality and rebuffering [5]. Evaluating the performance of a video streaming pipeline (codec, ABR, etc.) requires monitoring many video playbacks in many combinations of network conditions, players, devices, bitrate ladder, etc. [6]. This comes at a potentially prohibitive cost unless the testbed used is efficient in terms of cost and accuracy. Such a testbed must emulate multiple clients, each with a different player and a dedicated network trace approaching a real streaming scenario. In this way, monitoring streaming components’ behavior can be parallelized, and performance can be evaluated efficiently. For example, CAdViSE [7] is a cloud-based adaptive video streaming evaluation framework for the automated testing of adaptive media players (see Figure 1). To run an experiment in CAdViSE with 𝑛 number of players, we need to instantiate 𝑛 EC2 machines and load a video player script (e.g., dashjs [8], shaka [9]) in each. Therefore, the evaluation cost of CAdViSE is proportional to the number of EC2 instances used. Moreover, in this context, adhering to realistic network traces is crucial. For example, using an inefficient network traffic shaper may not generalize well for an ABR in the real world, leading to sub-optimal performance and low user satisfaction.
MMSys ’24, April 15–18, 2024, Bari, Italy
Origin Server Dockerized Player Player
Amazon S3
Origin/CDN Server
Internal Network
VRP PyShaper
Wondershaper
Physical Machine
Figure 1: CAdViSE architecture
Wondershaper [10] is a network shaper used not only in CAdViSE but also in many other emulators that need to control network speed [11, 12]. Wondershaper shapes the network bandwidth by connecting directly to a computer’s network interface card (NIC) and bottlenecking it to the required bandwidth using the Linux built-in command, i.e., the TC command [13]. In this paper, we address the cost and accuracy of video streaming emulators and introduce PyStream, a simple yet powerful emulator tool that initiates and monitors multiple players on a physical machine or a virtual machine (e.g., AWS EC2 [14]). PyStream is a docker-based emulator that, as Figure 2 shows, allows developers and researchers to run different numbers of players with various player types and network traces. Section 2 introduces the current state-of-the-art landscape regarding network shapers and video streaming monitoring tools and emulators. Section 3 describes the current problems with network emulation and motivates our work. Section 4 delves into the specific design and implementation of PyStream, while Section 5 provides benchmarks and observations about PyStream’s performance. Finally, Section 6 concludes our work with use cases and conclusions.
Related Work
Regarding the defined objectives for designing PyStream, we have identified several tools that handle one or more steps in the traffic emulation pipeline. Wondershaper [10] is a foundational network shaping tool that was released in 2002 and is still relevant nowadays. Conceptually simple, it is a script that allows the user to limit the bandwidth of one or more network adapters. It uses Linux’s iproute TC command [13] but greatly simplifies its operation. Mahimahi [15] is another network emulation and traffic shaping framework that works under the record-replay principle. HTTP traffic can be recorded, stored, and replayed with different network conditions. For this matter, Mahimahi provides several tools, such as the RecordShell and ReplayShell for recording and replaying traffic, respectively, the DelayShell for emulating delays and the LinkShell to emulate packet loss. Trickle [16] is a tool for rate limiting TCP connections
Network Traces
Player Scripts
Local Database
System Manager
System Manager
2
Coordinator
Player
Radler, Prüller, et al.
OR
Virtual Machine (AWS EC2)
Figure 2: Proposed PyStream architecture in smaller, unmanaged networks that runs in user-space. It utilizes the preloading functionality of the Unix dynamic loader to load its own socket library wrappers. It enables Trickle to delay and truncate socket I/Os without requiring administrator privileges; and, in this way it limits the bandwidth. Dummynet [17] is a network emulation tool widely used in the field of networking research. It includes a traffic shaper and a packet scheduler, enables the emulation of a whole set of network environments, and it has the ability to model delay, and packet loss, it is easy to integrate and provides consistent results. With Mininet [18], users can create, configure, and connect virtual network nodes using lightweight containerization technologies. This enables the emulation of complex network topologies on a single computer, in accordance with Mininet’s focus on SDN research, education, and testing. ns-3 [19] is a discrete-event network simulator targeted at research and education, allowing the modeling and simulation of computer networks with a focus on network protocols and scenarios. Sabre [20] is an open-source framework that emulates real-world conditions to evaluate ABRs via quality of experience (QoE) metrics. Sabre can also be used to test new ABR algorithms without learning the intricacies of implementing the production layer. AdViSE [21] is a framework that enables testing media players under various network conditions. It accomplishes this by creating software-defined networks that are then shaped into the desired network. This network is then hosted on a physical machine so the players can access it. CAdViSE [7] improves over AdViSE by removing the system’s physical deployment. This testbed can be operated on cloud-based systems such as Amazon AWS, making upscaling easier and standardizing the machine’s performance. This architecture is continuously updated with new features and improvements. One such improvement is LLL-CAdViSE[22], which enables the testing of low-latency metrics. Table 1 presents a quick overview of video streaming emulators and network traffic shaper tools.
3
Problem Description and Motivation
As mentioned earlier, PyStream aims to tackle two main challenges in video streaming emulators: the cost of emulation/evaluation and the need for accuracy. In Figure 1, we illustrated that as the
PyStream: Enhancing Video Streaming Evaluation
MMSys ’24, April 15–18, 2024, Bari, Italy
Table 1: Overview of related work over different metrics. (TS: traffic shaper, VSE: video streaming emulator, NE: network emulator, SE: switch emulator, and NS: network simulator)
Wondershaper [10] Trickle [16] Mahimahi [15] Dummynet [17] Mininet [18] ns-3 [19] Sabre [20] AdViSE [21] CAdViSE [7] PyStream
Tool
TS Type
Platform
TS TS NE, TS NE, TS SE NS VSE VSE VSE VSE
iproute TC iproute TC Wondershaper Wondershaper Wondershaper PyShaper
Linux Unix-like Linux/Windows Unix-like Linux Unix-like Linux Linux Linux Linux/Windows
number of players increases in CAdViSE, the cost also escalates. This increase occurs because we have to assign an EC2 virtual machine to each player, pulling its image from the Docker hub. The limitation arises from Wondershaper [10], the traffic shaper used in CAdViSE, which can only run on a physical or virtual network interface. Note that Wondershaper has been used as the main traffic shaper module in many emulators (see Table 1). From an accuracy standpoint, we compare the original network trace with the predicted bandwidth across various network traces1 [7] and two ABR algorithms: L2A [23] and throughput based [24]. As depicted in Figure 3, the predicted bandwidth by both ABR algorithms is not accurate. This behavior could be attributed to either the inaccuracy of the ABR bandwidth prediction module or the low performance of Wondershaper. While numerous studies have focused on optimizing the bandwidth prediction module in ABRs [25], the primary reason for this behavior is the low accuracy in shaping network traffic with Wondershaper. The evaluation section will demonstrate that ABRs can accurately predict the bandwidth if paired with an efficient traffic shaper. Therefore, in this paper, we introduce PyStream, a Docker-based video streaming emulator which includes an efficient Python-based network traffic shaper called PyShaper. Our system is developed not only to enhance the accuracy of traffic shaping but also to reduce the cost of streaming emulation.
4
PyStream Details
Figure 2 illustrates the conceptual architecture of PyStream, which can be executed on either a physical or virtual (e.g., AWS EC2) machine. PyStream operates through Docker and consists of the following modules: coordinator, virtual reverse proxy (VRP), PyShaper, and the local database. Before describing the details of these modules, let us explain how PyStream can be configured to run various video streaming scenarios. 1 The Ramp Up network trace cycles through the values {372, 990, 1608, 2226, 2844,
3462, 4080, 4698, 5316, 5934, 6552, 7170} (kbps) after every 10 seconds, with 80 ms of constant latency. The Ramp Down network trace cycles through the same values as Ramp Up, in inverse order after every 10 seconds, with 80 ms of constant latency. The Cascade network trace cycles through the values {1, 2, 3, 4, 5, 6, 5, 4, 3, 2, 1} (Mbps) after every 5 seconds, with 80 ms of constant latency. The Steps network trace cycles through the values {990, 1608, 2226, 4080, 4698, 6552, 4698, 4080, 2226, 1608, 990} (kbps), after every 10 seconds, with 80 ms of constant latency.
The system manager, orchestrating the experiments, first defines the quartet <[PID], PT, NT, L>, where [PID] is the list of player IDs that should run the player type PT (e.g., dashjs [8], shaka [9], etc.) operating over the network trace NT, that is similar to the proposed JSON format in CAdViSE. Moreover, 𝐿 is the duration of the emulation in seconds. We note here that PID can be a unique name for each player. When a player sends an HTTP request, its PID is embedded in the URL. Thus, we can easily customize the manifest address for each player by adding the PID. The quartets are defined and stored in the local database by the system manager. The main responsibility of the coordinator module is to fetch a docker from Docker Hub, update the determined player script (𝑃𝑇 ), and run the docker. The coordinator also instructs PyShaper to apply the assigned network traces (𝑁𝑇 ) to the players’ traffic. Dockerized players are configured to connect to the VRP module to start streaming. The VRP plays as a man-in-the-middle between the players and the origin/CDN server. In the VRP, PyShaper adjusts the bandwidth for each player by introducing synthetic delay when sending HTTP packets, according to the determined network traces. Whenever a player sends an HTTP request, the VRP downloads the requested segment from the origin/CDN server2 . The VRP then calculates the synthetic time the segment should have taken to download, according to the player’s network trace and waits that long before forwarding the response to the player. To accurately follow the given network traces, a separate CPU thread updates the variables used for bandwidth calculations for each player every second. Crucially, before setting the thread to sleep for the calculated delay, part of the segment must be sent to ensure that the TCP connection remains active and an accurate bandwidth estimation at the player is maintained. Once the thread’s sleep period concludes, the remaining segment content is transmitted to the player (see Figure 4). Algorithm 1 shows the logic for calculating the synthetic delay. To that end, it uses the following parameters: Input. • Player ID (PID): This uniquely identifies the player that the algorithm is calling. • Segment Size (SS): Represents the size of the currently requested segment by the player PID. • Network Trace (NT): The set of network trace values used for the player PID. • Streaming time (ST): A pointer that tracks the current position inside the set NT for the player PID. Global Variables. Leftover Bandwidth (LeftOverBW) [global variables]: A two-dimensional array that, with the length of the number of players, stores unused bandwidth from their previously requested segments. Output. Synthetic Delay (d)[output]: The determined delay by the algorithm for player PID. Let us demonstrate Algorithm 1 through a simple example. Consider a dockerized player with PID=1 requesting a segment with a size of SS=25 kb (see Figure 4). Additionally, we have a simple network trace NT which consists of three values NT={10, 20, 30} kbps, 2 In this study, we assume that all requested segments are available in the local database.
MMSys ’24, April 15–18, 2024, Bari, Italy
2 0
0
20
40
60
80
streaming time (s)
100
120
6 4 2 0
0
20
40
60
80
streaming time (s)
100
120
Predicted Bandwidth- ABR:Throughput Net. Trace: Cascade 8
6
Bitrate (Mbps)
4
Bitrate (Mbps)
6
Bitrate (Mbps)
Bitrate (Mbps)
Original Net. Trace Predicted Bandwidth- ABR:L2A Net. Trace: Ramp Down 8 8
Net. Trace: Ramp Up
8
Radler, Prüller, et al.
4 2 0
0
20
40
60
80
streaming time (s)
100
120
Net. Trace: Steps
6 4 2 0
0
20
40
60
80
streaming time (s)
100
120
Figure 3: Comparing the original network traces (Ramp Up, Ramp Down, Cascade, and Steps) [7] with the predicted bandwidth by different ABRs: L2A [23] and throughput based [24]. representing the available bandwidth over three seconds of streaming. It’s important to note that if the values of NT are set for a specific duration longer than a second, we should scale NT down to a one-second interval by repeating its values according to its duration. We also set LeftOverBW[1]=0. To calculate the correct delay for the requested segment with size SS, the algorithm first calculates the required bandwidth (RBW), which is 40 kb/1 sec (line 1). If bandwidth from an earlier call is available, it is stored in an available bandwidth variable, denoted by ABW. In the first call, we have ABW = 0. If enough bandwidth from the last call is available (the ’if’ condition in line 3), then the synthetic delay d is calculated in line 6. However, if ABW < RBW, the ’else’ path is executed. In the first three lines of the ’else’ block, we initialize delay d using the remaining ABW from the last function call, reduce SS by ABW × 1 sec, and set RBW=0, in lines 8-10, respectively. The ’while’ loop in line 11 is entered because RBW > ABW. ST is increased by one since we need to look for the next entry of the NT array and increase ABW accordingly. Thus, we have ABW = NT[ST] = 10 kbps. With the current state of the values, the ’if’ clause in line 15 is false, and the else block must be executed. Therefore, SS is decreased by ABW = 10 × 1 sec, which means SS is now 15, and d is increased by 1. The next execution of the while loop in Line 11 will once again increase ST by 1, set ABW to the next network trace entry NT[ST], and then recalculate RBW. Now, the ’if’ clause in line 15 is true, and then d is set to d = d + RBW / ABW = 1.75. It means that the segment transmission must be delayed by 1.75 seconds to emulate the selected network trace. The only thing left is to reduce the ABW by the RBW that was just used. This value is then saved into
Dockerized Player
VRP PyShaper
It receives the first chunk of the segment immediately. Wait for d seconds to receive the remaining data.
http://… ?CID=1
NT=[10, 20, 30] kbps SS=25kb
Synthetic delay d=1.75s
Figure 4: An example of requesting a segment by a dockerized player in PyStream.
LeftOverBW[1] followed by setting RBW to 0. Since RBW is now 0, the while loop terminates in the next iteration, and d is returned in line 23.
5
Performance
In order to assess PyStream’s performance, we setup an AWS EC2 instance of type c6a.32xlarge, which is equipped with the following components: • CPU: 64 cores @ 3.6 GHz with hyperthreading, for a total of 128 threads • RAM: 256 GB Algorithm 1: Synthetic Delay Calculator Function Global : LeftOverBW Input : 𝑃𝐼𝐷 (player ID), 𝑆𝑆 (segment size), 𝑁𝑇 (network trace), 𝑆𝑇 (streaming time) Output :𝑑 (synthetic delay) 1 𝑅𝐵𝑊 ← 𝑆𝑆 / 1 sec 2 𝐴𝐵𝑊 ← LeftOverBW[𝑃𝐼𝐷] 3 if 𝐴𝐵𝑊 ≥ 𝑅𝐵𝑊 then 4 𝐴𝐵𝑊 ← 𝐴𝐵𝑊 − 𝑅𝐵𝑊 5 LeftOverBW[𝑃𝐼𝐷] ← 𝐴𝐵𝑊 6 𝑑 ← 𝑆𝑆 / 𝑁𝑇 [𝑆𝑇 ] 7 else 8 𝑑 ← 𝐴𝐵𝑊 / 𝑁𝑇 [𝑆𝑇 ] 9 𝑆𝑆 ← 𝑆𝑆 − 𝐴𝐵𝑊 × 1 sec 10 𝐴𝐵𝑊 ← 0 11 while 𝑅𝐵𝑊 > 𝐴𝐵𝑊 do 12 𝑆𝑇 ← 𝑆𝑇 + 1 13 𝐴𝐵𝑊 ← 𝑁𝑇 [𝑆𝑇 ] 14 𝑅𝐵𝑊 ← 𝑆𝑆 / 1 sec 15 if 𝐴𝐵𝑊 ≥ 𝑅𝐵𝑊 then 16 𝑑 ← 𝑑 + 𝑅𝐵𝑊 / 𝐴𝐵𝑊 17 𝐴𝐵𝑊 ← 𝐴𝐵𝑊 − 𝑅𝐵𝑊 18 𝑅𝐵𝑊 ← 0 19 LeftOverBW[𝑃𝐼𝐷] ← 𝐴𝐵𝑊 20 else 21 𝑆𝑆 ← 𝑆𝑆 − 𝐴𝐵𝑊 × 1 sec 22 𝑑 ←𝑑 +1 23 return 𝑑
PyStream: Enhancing Video Streaming Evaluation
8
6 4
8
6 4
4
2
2
2
0
0
0
0
0
20
40
60
80 100 120
Streaming time (s)
0
20
40
60
80 100 120
Streaming time (s)
0
20
40
60
80 100 120
8
6
2
Streaming time (s)
Net. Trace: Ramp Up
10
Bitrate (Mbps)
4
8
Net. Trace: Steps
10
Bitrate (Mbps)
6
Net. Trace: Cascade
10
Bitrate (Mbps)
Bitrate (Mbps)
8
Net. Trace: Ramp Down
10
Bitrate (Mbps)
Net. Trace: Stable Original Net. Trace PyStream(PyShaper) TC
10
MMSys ’24, April 15–18, 2024, Bari, Italy
6 4 2
0
20
40
60
0
80 100 120
Streaming time (s)
0
20
40
60
80 100 120
Streaming time (s)
Figure 5: Comparing the variation between the original network trace and predicted bandwidth by L2A ABR using PyStream and Wondershaper to shape the network traffic across various network traces.
8
6 4
8
6 4
4
2
2
2
0
0
0
0
0
20
40
60
80 100 120
Streaming time (s)
0
20
40
60
80 100 120
Streaming time (s)
0
20
40
60
80 100 120
8
6
2
Streaming time (s)
Net. Trace: Ramp Up
10
Bitrate (Mbps)
4
8
Net. Trace: Steps
10
Bitrate (Mbps)
6
Net. Trace: Cascade
10
Bitrate (Mbps)
Bitrate (Mbps)
8
Net. Trace: Ramp Down
10
Bitrate (Mbps)
Net. Trace: Stable Original Net. Trace PyStream(PyShaper) TC
10
6 4 2
0
20
40
60
0
80 100 120
Streaming time (s)
0
20
40
60
80 100 120
Streaming time (s)
Figure 6: Comparing the variation between the original network trace and predicted bandwidth by Throughput ABR using PyStream and Wondershaper to shape the network traffic across various network traces.
8
6 4
8
6 4
4
2
2
2
0
0
0
0
0
20
40
60
80 100 120
Streaming time (s)
0
20
40
60
80 100 120
Streaming time (s)
0
20
40
60
80 100 120
Streaming time (s)
8
6
2
Net. Trace: Ramp Up
10
Bitrate (Mbps)
4
8
Net. Trace: Steps
10
Bitrate (Mbps)
6
Net. Trace: Cascade
10
Bitrate (Mbps)
Bitrate (Mbps)
8
Net. Trace: Ramp Down
10
Bitrate (Mbps)
Net. Trace: Stable Original Net. Trace PyStream(PyShaper) TC
10
6 4 2
0
20
40
60
80 100 120
Streaming time (s)
0
0
20
40
60
80 100 120
Streaming time (s)
Figure 7: Comparing the variation between the original network trace and predicted bandwidth by Throughput ABR using PyStream and Linux TC command to shape the network traffic across various network traces. • OS: Ubuntu 20.04 LTS The machine’s performance is monitored with the built-in Docker Stats [26], and a custom script logged each player’s load at every second. To benchmark these results, the same tests are run on CAdViSE, in order to emulate the same traces used in PyStream.
5.1
Accuracy Tests
The accuracy tests are performed in order to see how close the players’ predictions are to the real network traces. For this, 20 players run for 120 seconds, and the average bandwidth predictions of L2a and Throughput-based ABRs are plotted in Figures 5 and 6, respectively. Around the average value, the range between minimum and maximum player value is shown as an indication of variability. It highlights how predicated bandwidth by ABRs is closer and much more consistent among each other to the original network trace when we use Pyshaper compared to the experiments using wondershaper. To have a better evaluation on the performance of PyStream, in the next scenario, we use Linux TC command [13] to shape the
bandwidth for players. As shown in Figure 7, PyStream efficiently emulates the bandwidth of the given network trace in comparison with TC tool.
5.2
Load Tests
Having shown the PyStream’s accuracy on different ABRs and network traces, PyStream is then tested on its scalability. Specifically, how increasing the number of players would impact the accuracy. To this end, the experiment is repeated with 8, 16, 32, 64, and 128 players. The number of players was chosen due to the EC2 machine’s resources to check the response of PyStream when the number of threads needed for players became equal (64) and higher (128) than the physical number of cores (64). Figure 8 shows the results of this stress test, where the red dot corresponds to the mean absolute error (MAE), calculated as follows: Í𝑛 |𝑇𝑖 − 𝑡𝑖 | 𝑀𝐴𝐸 = 𝑖=0 𝑛 , where 𝑇𝑖 is the real bitrate for client i, 𝑡𝑖 is the bitrate predicted by the player for client i, and 𝑛 is the number of players. The
100 10 1 8
16
32
64
Number of Concurrent Players
128
103 102 101 100
8
16
32
64
Number of Concurrent Players
128
Net. Trace: Cascade
103
102
101 8
16
32
64
Number of Concurrent Players
128
Net. Trace: Steps
104 103 102 101 100 8
16
32
64
Number of Concurrent Players
128
Deviation from Real Bitrate (log. kbps)
101
Net. Trace: Ramp Down
104
Deviation from Real Bitrate (log. kbps)
102
Radler, Prüller, et al.
Deviation from Real Bitrate (log. kbps)
Net. Trace: Stable 103
Deviation from Real Bitrate (log. kbps)
Deviation from Real Bitrate (log. kbps)
MMSys ’24, April 15–18, 2024, Bari, Italy
Net. Trace: Ramp Up
104 103 102 101 100 10 1
8
16
32
64
Number of Concurrent Players
128
Figure 8: MAE and Variance in relation to concurrent players
CPU Usage [%]
1.0
Players Coordinator
0.8 0.6 0.4 0.2 0.0
8
RAM Usage [%]
0.20 0.15
16
32
64
128
16
32
64
128
Number of Concurrent Players Players Coordinator
0.10 0.05 0.00
8
Number of Concurrent Players
Figure 9: Resource usage in relation to concurrent players results confirm that, while the number of players is much lower than the number of physical cores, PyStream-caused deviations from the network traces remain small and increase exponentially (linearly in our log plot) up to the physical number of cores. Once the number of processes surpasses the number of cores, it appears that the performance stabilizes, but this is due to the machine being forced to freeze several threads. Therefore, performances above this threshold should not be considered.
5.3
Resouce Tests
In the final test, we investigate the connection between the machine’s resources and PyStream’s utilization of said resources. This is extrapolated from running the same experiments as above and logging each Docker container’s (𝑛 players + 1 controller) resource utilization over time. Figure 9 shows a linear dependence of both CPU and RAM usage to the number of players, but not from the same containers: while CPU usage is negligible for player containers but high for the controller, the opposite is true for RAM. This is because the controller is fundamentally a Python script that needs
to run calculations for each player, while each player does not perform complex calculations but has to store a player and, crucially, its streamed video segments. This is a fundamental point in favor of PyStream’s architecture over CAdViSE: typical virtual machines’ CPU count and RAM scale together. Therefore, a system that is able to leverage both from the same machine is a more efficient use of resources than needing to spin up multiple machines of which only one dimension is fully utilized. As can be seen from Figure 9, the CPU count is the bottleneck of this system, while the required RAM is much lower than the physical limit of an EC2 machine with a sufficient number of cores. Therefore, both PyStream and CAdViSE need the same-sized machine to perform an experiment with the same number of players. Quantifying the resource-saving is then a simple matter of summing the costs of EC2 machines required for CAdViSE’s players. Figure 9 shows how, for 64 players, the total utilization of RAM is around 7.5% of the total 256 GB available, resulting in 0.3 GB of RAM per player. This amount of RAM is serviced already by the smallest EC2 instance type, namely t2.nano, which costs 0.6 cents/hour, for a total cost of 38 cents/hour. Given that a 64-core EC2 instance type such as costs about 5 euros/hour, the savings in this experiment amount to around 8%. However, this calculation assumes that the t2.nano instance type is actually able to run a CAdViSE player instance and that CAdViSE’s server requires similar computational power than PyStream’s controller. Therefore, more experiments are necessary to quantify the cost comparisons between the two systems, and possibly others, with precision.
6
Conclusion
In this study, we presented PyStream, a novel way to emulate network traces through a Docker environment, making it possible to run multiple simultaneous tests locally. We showed the limitations of current approaches, especially with regard to traffic shaping, and how PyStream significantly improves the accuracy of the emulations by developing PyShaper, an algorithm that calculates how long each segment should take to reach each player given their assigned network trace, and keeps the player from receiving it beforehand. Furthermore, by combining the resource requirements of the different parts of a network emulator, PyStream is able to make full use of a single virtual/physical machine, reducing the cost of running experiments. While further exploration and more comprehensive tests could provide improvements, we showed how PyStream’s approach is a valuable addition to the current landscape of network emulation.
PyStream: Enhancing Video Streaming Evaluation
References [1]
[2]
[3]
[4]
[5]
[6]
[7]
[8] [9] [10] [11]
[12]
[13] [14]
2023. 2023 global internet phenomena report. Last accessed 30 November 2023. (2023). https://www.sandvine.com/hubfs/Sandvine_Redesign_2019/Downloa ds/2023/reports/Sandvine%20GIPR%202023.pdf. 2021. Amount of data created, consumed, and stored 2010-2020, with forecasts to 2025. Last accessed 30 November 2023. (2021). https://www.statista.com/sta tistics/871513/worldwide-data-created/. Farzad Tashtarian, Abdelhak Bentaleb, Alireza Erfanian, Hermann Hellwagner, Christian Timmerer, and Roger Zimmermann. 2022. Hxl3: optimized delivery architecture for http low-latency live streaming. IEEE Transactions on Multimedia. Farzad Tashtarian, Alireza Erfanian, and Amir Varasteh. 2018. S2VC: An SDNbased framework for maximizing QoE in SVC-based HTTP adaptive streaming. Computer Networks, 146, 33–46. Minh Nguyen, Daniele Lorenzi, Farzad Tashtarian, Hermann Hellwagner, and Christian Timmerer. 2022. Dofp+: an http/3-based adaptive bitrate approach using retransmission techniques. IEEE Access, 10, 109565–109579. Farzad Tashtarian, Abdelhak Bentaleb, Hadi Amirpour, Sergey Gorinsky, Junchen Jiang, Hermann Hellwagner, Christian Timmerer, et al. 2024. ARTEMIS: Adaptive bitrate ladder optimization for live video streaming. In USENIX Symposium on Networked Systems Design and Implementation, 1–21. Babak Taraghi, Anatoliy Zabrovskiy, Christian Timmerer, and Hermann Hellwagner. 2020. Cadvise: cloud-based adaptive video streaming evaluation framework for the automated testing of media players. In Proceedings of the 11th ACM Multimedia Systems Conference (MMSys ’20). Association for Computing Machinery, Istanbul, Turkey, 349–352. isbn: 9781450368452. doi:10.1145/33398 25.3393581. DASH Industry Forum (DASH-IF). 2012. Dash.js javascript reference client. Last accessed 30 November 2023. (2012). https://reference.dashif.org/dash.js/. Shaka Project. 2015. Shaka player. Last accessed 30 November 2023. (2015). https://github.com/shaka-project/shaka-player. Simon Séhier Bert Hubert Jacco Geul. 2002. Wondershaper. Last accessed 30 November 2023. (2002). https://github.com/magnific0/wondershaper. Hua-Jun Hong. 2017. From cloud computing to fog computing: unleash the power of edge and end devices. In 2017 IEEE International Conference on Cloud Computing Technology and Science (CloudCom), 331–334. doi:10.1109/CloudCo m.2017.53. Yacine Benchaib and Claude Chaudet. 2012. Virmanel: a mobile multihop network virtualization tool. In Proceedings of the Seventh ACM International Workshop on Wireless Network Testbeds, Experimental Evaluation and Characterization (WiNTECH ’12). Association for Computing Machinery, Istanbul, Turkey, 67–74. isbn: 9781450315272. doi:10.1145/2348688.2348703. Alexey N. Kuznetsov Bert Hubert. 2024. Tc(8). Last accessed 29 January 2024. (2024). https://linux.die.net/man/8/tc. Amazon. 2006. Amazon elastic compute cloud (ec2): secure and resizable compute capacity for virtually any workload. Last accessed 30 November 2023. (2006). https://aws.amazon.com/ec2/.
MMSys ’24, April 15–18, 2024, Bari, Italy
[15]
[16] [17]
[18]
[19]
[20]
[21]
[22]
[23]
[24]
[25]
[26]
Ravi Netravali, Anirudh Sivaraman, Keith Winstein, Somak Das, Ameesh Goyal, and Hari Balakrishnan. 2014. Mahimahi: a lightweight toolkit for reproducible web measurement. In Proceedings of the 2014 ACM Conference on SIGCOMM (SIGCOMM ’14). Association for Computing Machinery, Chicago, Illinois, USA, 129–130. isbn: 9781450328364. doi:10.1145/2619239.2631455. Marius A Eriksen. 2005. Trickle: a userland bandwidth shaper for unix-like systems. In USENIX Annual Technical Conference, FREENIX Track, 61–70. Luigi Rizzo. 1997. Dummynet: a simple approach to the evaluation of network protocols. SIGCOMM Comput. Commun. Rev., 27, 1, (Jan. 1997), 31–41. doi:10.1 145/251007.251012. Bob Lantz, Brandon Heller, and Nick McKeown. 2010. A network in a laptop: rapid prototyping for software-defined networks. In Proceedings of the 9th ACM SIGCOMM Workshop on Hot Topics in Networks (Hotnets-IX) Article 19. Association for Computing Machinery, Monterey, California, 6 pages. isbn: 9781450304092. doi:10.1145/1868447.1868466. Klaus Wehrle, Mesut Güneş, and James Gross, (Eds.) 2010. The ns-3 network simulator. Modeling and Tools for Network Simulation. Springer Berlin Heidelberg, Berlin, Heidelberg, 15–34. isbn: 978-3-642-12331-3. doi:10.1007/978-3-64 2-12331-3_2. Kevin Spiteri, Ramesh Sitaraman, and Daniel Sparacio. 2018. From theory to practice: improving bitrate adaptation in the dash reference player. In Proceedings of the 9th ACM Multimedia Systems Conference (MMSys ’18). Association for Computing Machinery, Amsterdam, Netherlands, 123–137. isbn: 9781450351928. doi:10.1145/3204949.3204953. Anatoliy Zabrovskiy, Evgeny Kuzmin, Evgeny Petrov, Christian Timmerer, and Christopher Mueller. 2017. Advise: adaptive video streaming evaluation framework for the automated testing of media players. In Proceedings of the 8th ACM on Multimedia Systems Conference (MMSys’17). Association for Computing Machinery, Taipei, Taiwan, 217–220. isbn: 9781450350020. doi:10.1145 /3083187.3083221. Babak Taraghi, Hermann Hellwagner, and Christian Timmerer. 2023. Lll-cadvise: live low-latency cloud-based adaptive video streaming evaluation framework. IEEE Access, 11, 25723–25734. doi:10.1109/ACCESS.2023.3257099. Theo Karagkioules, Rufael Mekuria, Dirk Griffioen, and Arjen Wagenaar. 2020. Online learning for low-latency adaptive streaming. In Proceedings of the 11th ACM Multimedia Systems Conference (MMSys ’20). Association for Computing Machinery, Istanbul, Turkey, 315–320. isbn: 9781450368452. doi:10.1145/33398 25.3397042. Abdelhak Bentaleb, Zhengdao Zhan, Farzad Tashtarian, May Lim, Saad Harous, Christian Timmerer, Hermann Hellwagner, and Roger Zimmermann. 2022. Low latency live streaming implementation in dash and hls. In Proceedings of the 30th ACM International Conference on Multimedia, 7343–7346. Abdelhak Bentaleb, Mehmet N Akcay, May Lim, Ali C Begen, and Roger Zimmermann. 2022. Bob: bandwidth prediction for real-time communications using heuristic and reinforcement learning. IEEE Transactions on Multimedia. Docker Inc. 2013. Docker container stats. Last accessed 30 November 2023. (2013). https://docs.docker.com/engine/reference/commandline/container_st ats/.