ASTRA: Low-Overhead Runtime Architecture for STReam Adaptation in Video Analytics Mahshid Ghasemi, Zoran Kostic, Javad Ghaderi, and Gil Zussman
arXiv:2609.07020v1 [cs.NI] 7 Sep 2026
Electrical Engineering, Columbia University {mahshid.ghasemi, zk2172, jghaderi, gil.zussman}@columbia.edu
Abstract—Real-time video analytics is crucial for smart city applications and cloud-connected vehicle control. To improve analytics accuracy, it is desirable to process the video at the highest resolution and frame rate. However, due to limited resources, streaming and processing video at the highest resolution and frame rate from all cameras is not feasible and adversely affects the analytics latency. Intelligent adaptation of cameras’ resolutions and frame rates based on network conditions and the video content is crucial in order to optimize the performance. In this paper, we present ASTRA, a lowoverhead runtime architecture for online adaptation of live camera analytics at the edge. ASTRA can execute various online algorithms as a black box. We deployed ASTRA in the realistic NSF COSMOS testbed and uniquely assessed its realtime performance using COSMOS’ street-level cameras. We further evaluated ASTRA with up to eight emulated cameras by streaming a comprehensive video dataset under real-world network conditions. We used ASTRA’s architecture to evaluate the practical performance of several classes of adaptation algorithms, including theoretical and empirical methods. The results indicate that ASTRA can provide system reliability (i.e., the probability of meeting accuracy and latency requirements) of more than 90% while maintaining performance within a deviation of less than 10% from optimal offline performance with average GPU utilization overhead of around 2% per camera. Index Terms—Video analytics, edge computing, online optimization, testbed evaluation
I. Introduction Video analytics, powered by Deep Neural Networks (DNNs), has gained importance in a wide range of applications, including smart city infrastructure, traffic signal optimization, and cloud-connected vehicle control [1], [2], [3]. The recent proliferation of traffic cameras, coupled with the availability of advanced edge/cloud computing resources, is expected to facilitate large-scale video analytics. Consider a system of geo-distributed cameras streaming to edge/cloud servers for real-time object detection. These cameras continuously generate a large volume of data that must be encoded, transmitted, decoded, and processed in real-time while adhering to network and computational capacity constraints. A microcosm of such a system is deployed in the NSF COSMOS wireless edge-cloud testbed in New York City (NYC) (see Fig. 1) [4], [2]. As part of this testbed, several traffic cameras are installed at busy streets and intersections in Manhattan, and their real-time video streams are processed at COSMOS’ edge servers
Edge servers
Edge servers
Edge servers
zzzzzz
zzzzzz
zzzzzz
Highperformance cloud servers
Figure 1: The NSF COSMOS testbed uses geo-distributed cameras and edge servers. It allows emulating real-world scenarios that require resource allocation for analyzing real-time video streams [5], [2], [4]. to extract traffic- and crowd-related information such as pedestrian and vehicle density. Two key performance metrics in such systems are accuracy and (end-to-end) latency. Accuracy is determined by evaluating how closely the output of the object detection (i.e., bounding boxes and labels) aligns with manually annotated ground-truth data. Latency comprises encoding latency, network latency, decoding latency and DNN inference latency. Accuracy and latency are significantly impacted by network conditions, computational capacity, and video content. Network conditions. Network variations can cause lower throughput, increased packet loss, and buffering. Therefore, if in such cases, resolution or frame rate are not reduced, (i) network latency increases, and/or (ii) quality of decoded images degrades which reduces accuracy. Computational capacity. Real-time processing of video streams using DNNs is computationally demanding. For example, the execution of a conventional object detection model on a single 1K resolution video stream at 30 fps can consume most of the processing capacity of an edge GPU. In such a case, processing additional streams leads to increased inference latency unless the resolution or frame rate of at least one of the video streams or the DNN model size is reduced. Video content. Depending on the density level and movements in the monitored scene, different resolutions, frame rates, and DNN complexities are required in order to achieve sufficient accuracy. For example, for cameras viewing urban streets, when vehicles are moving slowly (e.g., at a red traffic light or during heavy traffic) the frame rate can be reduced without losing accuracy. Similarly, when the pedestrians’ density is low, or they are close to the camera, the resolution can be reduced without
Figure 2: Deployment of online adaptation in a multi-camera setup connected to edge servers. The adaptation system updates the configuration of the system periodically in response to changes in the environment. affecting the accuracy. Therefore, optimizing parameters such as resolution and frame rate when feasible, can significantly save on bandwidth and computational power. Iterative adaptation algorithm. In this paper, the objective is to dynamically identify video streams’ resolutions and frame rates that maximize the overall performance (a function of accuracy and end-to-end latency), while meeting application-specific constraints on minimum accuracy and maximum end-to-end latency. We refer to a specific combination of resolution and frame rate for a stream as a configuration. Our measurements, consistent with prior research [6], [7], [8], [9], indicate that accuracy and latency are timevarying and unknown functions of the configuration. The characteristics of these functions depend on dynamic variables such as video content, available network and computing resources, and DNN models. Identifying an optimal configuration through deterministic or offline optimization with empirical modeling cannot capture the unknown and time-varying nature of the problem and often leads to poor performance. This necessitates iterative online adaptation in such unpredictable settings, where configurations can be evaluated only when deployed and processing real video streams, without knowledge of the future or an explicit model for the accuracy or latency functions. At each iteration, such algorithms select a configuration to evaluate based on previously observed performance. We refer to such algorithms as iterative algorithms. Key idea. The primary goal of online adaptation is to maximize performance under constrained resources. Achieving this requires that: (i) the adaptation runtime architecture incurs minimal overhead when executing the iterative algorithm, and (ii) given the rapid evolution of online adaptation algorithms, it is crucial that the architecture allows for easy integration of new algorithms, requiring only minimal changes. To meet these requirements, we introduce ASTRA, a low-overhead runtime Architecture for STReam Adaptation in real-time video analytics that enables execution of existing and future iterative algorithms as a black-box (as shown in Fig. 2). A patent based on ASTRA is pending [10]. Prior studies. Significant work has been done on exploring
online video analytics adaptation approaches. However, most previous studies focus solely on developing adaptation algorithms (the “content” of the black-box in Fig. 2) and are evaluated using recorded videos, not within an end-to-end real-time system [8], [6], [11], [12], [13]. Despite the growing number of online adaptation algorithms for video analytics, there remains a significant gap in understanding the system-level overheads and practical feasibility of deploying such algorithms within an edge/cloud network and operating them in real-time. ASTRA is designed to close this gap. To our knowledge, ASTRA is the first to provide a detailed architectural design and practical implementation for executing iterative adaptation algorithms in a real-time video analytics pipeline. A. Challenges Architecture challenges. Two main obstacles in executing online adaptation algorithms are (i) the cost of performance measurements at each iteration and (ii) the delay incurred in iterative switching between configurations and video streams. (i) Performance measurement cost: Online adaptation algorithms require feedback that indicates the observed performance of a specific configuration. For example, for object detection, to measure the accuracy associated with a given configuration, the detected objects (i.e., bounding boxes and their labels) obtained with that configuration must be compared against the ground truth. Since it is impossible to access manually annotated ground truth in real-time, a proxy ground truth obtained via the most resource-intensive configuration (highest resolution and frame rate) may be used instead. Generating proxies, by definition, requires significant computational and network overhead, and so does measuring the accuracy/latency of a given configuration. (ii) Switching delay: Executing iterative adaptation requires evaluating various configurations by sequentially running object detection models on short segments of the video stream under each configuration. This necessitates switching between video streams and analytics pipelines with different configurations. Each switching introduces non-trivial overhead that can significantly increase adaptation time. This overhead includes repeated decoding and
2
conditions and fully controlled, reproducible scenarios. (i) Deployment on the COSMOS testbed. We deployed ASTRA in the COSMOS testbed network and performed real-time adaptation for two live traffic cameras (depicted in Fig. 3) deployed on the 1st and 2nd floor of a building viewing an intersection in NYC [5]. (ii) Controlled and reproducible evaluation. A fair comparison among algorithms necessitates using identical and reproducible conditions. To achieve this, we deployed ASTRA on Google Cloud virtual machines (VMs) and conducted experiments with eight emulated cameras under realistic network latency conditions. Cameras were emulated by streaming videos via the Real-time Transport Protocol (RTP) [16] using GStreamer [17]. Unlike prior studies (e.g., [18], [19], [20], [21], [22]) that rely solely on pre-recorded videos, our emulated setup preserves the essential characteristics of live deployments while also enabling reproducibility and identical conditions across algorithms. Ethical consideration. The use of COSMOS’ live and recorded video streams [5], [15] was designated IRBexempt by Columbia University. These videos are utilized solely for academic and research purposes and will be distributed only after appropriate anonymization, including obscuring faces and license plates. More generally, the authors had several discussions with local community stakeholders, which point to the fact that the use of videos for improved traffic flow and safety (a common use by municipalities) is acceptable by the community. Results. Compared to baseline architectures, ASTRA reduces the per-iteration execution overhead of the iterative adaptation algorithm by an order of magnitude (e.g., from seconds to hundreds of milliseconds). ASTRA, when integrated with the GP-UCB-C algorithm, achieves performance within 10% of the optimal performance. It ensures system reliability (i.e., probability of meeting the accuracy and latency requirements) of more than 90%. ASTRA incurs low adaptation overhead, averaging around 2% GPU utilization per camera over time.
2nd floor camera 1st floor camera
6.95 m
3.95 m
Figure 3: COSMOS cameras deployed on the 1st and 2nd floor of a building and their view of 120th St. and Amsterdam Ave. intersection, NYC [5], [15]. encoding, buffering, pipeline reconfiguration, data transfer (e.g., from storage to memory or CPU to GPU), cache inefficiencies, and model loading. In real-time systems, even minor overheads, when recurrent, accumulate and compromise performance. Therefore, to ensure practical feasibility and scalability, these overheads must be mitigated. Algorithmic challenges. Due to the temporal and computational overhead of each iteration of online adaptation, the iterative algorithm must be capable of attaining a nearoptimal configuration in a minimum number of iterations to ensure near-optimal performance most of the time. B. Contributions Identifying and measuring runtime overheads in online video adaptation. We performed extensive profiling across memory, CPU, and GPU resources, and identified hidden bottlenecks and overheads in online adaptation systems. Quantifying these overheads and bottlenecks is crucial for verifying the feasibility of deployment. Efficient architecture design to mitigate runtime overheads. We designed and implemented ASTRA, an online adaptation framework that mitigates the identified overheads. Starting from a baseline design, we systematically refined the architecture to reduce inefficiencies. ASTRA incorporates asynchronous processing, OS-level signaling, efficient memory management, and modular design. ASTRA enables performance evaluation and fair comparison of different iterative algorithms within a unified framework. Executing representative adaptation algorithms as interchangeable black-box components. To compare different algorithmic choices, we used ASTRA to execute two iterative algorithms (the “content” of the black-box in Fig. 2): (i) the Gaussian Process Upper Confidence Bound with Constraints (GP-UCB-C) algorithm derived from [14], and (ii) Chameleon++ derived from [8]. These algorithms were chosen to represent two categories of common iterative adaptation algorithms: (1) theoretically guaranteed algorithms (GP-UCB-C), and (2) empirical approximated algorithms (Chameleon++). We then compared their performance against two baselines: (a) the offline optimal (upper bound on achievable performance) and (b) a ground-truth-free adaptation method (AutoML++) that estimates accuracy via confidence scores instead of groundtruth. Realistic evaluation: We evaluated ASTRA in two complementary settings designed to capture both real-world
II. Related Work Several studies presented various methods to enhance the performance of real-time video analytics despite the dynamics in the environment, such as network conditions variation, video content (e.g., density, lighting, and weather conditions) variations. Related work primarily falls into the following categories based on their optimization approaches. Configuration adaptation. Studies employ periodic adaptation, including DNN retraining or continuous updates [12], [23], [7], [21], [22], dynamic model switching [18], [11], and adjusting resolution, frame rate, or bitrate [8], [24], [25], [26], [27], [9], [28]. Specialized systems address specific constraints such as packet loss [29], satellite links [20], or camera orientation [19]. RIVA [30] uses hierarchical reinforcement learning to adapt bitrate
3
allocation and online retraining decisions for real-time industrial video analytics under dynamic network and video conditions. OAVS [31] uses hierarchical reinforcement learning to update bandwidth prediction and bitrate allocation under dynamic network and scene conditions for drone-sourced video analytics. A benchmark provided by [32] evaluates how decoder and model-size choices affect latency in cloudlet-offloaded drone video analytics.
III. Online Configuration Adaptation Consider a video analytics system where real-time streams from several cameras are transmitted to edge servers for object detection using DNN models (see Fig. 2). The key performance metrics of such a system are accuracy, (end-to-end) latency, and bandwidth consumption. Bandwidth manifests itself in accuracy and network latency from the camera to the edge server when streaming over Transmission Control Protocol (TCP) connections. As mentioned in Section I, accuracy and latency are timevarying, unknown functions of the system’s configurations (e.g., resolution and frame rate of the cameras). The characteristics of these functions depend on dynamic parameters such as video content, model complexity, network conditions, and available computational resources. Adaptation objective. The objective is to dynamically find the configurations that maximize a value function of accuracy and latency, subject to constraints on maximum latency and minimum accuracy. The unpredictable and dynamic nature of the problem limits the effectiveness of empirical modeling and deterministic optimization in identifying near-optimal configurations. Hence, an iterative algorithm is required to estimate the system behavior in an online manner. At each iteration t of such an algorithm, one configuration xt is sampled. Here, xt ∈ X , where X is a discrete set consisting of a finite number of supported resolutions and frame rates. Then, xt ’s achieved accuracy, At (xt ), and latency, Lt (xt ) on a short segment (e.g., τ =0.2 seconds) of the camera’s video stream is evaluated. The results up to iteration t are used to decide on the subsequent configuration xt+1 to be evaluated. The algorithm must be terminated upon identification of a near-optimal configuration. We refer to such algorithms as iterative algorithms. Performance metrics. Accuracy At (xt ) is calculated by comparing the obtained object detection results (i.e., bounding boxes and labels) using configuration xt on a video segment against the ground truth. Following previous work [8], [60], [27], we use the F1-score metric. The detection F1-score is computed by verifying if a bounding box shares the same label and has sufficient spatial overlap with the associated ground truth. To account for the impact of frame rate on the accuracy, following the prior research [8], [66], [67], we use the location of objects from the previous sampled frame for a frame that is not sampled by the configuration. Latency, Lt (xt ) is defined as the sum of the encoding latency by the camera, the network latency to transfer video segment from the camera to the server, the decoding latency on the edge server, and the inference latency to run object detection on the video segment with configuration xt . Proxy ground-truth. The precise ground truth required for calculating accuracy At (xt ) can only be obtained through manual annotation, which is infeasible in realtime. Replacing manual annotation with analytical ap-
Distributed inference/learning. Studies leverage distributed computing via DNN partitioning [33], [34] and efficient edge-cloud resource allocation or load balancing [35], [36], [37], [38], [39], [40], [41], [42], [43], [13]. Concurrent DNN execution on heterogeneous processors is also explored [44], [42], [45]. VideoJam [13] designs a pipeline-level load-balancing architecture for live video analytics. ViEdge [43] distributes inference across edge resources. The study [46] benchmarks on-device versus edge-offloaded real-time vision tasks and shows that edge offload remains important for running high-accuracy models within tight latency constraints. PIB [47] extracts and compresses task-relevant feature maps on edge cameras and fuses them at the edge server. In another study [48], smart cameras offload tasks to an Unmanned Aerial Vehicle (UAV)-enabled edge server. Video frames filtering. Studies focus on on-camera frame filtering [49], [50], [51], [52], region of interest determination [53], [54], [55], [56], [57], [58], [59], and adaptive encoding [60], [61], [62]. Additionally, frame or region-based enhancement techniques are used to improve throughput and accuracy [63], [64]. RegenHance [64] enhances important regions of a frame, achieving higher throughput than naive per-frame super-resolution. JIGSAW [65] exploits cross-camera overlap to pack high-utility tiles for efficient edge inference. Unlike previous studies that primarily focus on developing adaptation methods to maintain video analytics performance under environmental variability, our work centers on the broader system challenges of executing configuration adaptation in real-time within a deployed video analytics pipeline. While prior work often models or optimizes specific costs such as bandwidth, latency, or retraining overhead, many practical execution overheads remain underexplored, including switching delays, pipeline reconfiguration, buffering behavior, resource contention, and model switching. Through extensive profiling and measurements, we show that these overheads can significantly affect adaptation performance and, if left unaddressed, can make online adaptation counterproductive by adding latency and computation that degrade overall performance. ASTRA directly targets these challenges: its architecture is explicitly designed to mitigate these realtime execution overheads, enabling online adaptation to deliver meaningful performance gains in real deployments.
4
configuration xt . This requires the system to momentarily process a short segment of video (for example, a fraction of a second such as 0.2 s) using the channel associated with configuration xt . Therefore, the system needs to switch rapidly between different channels throughout the adaptation period until the algorithm converges to a nearoptimal configuration. To understand the feasibility of such rapid switching, we measured channel-switching latency for three types of cameras deployed in the COSMOS testbed: HikVision (DS-2CD5585G0-IZHS 8 MP Outdoor Dome) camera, Bosch (FLEXIDOME 8100i NDE-8704R 4K UHD Outdoor Network PTRZ Dome) camera, and Axis (Q3628-VE 8 MP Outdoor Dome) camera. This latency is defined as the time from issuing a request for a new channel to the moment the first packet of the first frame from that channel arrives at the edge server. The edge server connects to the cameras via fiber optics, resulting in negligible network delay (under 10 ms), which means this latency primarily reflects the internal switching latency of the cameras. Each camera is programmed with three channels, with the following resolution and frame per second (fps): ((3840×2160), 30 fps), ((1920×1440), 15 fps), ((800×600), 10 fps).
Figure 4: Adaptation timeline: The iterative algorithm starts adaptation at the beginning of each time window until a near-optimal configuration is determined, then that configuration is used for the rest of the window. proximations [26] or with confidence scores generated by the DNN [9] introduces unknown errors into the accuracy measurements. Our experiments (Section VI) show that such imprecise accuracy estimates can lead the system to pick configurations that deviate significantly from the actual optimum. DNN models generally achieve their highest accuracy at the highest resolution and frame rate, which is also the most resource-intensive configuration. We refer to this as the golden configuration. Because of this, many prior works use the model’s output under the golden configuration as the reference for assessing accuracy under less resource-intensive settings [20], [6], [68]. Throughout the paper, we refer to the object detector’s output under the golden configuration as the proxy ground truth, or simply the proxy. IV. Architecture Design As discussed in Section I-A, each iteration of an iterative adaptation algorithm incurs substantial overhead. In particular, generating proxies (by definition) and sequentially switching between configurations and video streams incur network, computational, and temporal overhead. To execute such iterative algorithms efficiently, these overheads must be minimized. Therefore, an effective adaptation architecture should ensure that the system uses a nearoptimal configuration majority of the time while incurring low overhead. Due to unpredictable variations in network conditions and video content, the optimal configuration shifts over time, requiring the adaptation process to be invoked periodically. The appropriate adaptation frequency depends on environment volatility. We refer to each adaptation cycle as a time window. A typical value for the time window is 1-2 minutes. As illustrated in Fig. 4, at the beginning of each time window, adaptation resumes until a near-optimal configuration is found by the adaptation algorithm, which is then used for the remainder of the window. To keep adaptation overhead low, the adaptation period should remain small relative to the window length (e.g., below 10%). In this section, we illustrate how we designed ASTRA, accordingly starting from a baseline design.
Fig. 6 shows the measured switching latencies for these cameras. The channel switching measurements (denoted by Channel SW) are based on around 100 channel switches per camera, switching among the three available channels (e.g., 1→2, 2→3, 3→1). As Fig. 6 indicates, the channel switching latency strongly depends on the camera type. Those cameras (e.g., the Bosch camera) that employ advanced mechanisms and/or larger internal buffers to maintain high-quality and smooth real-time streaming exhibit higher switching delays. These results show that channel switching latency can range from ∼400 ms to more than 1.2 s, which is significant when the adaptation algorithm needs to switch channels every fraction of a second, such as 0.2 s. To mitigate this latency and reduce the dependency on camera-specific buffering behavior, we improve upon the baseline architecture shown in Fig. 5(a) by introducing the modified design shown in Fig. 5(b). We replace the “channel switcher” component that repeatedly requests different channels from the cameras with a down-sampler component. This component only receives the highestresolution, highest-frame-rate stream (Channel 1, indicated by green arrows in Fig. 5) and locally generates other configurations sampled by the iterative adaptation algorithm. Since Channel 1 with golden configuration, hereafter referred to as the golden stream, is already necessary to compute the proxy ground-truth, it can be mirrored internally at the edge server to perform the down-sampling, eliminating the need to request two separate streams from the camera.
A. Baseline Architectures To perform online adaptation by running an iterative algorithm described in Section III, the baseline architecture for real-time configuration adaptation is shown in Fig. 5(a). In this architecture, during an adaptation period for each iteration t, the algorithm needs to observe both the accuracy and latency obtained under the sampled
Switching latency. The effectiveness of the architecture in Fig. 5(b) is reflected in the latency measurements shown
5
(a) Channel switching: Baseline architecture for online adaptation by executing an iterative algorithm.
(b) Configuration switching: Improved adaptation architecture that eliminates channel switching delay. Figure 5: Comparison of (a) baseline and (b) improved architectures for online adaptation. in Fig. 6. The figure compares two switching approaches across three camera types (Bosch, HikVision, and Axis): channel switching (left boxes in each camera section) and configuration switching (right boxes in each camera section) achieved through down-sampling. Similarly to the channel switching measurements, the configuration switching measurements are also based on around 100 down-sampling (configuration switches) from the golden configuration to the other two configurations (i.e., 1→2 and 1→3). Fig. 6 indicates that the improved architecture in Fig. 5(b) could reduce the switching delay by approximately 20×. Importantly, it also mitigates the influence of camera-specific buffering and jitter-control mechanisms, making the system’s behavior more predictable across different camera models. Bandwidth consumption. The design in Fig. 5(b) also reduces network bandwidth usage by removing the additional camera streams required by the architecture in Fig. 5(a), denoted by red arrows. Precise F1-score calculation. The improved architecture in Fig. 5(b) also enhances synchronization between the ground-truth pipeline (green box) and the object detection pipeline (dark purple box). By using a single stream from Channel 1 instead of two separate camera streams, both pipelines operate on identically timestamped frames. This eliminates the need to align frames originating from different channels, which often have inconsistent timing due to independent buffering and Network Time Protocol (NTP) drifts. In some cameras, Real-Time Control Protocol (RTCP) packets, which normally carry timing
and synchronization information for media streams, are either absent or incomplete. Consequently, the timestamps embedded in frame packets are not sufficiently reliable for exact frame matching between the two pipelines, making it difficult to compute frame-level accuracy metrics such as the F1-score. The single-stream design solves this issue, ensuring accurate temporal alignment and reliable accuracy calculation. Importance of temporal alignment. To quantify the impact of temporal misalignment between the ground-truth pipeline and the object detection pipeline on accuracy calculation during adaptation, we analyzed a 10-minute recording from one of the COSMOS cameras (2nd -floor view in Fig. 3). For each configuration, we computed the average calculated F1-score over consecutive adaptation iterations of 0.2 s video duration. The results are shown in Fig. 7. Accurate F1-score measurement is critical for the adaptation algorithm to correctly model the relationship between configuration and accuracy. Even small frame misalignments between the two pipelines can distort this relationship. As shown in Fig. 7, a drift of just one frame can substantially alter the computed F1-score. For example, in both resolution settings (Figs. 7(a) and (b)), although the highest frame rate configuration truly achieves the best accuracy, a minor frame drift causes it to appear as the lowest-performing configuration. Such distortion can mislead the iterative algorithm, resulting in the selection of a suboptimal configuration and defeating the purpose of online adaptation.
6
Figure 6: Latency comparison of channel switching: transitioning between different camera streams (denoted by Channel SW), and configuration switching: changing resolution/frame rate parameters (denoted by Config SW), corresponding to the architectures in Fig. 5(a) and Fig. 5(b), respectively.
(b) Resolution: 960×540
(a) Resolution: 3840×2160
Figure 7: Impact of frame-drift between the ground-truth pipeline and the object detection pipeline on F1-score measurement for two resolution levels: (a) 3840×2160, and (b) 960×540, with different frame rates. B. Our Architecture: ASTRA
processes are decoupled through queued data flow. By structuring the adaptation processes as a Fork-Join system with input/output queues, the architecture decouples the concurrent tasks and eliminates global synchronization barriers, which has been shown to significantly improve throughput and resource utilization in parallel processing systems [69], [70].
The architecture should also manage differences in computational and time requirements among concurrent processes effectively to reduce their waiting (idle) time and ensure efficient resource utilization. For example, in Fig. 5(b), two major processes (outlined by white dashed contours) run in parallel during adaptation. Process 1 (the down-sampling and object detection pipeline) operates on sampled configurations, while Process 2 (the ground-truth pipeline) processes the golden-stream frames required to generate the proxy ground truth during a brief interval within the adaptation period. Because Process 1 includes additional steps such as down-sampling and switching between configurations and object detection models, its runtime can fluctuate depending on the chosen configuration. Process 2 runtime can also fluctuate due to variations in available GPU capacity. Consequently, either process can become the bottleneck at different iterations. To prevent one process from stalling the other and to maintain a high throughput, the architecture should execute these concurrent tasks asynchronously by introducing first-in-first-out (FIFO) queues at their inputs and outputs. This transforms the system into a classic Fork-Join model, where concurrent
In practice, such an asynchronous design can be implemented efficiently using programming constructs such as futures and promises [71], [72]. These mechanisms enable non-blocking execution of concurrent processes: a future represents the result of a computation that is not yet complete, such as a proxy ground-truth, sampled configuration, or measured accuracy, while a promise provides a way to deliver that result immediately once it becomes available. Leveraging these constructs allows both pipelines to progress at their own pace while ensuring that the adaptation loop receives synchronized inputs as soon as they are ready, maximizing responsiveness and overall system efficiency. ASTRA is an asynchronous version of the architecture in Fig. 5(b) that incorporates this queuing method. This architecture is illustrated in Fig. 8. It incorporates four queues: (i) video chunks queue, (ii) samples’ output queue, (iii) proxy ground-truth queue,
7
Figure 8: Overview of ASTRA, showing how the design in Fig. 5(b) is converted into an asynchronous Fork–Join–style pipeline that reduces idle time and improves throughput. and (iv) measured network latency queue. Chunkification. To construct the video chunk queue, ASTRA includes a chunkification component that receives decoded frames and groups a small number of consecutive frames (e.g., 5) into a single video chunk. A natural choice for the chunk length is the I-frame interval, the number of frames between two successive I-frames in an H264/AVC bitstream. Since all frames within an I-frame interval are decoded together and become available at nearly the same time, aligning the chunk size with the I-frame interval minimizes buffering and idle time. Short I-frame intervals (e.g., 5-6 frames) align naturally with the desired chunk size. Larger intervals, however, require the system to buffer the full interval before generating chunks. For instance, with a 15-frame interval and a 5-frame chunk size, the system must wait for all 15 frames to arrive before producing three chunks. ASTRA trades this minor latency in favor of small, consistent chunk sizes, which are crucial for iterative adaptation algorithm’s convergence. Golden Stream Bandwidth Consumption. In ASTRA (Fig. 8), the golden stream is transmitted only briefly during each adaptation period. This duration is substantially shorter than the adaptation period, which itself occupies only ∼10% of the time window. In our evaluation setup (Section VI), for each camera, the golden stream is transmitted for only ∼1 s, while adaptation takes 4– 7 s within a 60 s time window. The golden stream consumes approximately 10 Mbps of bandwidth while active. Moreover, ASTRA’s modular design allows the adaptation components, particularly Process 1 (the down-sampling and object detection pipeline) and Process 2 (the groundtruth pipeline), to be deployed near the cameras with highbandwidth connectivity, while the continuously running analytics components (the purple box labeled “Realtime Processing” in Fig. 8) can connect over bandwidthconstrained links.
eras lack programmable compute units. This measurement is complicated by three factors: (1) Lack of precise synchronization: Commodity cameras rely on NTP [73], which often exhibits clock offsets of tens of milliseconds due to network asymmetry [74], [75]. This margin of error is too large for precise per-chunk latency measurement. (2) RTP limitations: The Real-time Transport Protocol (RTP) protocol [16] is designed for playback jitter control and does not provide a mechanism for the receiver (edge server) to derive one-way network latency without tight clock synchronization. (3) Variable encoded chunk sizes: Dynamic compression rates cause encoded chunk sizes to fluctuate unpredictably based on video content and camera model. Consequently, the edge server cannot accurately estimate transmission time based solely on the selected configuration. ASTRA’s estimation of chunk-level network latency. It is important for the architecture to avoid incorporating latency measurements with high error margins into the adaptation algorithm. Otherwise, the iterative adaptation algorithm may converge to a configuration far from optimal and severely degrade performance. To estimate the end-to-end network latency of each video chunk, ASTRA relies only on receiver-side timestamps at the edge server. Let t0 and t1 denote the times the camera transmits the first and last bits of a video chunk, and t2 and t3 the times the edge receives those bits. Let dfirst and dlast denote the one-way path delays experienced by the first and last bits, respectively. By definition t2 = t0 + dfirst and t3 = t1 + dlast . The true chunk-level network latency that we aim to measure is Lnet total = t3 − t0 = (t1 − t0 ) + dlast ,
(1)
which is equal to the chunk’s transmission time plus the one-way network delay of its last bit.
C. Latency Measurement
The edge cannot observe t0 , but it can observe t3 and t2 , and hence it can measure
The iterative algorithm requires the total latency (network + inference) for each iteration t. While inference latency is measured directly within the object detection pipeline (see Fig. 8), measuring the network latency of video chunk Cht is challenging because commodity cam-
t3 − t2 = (t1 − t0 ) + (dlast − dfirst ),
(2)
Since the chunk sizes are small (∼ 0.2 s) the jitter term
8
(dlast − dfirst ) is negligible, making t 3 − t2 ≈ t 1 − t0 .
lead to less efficient memory access and extra delays [79], [80] (a few to tens of milliseconds). These delays, even though small, mostly ranging from a few to hundreds of milliseconds, could severely impact the adaptation when happening recurrent. Consider an adaptation algorithm that can determine a near-optimal configuration in ten iterations. With a total switching delay of even a few hundred milliseconds per iteration, the system would spend several seconds solely on configuration switching in the object detection pipeline.
(3)
Thus, using (1) and (3) and the fact that one-way delay dlast is roughly half of Round Trip Time (RTT), ASTRA approximates the per-chunk network latency as: RTT , (4) 2 The error introduced by assuming a small jitter term dlast −dfirst and by approximating one-way delay as RTT/2 is much smaller than the tens of milliseconds errors introduced by NTP timestamps [74], [75] or RTP sender timestamps. This way, ASTRA obtains a practical and sufficiently accurate per-chunk network-latency estimate using only receiver-side timestamps without requiring any programmable device on the camera. To measure RTTs, ASTRA’s network latency measurement component sends periodic ping requests to the camera during the transmission of each video chunk and computes the average RTT. Lnet total ≈ (t3 − t2 ) +
A. Mitigating Switching Overhead We applied the following optimizations to the dark purple Object detection pipeline in Fig. 8, reducing switching latency from hundreds of milliseconds to 10-20 ms. Efficient buffer management: Since video chunks are small, large inter-element queues only add switching delay. We cap these buffers to one frame and reuse small buffer pools to reduce draining, flushing, and reallocation overhead. Device memory versus host memory: In ASTRA’s object detection pipelines (dark purple and green boxes in Fig. 8), frames remain in GPU device memory throughout the hot path (decoder → preprocess → DNN → postprocess). We use NVMM/device-resident buffers and custom kernels on cudaMalloc pointers, avoiding CUDA zero-copy host memory because PCIe reads add higher per-chunk latency1 . Dynamic pipeline reconfiguration: ASTRA reconfigures the object-detection pipeline at runtime without recreating elements, reducing switching delay from hundreds of milliseconds to about ten milliseconds. It uses dynamic CPU/GPU memory allocation to update configurationdependent state and preloads all supported TensorRT engines into GPU memory so switching only selects an already initialized model. Idle models use memory but add no computation.
V. Configuration Switching in Object Detection Pipeline In addition to the overhead of switching video stream configurations, the frequent switching of configurations in the object detection pipeline (dark purple box in Fig. 8) also incurs non-negligible overhead. As discussed in Sec. III and Sec. IV-B, the adaptation iterations needs to occur back-to-back in short time slots (e.g., ∼ 0.2 s) which requires high frequency of switching between configurations in the object detection pipeline (dark purple box in Fig. 8). In this section, we discuss the overheads of configuration switching in this pipeline and how these overheads can be mitigated. To understand inherent delays of configurations switching in the object detection pipeline, we conducted a thorough profiling of the pipeline at runtime. The results of our profiling, combined with the measurements presented in previous studies such as [76], [77], [78], reveal that the primary overhead of such switching includes: Pipeline reconfiguration overheads: Processing video chunks with different resolutions and frame rates requires recompiling and initializing pipeline elements. For example, switching to a configuration with a different model variant or input shape requires reloading and initializing the inference engine and allocating model-specific buffers which can take tens to hundreds of milliseconds per switch. Memory allocation and deallocation overheads: Processing a sequence of video chunks with different configurations triggers frequent memory allocation and deallocation for intermediate buffers across the pre-processing, inference, and post-processing stages, introducing additional delays of a few to tens of milliseconds. Cache inefficiencies: Switching between video chunks and configurations can disrupt cache continuity, resulting in increased cache misses. Since GPU performance is highly dependent on the memory hierarchy, these cache misses
VI. Evaluations In this section, we evaluate ASTRA’s architecture in terms of accuracy, latency, and GPU utilization under diverse network conditions and video contents. We plug two iterative algorithms into the architecture (the ”content” of the black box in Fig. 8): (i) GP-UCB-C derived from [14] and (ii) Chameleon++ derived from [8]. These algorithms represent two common classes of adaptation methods: (1) theoretically guaranteed algorithms (GP-UCB-C) and (2) empirically estimated algorithms (Chameleon++). We compare their performance against two baselines: (a) the offline optimal (upper bound on achievable performance) and (b) a ground-truth-free adaptation approach that estimates accuracy using confidence scores rather than ground truth. 1 We have published measurements of memory-type impact on latency in [3]
9
As described in Section III, the goal of the iterative algorithm is to select a configuration x ∈ X that maximizes a weighted accuracy–latency utility, max ca A(x) − cl L(x) x∈X
s.t.
It assumes that configuration parameters (e.g., resolution and frame rate) contribute independently to accuracy in order to estimate the accuracy of all remaining configurations, without exhaustive profiling. Chameleon++ extends this approach by also measuring inference latency during profiling. For configurations that are not profiled, their latency is estimated using the same parameterindependence assumption used by Chameleon for accuracy estimation. This allows Chameleon++ to jointly consider accuracy and latency without modifying Chameleon’s core algorithm. AutoML++ (ground-truth-free adaptation): In this benchmark, the iterative algorithm is the GP-UCB-C algorithm that uses confidence scores generated by the object detection model as the accuracy metric instead of F1-scores that need proxy ground truth. This approach is inspired by the AutoML adaptation platform [9], a GP-UCB-based method that uses confidence score as the accuracy metric. Offline optimal: This benchmark uses, for each time window, the optimal configuration given full offline knowledge of A(x) and L(x). For multi-camera experiments, A(x) and L(x) in (P1) are computed as averages over all cameras. The optimal configuration is found by exhaustively evaluating all configurations over the entire time window and selecting the best feasible one according to (P1). Importantly, optimality is defined with respect to (P1), including its accuracy and latency constraints, rather than with respect to either metric individually. For example, in the accuracy-constrained setting, the offline optimal is the lowest-latency configuration that satisfies the minimumaccuracy threshold. An online method may show lower latency only by violating the corresponding accuracy threshold. Therefore, appearing better on one individual metric does not indicate a better feasible solution to P1. This benchmark has no adaptation period, incurs no adaptation overhead, and serves as an upper bound benchmark on achievable performance.
A(x) ≥ α, L(x) ≤ β, (P1)
where A(x) and L(x) denote average accuracy and latency, ca , cl ≥ 0 are application-specific weights, and α, β are application-specific accuracy and latency constraints. GP-UCB-C: Following the constrained Bayesian optimization method introduced in [14], we model the performance of each configuration using Gaussian Processes (GPs) [81]. Specifically, the accuracy function A(x) and latency function L(x) over the discrete configuration set X are treated as correlated, unknown functions. For each configuration x and iteration t, we denote by UCBA t (x) L L and LCBA (x) (resp. UCB (x) and LCB (x)) the upper t t t and lower confidence bounds on the accuracy A(x) (resp. latency L(x)) given by the GP posterior. After each observation, the GP posterior provides predictive means µA t (x), A A µL (x) and confidence intervals [LCB (x), UCB (x)] and t t t L [LCBL (x), UCB (x)] for accuracy and latency, respect t tively. At each iteration t, the algorithm selects the next configuration xt by maximizing the UCB on the utility, L xt = arg max ca UCBA t (x) − cl LCBt (x), x∈At
over the feasible set L At = x ∈ X UCBA t (x) ≥ α, LCBt (x) ≤ β . Configurations that are likely to violate the accuracy or latency constraints are therefore removed from At . After evaluating At (xt ) and Lt (xt ), the GP posterior is updated and the feasible set At+1 is recomputed. The algorithm then identifies the current best feasible configuration L ht = arg max ca µA t (x) − cl µt (x) x∈At
and the most promising alternative
L A L ca LCBA t (ht )−cl UCBt (ht ) ≥ ca UCBt (rt )−cl LCBt (rt ),
Object detection models. In our experiments, the adaptation is performed for the object detection task, specifically pedestrians and vehicles detection, using YOLOv8 [82] (Section VI-A) and YOLOv4 [83] (Sections VI-B and VI-C) models. Importantly, ASTRA is model-agnostic: any DNN-based object detection model, including newer versions of YOLO, can be integrated into its architecture.
and returns ht as the selected configuration. We refer to this constrained UCB-based selection rule, together with the feasibility update and the early stopping condition above, as GP-UCB-C. Chameleon++: This algorithm is an improvement of the Chameleon algorithm [8] that adds latency-awareness while preserving the original search logic. Chameleon is a profiling-based configuration search algorithm. Namely, it measures the accuracy of a small subset of configurations.
Optimization setup and parameters. The evaluations are performed for two sets of parameters in (P1): (i) accuracy constrained, i.e., ca = 0, cl = 1, β = +∞, and α ≫ 0 (Sections VI-A and VI-B) and (ii) latency constrained, i.e., ca = 1, cl = 0, α = 0, and β ≪ +∞ (Section VI-C). The configuration set X used consists of 12 elements, each of which is a tuple (ri , fi ) where ri is one of the values {832, 608, 416} and fi is one of the values {30, 15, 10, 5}. A video stream with configuration (r, f ) has resolution r × r
rt = arg
max
x∈At \{ht }
L ca UCBA t (x) − cl LCBt (x).
The procedure stops at iteration t when the lower confidence bound on the utility of ht is no smaller than the upper confidence bound on the utility of rt , i.e.,
10
and frame rate f fps. These three resolution levels and four frame-rate levels provide sufficiently distinct operating points for the adaptation algorithm. Configurations with only small differences in resolution or frame rate typically exhibit similar accuracy and latency, and therefore using a much denser configuration space would increase the adaptation search space with limited additional performance diversity. Increasing the number of sufficiently distinct candidate configurations can increase adaptation time by enlarging the search space. All experiments use an F1-score threshold of 0.7 and chunk size τ =0.2 s (also referred to as time slot). The time window size is determined by the time required for adaptation and the video content/network volatility. A smaller time window is needed for higher volatility, but it is essential to keep time window size notably larger than the adaptation period, e.g., ∼10x (see Fig. 4). ASTRA’s adaptation period is typically 4-7 s when considering 12 configurations. Accordingly, time window length of 60 s is used in all the experiments. Consequently, for a 60second time window and 0.2-second slot duration, the total 60 = 300. number of time slots is T = 0.2 Accuracy and latency per time window. During the adaptation period of camera i, in the n-th time window, the system (object detection pipeline denoted by light orange box in Fig. 8) uses the configuration chosen in the (n − 1) time window (denoted by x(i,n−1) ). For the remainder of the time window, it uses the configuration chosen by the adaptation algorithm for time window n (i.e., x(i,n) ). Therefore, for K cameras, the average achieved accuracy and latency at time window n are calculated as follows: ! ηP i,n K 1 T 1 P P (i,n−1) (i,n) An = At,i (x )+ At,i (x ) K i=1 T t=1 t=ηi,n +1 ! ηP i,n K 1 T 1 P P Ln = Lt,i (x(i,n−1) ) + Lt,i (x(i,n) ) K i=1 T t=1 t=ηi,n +1 (5) where T is the total number of slots in a time window, and ηi,n is the number of slots in the adaptation period of camera i within the nth window. For camera i, At,i (x(i,n−1) ) and Lt,i (x(i,n−1) ) denote the accuracy and latency achieved using configuration x(i,n−1) at time slot t. We performed evaluations for two settings: (i) realtime adaptation of two live cameras within the COSMOS testbed (Section VI-A), and (ii) real-time adaptation of up to eight emulated cameras using a comprehensive video dataset on the Google Cloud Platform (Sections VI-B and VI-C). Network latency dataset. In both settings, the live and emulated cameras are connected to the servers via highspeed links (e.g., COSMOS cameras use fiber optics). This ideal connectivity is atypical for distributed edge computing systems. Therefore, to evaluate the performance of ASTRA under realistic network conditions between
Figure 9: CDF of (a) mean RTT and (b) jitter from end users to the nearest Alibaba Cloud sites under 4G [84]. the cameras and the edge servers, we used a dataset of extensive real network latency measurements provided in [84] to induce extra latency. The dataset contains the RTT between end users and the nearest Alibaba Cloud site under 4G. The empirical Cumulative Distribution Function (CDF) of the mean RTT and jitter (RTT coefficient of variation measured as the standard deviation divided by the empirical mean) are presented in Fig. 9. The detailed dataset of these measurements was graciously provided by the authors of [84]. A. Deployment in COSMOS Testbed We deployed ASTRA within the COSMOS testbed [15] and assessed its adaptation capabilities using two live cameras. This deployment validates the system’s real-time performance under actual physical constraints. This effectively addresses a gap in prior research, where evaluations involving live cameras are rarely conducted. COSMOS cameras are located on the 1st and 2nd floor of a building, viewing 120th St. and Amsterdam Ave. intersection, NYC (see Fig. 3). Moreover, two COSMOS servers running Ubuntu 20.04 with Intel Xeon [email protected] and an NVIDIA V100 GPU, were used to deploy ASTRA. One server hosts the real-time processing (denoted by the light purple container in Fig. 8) and the other hosts the periodic adaptation (denoted by the pale red container in Fig. 8). We emulated realistic network latencies between cameras and the servers using the dataset described above. Parameter tuning (α). We experimented with α ranging from 0.4 to 1 (ideal). We observed that values below 0.85 compromised pedestrian detection accuracy, while values above 0.85 incurred high latency and GPU utilization. Consequently, we selected α = 0.85 as the threshold for the accuracy-constrained scenario to balance accuracy and overhead. We executed ASTRA with the GP-UCB-C algorithm during two 30-minute periods, representing peak (4 PM) and off-peak (4 AM) hours. Fig. 10 shows the average achieved accuracy, An , and latency, Ln , for each time window, as defined in (5). The results confirm that ASTRA consistently maintains accuracy above or close to the required threshold of 0.85 in both periods. Notably, latency is lower during off-peak hours, as reduced pedestrian and vehicle density allows the system to select lower resolutions and frame rates. To illustrate ASTRA’s adaptability, we provide snapshots of the intersection at specific timestamps. For instance, at window n = 19 (day),
11
n = 19 (Day)
n = 24 (Day)
n = 8 (Night)
n = 27 (Night)
Figure 10: Left: ASTRA’s (with GP-UCB-C as the algorithm) achieved average accuracy and latency at each time window, during two 30-minute periods (peak/day and off-peak/night). Right: Snapshots of the intersection at circled time windows. ASTRA leveraged the low traffic density to reduce the resolution/frame rate, lowering latency while meeting the accuracy target. Conversely, at n = 24, increased density and speed prompted ASTRA to upgrade the configuration to preserve accuracy, resulting in higher latency. A similar adaptive pattern is observed at night (n = 8 and n = 27).
to represent a range of traffic densities and weather conditions (sunny, rainy, etc.). Multi-camera. In our setting, in a single instance of ASTRA, real-time adaptation of one camera takes 4–7 s. Increasing the number of cameras does not increase the per-camera configuration search space, but increases the aggregate adaptation workload of an ASTRA instance, since each camera requires an independent adaptation process. Depending on the available adaptation resources, cameras can be adapted individually or in batches. In our evaluation, cameras are adapted sequentially, with each camera adapted once during a repeating 60 s time window and with a different adaptation offset within the window. Therefore, a 60 s time window allows for ⌊60/7⌋ = 8 cameras. To increase the number of cameras, multiple instances of ASTRA architecture with similar performance can be set up, thereby scaling/multiplying the 8camera system. Accordingly, we conducted adaptation for up to eight emulated cameras using all the benchmark algorithms for the accuracy-constrained scenario. Each emulated camera is associated with a set of videos (26 videos for four or more emulated cameras and 52 videos for two emulated cameras). GPU utilization was measured as the average utilization over both GPU-equipped edge servers in the entire time window. Performance comparison. Fig. 11 shows the system-level impact of increasing the number of cameras, and consequently the aggregate adaptation workload, from 2 to 4 and 8 cameras. Despite this increase, compared to offline optimal, ASTRA with GP-UCB-C algorithm achieves near-optimal performance (less than 10% deviation) with GPU utilization around 2%, 4%, and 9% for 2, 4, and 8 cameras, respectively, which means around 2% per camera. Moreover, it satisfies the accuracy threshold of 0.85 in more than 90% of the time in all cases. Note that an online method can occasionally appear better than the offline optimal on one individual metric ((in this case, lower latency), but such cases correspond to a violation of the associated constraint (in this case, minimum accuracy of 85%) and therefore do not represent a better feasible
B. Controlled and Reproducible Evaluation To thoroughly evaluate ASTRA and ensure a fair comparison of iterative algorithms, it is necessary to use identical and reproducible conditions, such as consistent video content and network conditions. Therefore, we deployed ASTRA in Google Cloud VM instances and emulated multiple cameras using a comprehensive video dataset. Experiments were conducted using three VM instances hosted in the us-east1 region (Moncks Corner, South Carolina). Two of these instances were equipped with an NVIDIA T4 GPU and were used to emulate the real-time processing (denoted by the light purple container in Fig. 8) and the periodic adaptation (denoted by the pale red container in Fig. 8). The third instance was dedicated to streaming H264-encoded videos over RTP (RTSP over TCP) using GStreamer, thereby emulating live cameras. This setup contrasts with prior studies (e.g., [18], [19], [20], [21], [22]), which typically rely on reading pre-recorded video files directly from local storage. By streaming videos over the network, we subject ASTRA to real-time latency and throughput constraints, whereas filebased approaches do not capture these challenges inherent to live video ingestion. To emulate realistic network latencies between the (emulated) cameras and the edge servers, we used the network latency dataset introduced earlier in this section. Video dataset. The dataset consists of videos recorded simultaneously from the two COSMOS testbed’s cameras (used in Section VI-A). It includes 52 pairs of ten-minute videos (104 in total). Each pair is recorded from the two cameras simultaneously. The video pairs are captured eight times daily from morning to evening, in June, July, and Sept. 2022. The videos were purposely selected
12
Figure 11: Empirical CDF of average achieved accuracy (with lower bound of 0.85), average achieved latency, and average GPU utilization per time window for the setting with (a) two, (b) four, and (c) eight cameras.
Figure 12: Performance comparison under varying network latency conditions. solution to P1. GPU utilization of ASTRA with Chameleon++ is close to GP-UCB-C. However, Chameleon++ does not meet the specified accuracy target, α =0.85, in more than 30% of the times across all the evaluated scenarios. AutoML++ (GPU-UCB-C with no ground-truth) has the lowest GPU utilization as it does not require proxy ground truth. Despite having low GPU overhead, AutoML++ fails to achieve the minimum accuracy required around 75%, 85%, and 90% of times for the 2, 4, and 8 camera scenarios, respectively. This confirms the importance of incorporating proxy ground-truth for adaptation.
C is more consistent. Notably, GP-UCB-C’s performance is closest to the optimal and demonstrates the least variation. These evaluations show that ASTRA, when instantiated with a principled iterative algorithm such as GP-UCBC, can track the best configuration under dynamic workloads with minimal overhead. Across single- and multicamera experiments, ASTRA consistently approaches the performance of an offline optimal benchmark that has full knowledge of accuracy and latency and incurs no adaptation overhead, demonstrating that fast online adaptation is both feasible and effective. Relationship to Network Conditions. The experiments in Fig. 12 explicitly vary network latency and jitter. However, reduced bandwidth and packet loss can affect ASTRA in a similar way by increasing the time required for a video chunk to arrive at the edge server. Since ASTRA receives video streams over TCP, reduced available bandwidth increases chunk transmission and queueing time, while packet loss can trigger retransmissions and further delay chunk delivery. Throughput limitations may also reduce the delivered video quality, which can lower the measured analytics accuracy. Therefore, the effects of adverse network conditions are reflected in ASTRA’s measured chunk-level latency and accuracy and, consequently, in the configuration selected by the adaptation algorithm. Thus, the experiments in Fig. 12 evaluate ASTRA’s response to increased chunk-delivery latency, which is a common consequence of several adverse network conditions, including increased propagation delay, jitter, bandwidth limitations,
C. Varying Network Latencies To assess adaptation robustness to varying network conditions, we conducted another experiment with increased latency and jitter. The measured latencies in [84] and similar network latency measurement studies [85], [86] are relatively small. Therefore, following the method used in [25] to evaluate AWStream, we added Gaussiandistributed delay with a mean between 10 ms and 150 ms and a variance of 20% between the emulated camera and the edge servers. To determine how effectively the GPUCB-C algorithm and other benchmarks adapt to network variation, we consider latency constrained optimization with β = 250 ms. The results of these experiments are presented in Fig. 12. As can be observed, the performance of AutoML++ is unstable under varying network conditions. In contrast, the performance of Chameleon++ and GP-UCB-
13
and packet loss.
[13] Y. Faye, F. Faticanti, S. Jain, and F. Bronzino, “VideoJam: Self-balancing architecture for live video analytics,” in Proc. IEEE/ACM SEC, 2024. [14] W. Xu, Y. Jiang, B. Svetozarevic, and C. Jones, “Constrained efficient global optimization of expensive black-box functions,” in ICML, 2023. [15] S. Yang, E. Bailey, Z. Yang, J. Ostrometzky, G. Zussman, I. Seskar, and Z. Kostic, “COSMOS smart intersection: Edge compute and communications for bird’s eye object tracking,” in Proc. SmartEdge, 2020. [16] H. Schulzrinne, S. Casner, R. Frederick, and V. Jacobson, “RFC3550: RTP: A transport protocol for real-time applications,” 2003. [17] W. Taymans, S. Baker, A. Wingo, R. S. Bultje, and S. Kost, “Gstreamer application development manual (1.2. 3),” Publicado en la Web, vol. 72, 2013. [18] K. Du, Y. Liu, Y. Hao, Q. Zhang, H. Wang, Y. Huang, G. Ananthanarayanan, and J. Jiang, “Oneadapt: Fast adaptation for deep learning applications via backpropagation,” in ACM SoCC, 2023. [19] M. Wong, M. Ramanujam, G. Balakrishnan, and R. Netravali, “MadEye: Boosting live video analytics accuracy with adaptive camera configurations,” in USENIX NSDI, 2024. [20] M. Zhang, J. Li, H. Zhao, L. Shen, and J. Liu, “Starstream: Live video analytics over space networking,” in Proc. ACM MM, 2024. [21] L. Wang, K. Lu, N. Zhang, X. Qu, J. Wang, J. Wan, G. Li, and J. Xiao, “Shoggoth: towards efficient edge-cloud collaborative real-time video inference via adaptive online learning,” in Proc. ACM/IEEE DAC, 2023. [22] M. Zhao, S. Liu, F. Wu, and G. Chen, “Responsive DNN adaptation for video analytics against environment shift via hierarchical mobile-cloud collaborations,” in Proc. ACM SenSys, 2025. [23] Y. Kong, P. Yang, and Y. Cheng, “Edge-assisted on-device model update for video analytics in adverse environments,” in Proc. ACM MM, 2023. [24] K. Wu, Y. Jin, W. Miao, Z. Zeng, Z. Qian, J. Wang, M. Zhou, and T. Cao, “Soudain: Online adaptive profile configuration for real-time video analytics,” in Proc. IEEE IWQOS, 2021. [25] B. Zhang, X. Jin, S. Ratnasamy, J. Wawrzynek, and E. A. Lee, “Awstream: Adaptive wide-area streaming analytics,” in Proc. ACM SIGCOMM, 2018. [26] W.-J. Kim and C.-H. Youn, “Lightweight online profiling-based configuration adaptation for video analytics system in edge computing,” IEEE Access, vol. 8, pp. 116 881–116 899, 2020. [27] C. Wang, S. Zhang, Y. Chen, Z. Qian, J. Wu, and M. Xiao, “Joint configuration adaptation and bandwidth allocation for edge-based real-time video analytics,” in Proc. IEEE INFOCOM, 2020. [28] J. Yi, G. Lee, M. Jeong, S. Shin, D. Kim, and Y. Lee, “Towards end-to-end latency guarantee in MEC live video analytics with App-RAN mutual awareness,” in Proc. ACM MobiSys, 2025. [29] K. Yang, M. Jeong, J. Yi, J. Lee, K. Park, and Y. Lee, “Logan: Loss-tolerant live video analytics system,” in Proc. ACM MobiCom, 2024. [30] Z. Li, Y. Zhu, S. Mumtaz, L. Kong, and B. Li, “RIVA: Communication-efficient streaming control for real-time industrial video analytics,” IEEE J. Sel. Areas Commun., vol. 43, no. 10, pp. 3548–3563, 2025. [31] Z. Li, M. Zhang, and Y. Zhu, “OAVS: Efficient online learning of streaming policies for drone-sourced live video analytics,” in Proc. IEEE/ACM IWQoS, 2024. [32] M. Bala, A. Chanana, X. Chen, Q. Dong, T. Eiszler, J. Xu, P. Pillai, and M. Satyanarayanan, “The OODA loop of cloudletbased autonomous drones,” in Proc. IEEE/ACM SEC, 2024. [33] S. P. Chinchali, E. Cidon, E. Pergament, T. Chu, and S. Katti, “Neural networks meet physical networks: Distributed inference between edge devices and the cloud,” in Proc. ACM HotNets, 2018. [34] Z. Zhao, K. M. Barijough, and A. Gerstlauer, “Deepthings: Distributed adaptive deep learning inference on resourceconstrained IoT edge clusters,” IEEE Trans. CAD, vol. 37, no. 11, pp. 2348–2359, 2018.
VII. Conclusions We presented ASTRA, a lightweight architecture for executing a broad class of iterative adaptation algorithms for real-time video analytics. ASTRA enables fast configuration adaptation with minimal overhead and supports algorithms ranging from theoretically grounded methods to empirical, profile-driven approaches. Our evaluations show that ASTRA, when paired with a principled algorithm, consistently tracks near-optimal configurations and achieves performance close to an offline optimal benchmark that has full knowledge of accuracy and latency and incurs no overhead (upper bound on achievable performance). These results demonstrate that principled online adaptation is both practical and effective for live video analytics systems. Future work will focus on several directions: (i) scaling live evaluations to significantly larger camera networks, (ii) developing adaptive time windows to better handle rapid environmental fluctuations, and (iii) extending ASTRA to large-scale geo-distributed edge/cloud systems with dynamic resource allocation. References [1] M. Ghasemi, S. Kleisarchaki, T. Calmant, L. Gürgen, J. Ghaderi, Z. Kostic, and G. Zussman, “Demo: Real-time camera analytics for enhancing traffic intersection safety,” in Proc. ACM MobiSys, 2022. [2] M. Ghasemi, S. Kleisarchaki, T. Calmant, J. Lu, S. Ojha, Z. Kostic, L. Gürgen, G. Zussman, and J. Ghaderi, “Real-time multi-camera analytics for traffic information extraction and visualization,” in Proc. IEEE PerCom, 2023. [3] M. Ghasemi, Y. Fu, X. Ouyang, P. Wang, M. K. Turkcan, J. Tavori, S. Kleisarchaki, T. Calmant, L. Gürgen, Z. Kostic et al., “Real-time video analytics for urban safety: Deployment over edge and end devices,” in Proc. ACM/IEEE SEC, 2025. [4] D. Raychaudhuri, I. Seskar, G. Zussman, T. Korakis, D. Kilper, T. Chen, J. Kolodziejski, M. Sherman, Z. Kostic, X. Gu et al., “Challenge: COSMOS: A city-scale programmable testbed for experimentation with advanced wireless,” in Proc. ACM MobiCom, 2020. [5] COSMOS Project, “Hardware: Cameras,” 2023. [Online]. Available: https://wiki.cosmos-lab.org/wiki/Hardware/Cameras [6] M. Zhang, F. Wang, and J. Liu, “Casva: Configuration-adaptive streaming for live video analytics,” in Proc. IEEE INFOCOM, 2022. [7] M. Khani, G. Ananthanarayanan, K. Hsieh, J. Jiang, R. Netravali, Y. Shu, M. Alizadeh, and V. Bahl, “RECL: Responsive resource-efficient continuous learning for video analytics,” in Proc. USENIX NSDI, 2023. [8] J. Jiang, G. Ananthanarayanan, P. Bodik, S. Sen, and I. Stoica, “Chameleon: Scalable adaptation of video analytics,” in Proc. ACM SIGCOMM, 2018. [9] A. Galanopoulos, J. A. Ayala-Romero, D. J. Leith, and G. Iosifidis, “Automl for video analytics with edge computing,” in Proc. IEEE INFOCOM, 2021. [10] M. Ghasemi, Z. Kostic, G. Zussman, and J. Ghaderi, “Systems and methods for adaptive streaming for real-time video analytics (ASTRA),” U.S. Patent Application Publication US 2026/0163998 A1, Jun. 2026. [11] V. Nigade, P. Bauszat, H. Bal, and L. Wang, “Jellyfish: timely inference serving for dynamic edge networks,” in Proc. IEEE RTSS, 2022. [12] R. Bhardwaj, Z. Xia, G. Ananthanarayanan, J. Jiang, Y. Shu, N. Karianakis, K. Hsieh, P. Bahl, and I. Stoica, “Ekya: Continuous learning of video analytics models on edge compute servers,” in Proc. USENIX NSDI, 2022.
14
[35] G. Bartolomeo, M. Yosofie, S. Bäurle, O. Haluszczynski, N. Mohan, and J. Ott, “Oakestra: A lightweight hierarchical orchestration framework for edge computing,” in Proc. USENIX ATC, 2023. [36] A. Jano, M. Mert Bese, N. Mohan, W. Kellerer, and J. Ott, “nextGSIM: Towards simulating network resource management for beyond 5G networks,” in IEEE Future Netw. World Forum., 2023. [37] K.-J. Fu, Y.-T. Yang, and H.-Y. Wei, “Split computing video analytics performance enhancement with auction-based resource management,” IEEE Access, vol. 10, pp. 106 495–106 505, 2022. [38] Z. Gao, S. Sun, Y. Zhang, Z. Mo, and C. Zhao, “Edgesp: Scalable multi-device parallel DNN inference on heterogeneous edge clusters,” in Proc. ICA3PP, 2021. [39] Z. Yang, W. Ji, Q. Guo, and Z. Wang, “JAVP: Joint-aware video processing with edge-cloud collaboration for DNN inference,” in Proc.ACM MM, 2023. [40] W. Liu, J. Geng, Z. Zhu, J. Cao, and Z. Lian, “Sniper: Cloudedge collaborative inference scheduling with neural network similarity modeling,” in Proc. ACM/IEEE DAC, 2022. [41] T. Murad, A. Nguyen, and Z. Yan, “DAO: Dynamic adaptive offloading for video analytics,” in Proc. ACM MM, 2022. [42] J. Wei, T. Cao, S. Cao, S. Jiang, S. Fu, M. Yang, Y. Zhang, and Y. Liu, “NN-stretch: Automatic neural network branching for parallel inference on heterogeneous multi-processors,” in Proc. ACM MobiSys, 2023. [43] X. Hou, Y. Guan, and T. Han, “ViEdge: Video analytics on distributed edge,” ACM Trans. IOT, 2025. [44] F. Jia, D. Zhang, T. Cao, S. Jiang, Y. Liu, J. Ren, and Y. Zhang, “Codl: Efficient CPU-GPU co-execution for deep learning inference on mobile devices,” in Proc. ACM MobiSys, 2022. [45] N. Ling, X. Huang, Z. Zhao, N. Guan, Z. Yan, and G. Xing, “BlastNet: Exploiting Duo-Blocks for cross-processor real-time DNN inference,” in Proc. ACM SenSys, 2022. [46] Q. Dong, J. Xu, P. Pillai, and M. Satyanarayanan, “Does accurate real-time AI need edge offload?” in Proc. ACM/IEEE SEC, 2025. [47] Z. Fang, S. Hu, L. Yang, Y. Deng, X. Chen, and Y. Fang, “PIB: Prioritized information bottleneck framework for collaborative edge video analytics,” in Proc. IEEE GLOBECOM, 2024. [48] H. Sun, X. Zhang, B. Zhang, K. Sha, and W. Shi, “Optimal task offloading and trajectory planning algorithms for collaborative video analytics with UAV-assisted edge in disaster rescue,” IEEE Trans. Veh. Technol., vol. 73, no. 5, pp. 6811–6828, 2024. [49] F. Bastani, S. He, A. Balasingam, K. Gopalakrishnan, M. Alizadeh, H. Balakrishnan, M. Cafarella, T. Kraska, and S. Madden, “MIRIS: Fast object track queries in video,” in Proc. ACM SIGMOD, 2020. [50] O. Moll, F. Bastani, S. Madden, M. Stonebraker, V. Gadepally, and T. Kraska, “Exsample: Efficient searches on video repositories through adaptive sampling,” in Proc. IEEE ICDE, 2022. [51] Y. Li, A. Padmanabhan, P. Zhao, Y. Wang, G. H. Xu, and R. Netravali, “Reducto: On-camera filtering for resourceefficient real-time video analytics,” in Proc. ACM SIGCOMM, 2020. [52] S. Paul, U. Drolia, Y. C. Hu, and S. T. Chakradhar, “Aqua: Analytical quality assessment for optimizing video analytics systems,” in Proc. IEEE/ACM SEC, 2021. [53] H. Wang, Q. Li, H. Sun, Z. Chen, Y. Hao, J. Peng, Z. Yuan, J. Fu, and Y. Jiang, “VaBUS: Edge-cloud real-time video analytics via background understanding and subtraction,” IEEE J. Sel. Areas Commun., vol. 41, no. 1, pp. 90–106, 2022. [54] I. Xarchakos and N. Koudas, “SVQ: Streaming video queries,” in Proc. ACM SIGMOD, 2019. [55] K. Apicharttrisorn, X. Ran, J. Chen, S. V. Krishnamurthy, and A. K. Roy-Chowdhury, “Frugal following: Power thrifty object detection and tracking for mobile augmented reality,” in Proc. ACM SenSys, 2019. [56] L. Liu, H. Li, and M. Gruteser, “Edge-assisted real-time object detection for mobile augmented reality,” in Proc. ACM MobiCom, 2019.
[57] H. Guo, S. Yao, Z. Yang, Q. Zhou, and K. Nahrstedt, “CrossRoI: Cross-camera region of interest optimization for efficient realtime video analytics at scale,” in Proc. ACM MMSys, 2021. [58] S. Liu, T. Wang, J. Li, D. Sun, M. Srivastava, and T. Abdelzaher, “Adamask: Enabling machine-centric video streaming with adaptive frame masking for DNN inference offloading,” in Proc. ACM MM, 2022. [59] W. Zhang, Z. He, L. Liu, Z. Jia, Y. Liu, M. Gruteser, D. Raychaudhuri, and Y. Zhang, “Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloading,” in Proc. ACM MobiCom, 2021. [60] K. Du, A. Pervaiz, X. Yuan, A. Chowdhery, Q. Zhang, H. Hoffmann, and J. Jiang, “Server-driven video streaming for deep learning inference,” in Proc. ACM SIGCOMM, 2020. [61] K. Du, Q. Zhang, A. Arapin, H. Wang, Z. Xia, and J. Jiang, “Accmpeg: Optimizing video encoding for accurate video analytics,” in Proc. MLSys, 2022. [62] S. Kim, K. Bin, D. Yang, S. Ha, S. Chong, and K. Lee, “ENTRO: Tackling the encoding and networking trade-off in offloaded video analytics,” in Proc. ACM MM, 2023. [63] Y. Lu, S. Jiang, T. Cao, and Y. Shu, “Turbo: Opportunistic enhancement for edge video analytics,” in Proc.ACM SenSys, 2022. [64] W. Wang, L. Mi, S. Cen, H. Dai, Y. Li, X. Fu, and Y. Liu, “Region-based content enhancement for efficient video analytics at the edge,” in USENIX NSDI, 2025. [65] I. Gokarn, Y. Hu, T. Abdelzaher, and A. Misra, “JIGSAW: Edge-based streaming perception over spatially overlapped multi-camera deployments,” in Proc. IEEE ICME, 2024. [66] D. Kang, J. Emmons, F. Abuzaid, P. Bailis, and M. Zaharia, “Noscope: Optimizing neural network queries over video at scale,” arXiv preprint arXiv:1703.02529, 2017. [67] H. Zhang, G. Ananthanarayanan, P. Bodik, M. Philipose, P. Bahl, and M. J. Freedman, “Live video analytics at scale with approximation and delay-tolerance,” in USENIX NSDI, 2017. [68] D. Wu, D. Zhang, M. Zhang, R. Zhang, F. Wang, and S. Cui, “ILCAS: Imitation learning-based configurationadaptive streaming for live video analytics with cross-camera collaboration,” IEEE Trans. Mobile Comput., 2023. [69] A. Marin, S. Rossi, and C. Williamson, “Speed scaling in forkjoin queues: a comparative study,” in Proc. ACM VALUETOOLS, 2020. [70] G. Pinto, A. Canino, F. Castor, G. Xu, and Y. D. Liu, “Understanding and overcoming parallelism bottlenecks in forkjoin applications,” in Proc. IEEE/ACM ASE, 2017. [71] J. Bachan, S. B. Baden, S. Hofmeyr, M. Jacquelin, A. Kamil, D. Bonachea, P. H. Hargrove, and H. Ahmed, “UPC++: A high-performance communication framework for asynchronous computation,” in Proc. IEEE IPDPS, 2019. [72] S. R. Paul, A. Hayashi, K. Chen, Y. Elmougy, and V. Sarkar, “A fine-grained asynchronous bulk synchronous parallelism model for PGAS applications,” J. Comp. Sc., vol. 69, p. 102014, 2023. [73] D. L. Mills, “Internet time synchronization: the network time protocol,” IEEE Trans. Commun., vol. 39, no. 10, pp. 1482– 1493, 1991. [74] C. Vélez, J. Díaz, A. Osuna, H. Álvarez-Martínez, and H. Esteban, “Assessment of an NTP service calibration over a local area network,” Meas. Sci. Technol., vol. 35, no. 6, p. 065013, 2024. [75] K. Liu, Z. Lu, Y. Zheng, and G. Tang, “Research on TDOA multiple stations time synchronization based on RTL-SDR,” in Proc. ACM ICTCE, 2022. [76] J. Yi, M. R. Islam, S. Aggarwal, D. Koutsonikolas, Y. C. Hu, and Z. Yan, “An analysis of delay in live 360° video streaming systems,” in Proc. ACM Multimedia, 2020. [77] Z. Yan and J. Yi, “Dissecting latency in 360 video camera sensing systems,” MDPI Sensors, vol. 22, no. 16, p. 6001, 2022. [78] C. Yao, W. Liu, W. Tang, and S. Hu, “EAIS: Energy-aware adaptive scheduling for CNN inference on high-performance GPUs,” Elsevier Future Gen. Comput. Sys., vol. 130, pp. 253– 268, 2022.
15
[79] F. Candel, S. Petit, A. Valero, and J. Sahuquillo, “Improving GPU cache hierarchy performance with a fetch and replacement cache,” in Proc. Euro-Par ICPDC, 2018. [80] B. Sepanski, T. Zhao, H. Johansen, and S. Williams, “Maximizing performance through memory hierarchy-driven data layout transformations,” in Proc. IEEE/ACM MCHPC, 2022. [81] N. Srinivas, A. Krause, S. M. Kakade, and M. Seeger, “Gaussian process optimization in the bandit setting: No regret and experimental design,” arXiv preprint arXiv:0912.3995, 2009. [82] J. Terven and D. Cordova-Esparza, “A comprehensive review of YOLO: From YOLOv1 to YOLOv8 and beyond,” arXiv preprint arXiv:2304.00501, 2023. [83] A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao, “YOLOv4: Optimal speed and accuracy of object detection,” arXiv preprint arXiv:2004.10934, 2020. [84] M. Xu, Z. Fu, X. Ma, L. Zhang, Y. Li, F. Qian, S. Wang, K. Li, J. Yang, and X. Liu, “From cloud to edge: a first look at public edge platforms,” in Proc. ACM IMC, 2021. [85] T. K. Dang, N. Mohan, L. Corneo, A. Zavodovski, J. Ott, and J. Kangasharju, “Cloudy with a chance of short RTTs: Analyzing cloud connectivity in the internet,” in Proc. ACM IMC, 2021. [86] B. Schlinker, I. Cunha, Y.-C. Chiu, S. Sundaresan, and E. KatzBassett, “Internet performance from facebook’s edge,” in Proc. ACM IMC, 2019.
16