Conceptio › Archive › arXiv CS
arXiv CSopen access

A Kubernetes-Native Request Router for Quality-Aware Inference Serving in the Computing Continuum

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

A Kubernetes-Native Request Router for Quality-Aware Inference Serving in the Computing Continuum Ignjat Karanovic1 , Pantelis A. Frangoudis1 , Ivan Čilić2 , Ivana Podnar Žarko2 , and Schahram Dustdar1 1

arXiv:2609.20497v1 [cs.DC] 17 Sep 2026

2

Distributed Systems Group, TU Wien, Austria Faculty of Electrical Engineering and Computing, University of Zagreb, Croatia

Abstract

1

We introduce Adaptive Score-based Routing Balancer (ASRB), a dynamic, score-based request routing mechanism for Kubernetes-based service deployments over the computing continuum. ASRB jointly considers infrastructure-level information, response time measurements, and application-level quality indicators, with a particular focus on serving Machine Learning (ML) workloads. For these workloads, ASRB balances requests over service instances deployed in the continuum, following service provider-defined policies encoded as weighted combinations of QoS criteria to flexibly address latency-accuracy trade-offs. To drive routing decisions and swiftly adapt to changes in the operating environment, ASRB monitors a range of runtime metrics across multiple system layers. To deal with the associated monitoring overhead, particularly important for large-scale deployments, it selectively and adaptively controls monitoring intensity without sacrificing on routing quality. ASRB is implemented without requiring any modifications to Kubernetes, making it straightforward to deploy and operate in existing cluster environments. Our testbed experiments demonstrate the versatility of ASRB: When tuned for latency reduction, it achieves at least 10 ms lower mean response time compared with latencyoriented state-of-the-art routing mechanisms, while it achieves higher accuracy when this is prioritized through specific configurations, thus enabling flexible and operator-controllable trade-offs. At the same time, it attains reduced failure rates, higher responsiveness to changes in the operating environment, and up to ∼70% less monitoring cost than relevant stateof-the-art solutions, at the potential expense of only a modest latency penalty in some configurations.

Modern IoT services are becoming increasingly datadriven, integrating Artificial Intelligence/Machine Learning (AI/ML) workflows as core components, and are characterized by multi-dimensional and often conflicting performance requirements. On the one hand, domains such as augmented reality [1], autonomous mobility [2], and real-time analytics [3] require consistently low response times to remain usable. The strive for bounded latency drives application deployment towards the edge of the computing continuum [4], as cloud deployment can be prohibitive from a responsetime perspective. On the other hand, the prediction quality of the deployed ML models, typically expressed in terms of accuracy metrics, is crucial in domains like medical diagnostics [5] and safety-critical automation [6].

Introduction

Achieving higher accuracy often incurs additional computational cost, which in turn translates to increased compute-induced latency and energy consumption. For example, deep neural networks for vision tasks, such as high-depth ResNet [7] variants and Vision Transformers [8] deliver stronger prediction performance but require more processing time compared to lightweight architectures like MobileNet [9] or to pruned and quantized model variants that reduce computational demand at the cost of accuracy [10]. Cloud nodes may run such heavy models faster, but the additional network delay can offset these benefits. Conversely, edge nodes typically offer lower network latency but may either host lightweight models with reduced accuracy or, if capacity allows, run more complex models with increased computation latency due to resource limitations, thereby offsetting proximity benefits. 1

This reveals a trade-off between latency and predic- proxy, which distributes requests using basic roundtion quality, which is important for service providers. robin or random strategies [11]. Solutions like serDistributed service deployment over the computing vice meshes (e.g., Istio2 ) offer customizable routcontinuum makes addressing this trade-off more chal- ing rules and observability, but often introduce oplenging for the following reasons: (i) Compute nodes erational overhead, making them less practical for with diverse capabilities (e.g., resource-constrained lightweight or resource-constrained distributed edge edge devices vs. cloud servers) may be hosting repli- deployments [12]. QEdgeProxy [13], upon which this cated service instances; this implies inconsistent re- work builds, maintains per-service dynamic pools of sponse times across these instances, while resource QoS-satisfying service instances which are candidates limitations can translate to service placement con- for routing, while proxy-mity [14] routes based on a straints. (ii) Similarly, different service instances come weighted combination of proximity (latency) and load with end-to-end network paths of varying distance and balancing criteria. ProxyDWRR [15], on the other thus latency. (iii) The volatile operating environment hand, treats CPU load as the sole decision factor. of the continuum requires intensive monitoring to These works integrate well with K8s, but do not ackeep a precise view of service and infrastructure state, count for ML model-specific QoS attributes in routing which is critical input for orchestration decisions. decisions. We address this particular trade-off in the context In contrast, ASRB unifies latency, model accuracy, of inference serving over the continuum, focusing on the following orchestration problem: Given a number and resource metrics in its routing strategy, enabling of service instances deployed over the continuum, each richer QoS trade-offs. Unlike centralized ProxyDserving a given task (e.g., image classification) via WRR scheduling, ASRB uses decentralized proxies, an appropriate (but potentially different) ML model, whose routing decisions introduce negligible overhead what is the optimal service instance to handle a service on the “hot path” (i.e., per request). This contrasts invocation, considering response time and prediction approaches based on (Deep) Reinforcement Learning [16, 17, 18], which, state management and training quality objectives? overheads aside, may require neural network execuWe make the following contributions: 1 We present tion per request. Nevertheless, ideas from RL, such Adaptive Score-based Routing Balancer (ASRB), a reas adapting parameters based on long-term QoS feedquest routing mechanism for ML inference serving, opback, could complement ASRB. We have applied such erating within a decentralized proxy architecture (§ 3) concepts in another line of work [19], albeit with pure and allowing service providers to flexibly specify quallatency orientation. ity objectives within the latency-accuracy trade-off space (§ 4). ASRB natively integrates with KuberNotably, latency and accuracy trade-offs are adnetes (K8s) and is open-source.1 2 We devise an dressed in model selection systems. MDInference [20] adaptive monitoring scheme (§ 5) that dynamically hosts a fast low-accuracy model on-device and a set tunes monitoring intensity, considering the current fit of high-accuracy models in the cloud, and duplicates of different nodes and their hosted service instances to requests to ensure bounded response time, at the attain given QoS goals, thus drastically reducing mon- expense of energy due to duplication. RAMSIS [21] itoring cost without significantly impacting ASRB’s exploits periods of low load to direct requests to higher routing performance. 3 We demonstrate the flexibil- accuracy models. However, it requires a central conity of ASRB in addressing the latency-accuracy trade- troller over which requests are routed. Jellyfish [22] off at reduced monitoring cost by evaluating it over a jointly instruments clients to adapt the size of their K8s-based testbed in a computer vision task (§ 6). input data and adaptively maps clients to ML model

2

instances. Other differences aside, these works do not address Kubernetes integration nor deal with monitoring cost reduction.

Related work

In Kubernetes-based environments, load balancing is typically handled by the built-in component kube1

2

https://github.com/Ignjat96/asrb-proxy

2

https://istio.io/latest/docs/

3

System Design

(i) scoring of service pods using a scalar metric that jointly considers the latency and application-specific 3.1 Architectural Elements and Deploy- quality (in our case, ML model accuracy) each pod ment Model offers, (ii) dynamic pod discovery and health checking, (iii) instrumenting the adaptive monitoring of nodes, The deployment model we assume involves clients, and (iv) fallback request routing when the initially presuch as IoT devices, user applications, or other coferred target is unavailable or does not satisfy current existing services, directing service invocation requests QoS constraints. It interfaces with the K8s control (typically, HTTP calls to application API endpoints, plane via the K8s API to retrieve lists of eligible pods such as to an image classification service) to routing and node/pod metadata, and updates internal caches proxies. Each proxy independently decides on a pervia informers, i.e., a K8s mechanism that allows to request basis on the most appropriate service instance watch for resource updates. Pod discovery, score comto forward the request to, out of a number of candidate putation, and monitoring are executed asynchronously. instances deployed over a cluster of compute nodes Notably, ASRB dynamically assigns quality attributes managed by Kubernetes, the de facto framework for to nodes via K8s labels: each node receives a continuorchestrating containerized applications. Nodes may ously updated label that characterizes its suitability span cloud VMs and edge devices. Our approach folfor serving requests. lows the QEdgeProxy design introduced in our prior ASRB tracks both work [13]. We adopt the way it integrates with K8s Monitoring Subsystem. resource-oriented metrics (CPU/memory use from K8s and build on its code base to implement ASRB. However, we depart from its latency-centric design by Metrics API) and latency-related data (by recording supporting multi-dimensional, ML-oriented QoS ob- the latency of HTTP probes and observing response jectives, and reduce its monitoring cost. Figure 1 times), introducing a low-overhead adaptive monitorprovides an overview of our system design and func- ing strategy (§ 5). Monitoring information is made available through a metrics endpoint, which is contionality. Service Pods. We focus on ML inference serving, sumed by Prometheus [23], a widely used monitoring where ML models for a specific task (e.g., image classi- system in K8s-based environments. fication) are deployed as containers over a K8s cluster. These service instances (pods, in K8s terminology) may run different models that are functionally equivalent, i.e., capable of handling the same task and accepting the same input over a unified interface, but 3.2 Node Labeling Mechanism exhibiting different performance characteristics and resource requirements. Information about these char- We use Kubernetes labels, i.e., lightweight key–value acteristics (e.g., expected ML prediction accuracy) is metadata that we attach to each node and update encoded in instance metadata. Note that instance based on observations made by proxies. A node may placement decisions are beyond the responsibilities of be labeled fit, unfit, or overloaded, where fit and unfit our scheme. reflect the aggregated score derived from latency and ASRB Node-Local Routing Proxy. The ASRB model accuracy, and overloaded indicates excessive proxy is deployed on each node as a DaemonSet, a CPU or memory utilization. Labels are assigned by Kubernetes construct ensuring that a copy of a pod proxies independently via calls to the K8s API and runs on all (or select) nodes in the cluster. This de- are maintained at the Kubernetes control plane level, ployment choice enables node-local request handling allowing proxies to indirectly share node state (e.g., and avoids a centralized routing bottleneck. Different an observation that a node is overloaded) and guide deployment options are possible: the service provider subsequent routing and monitoring decisions, as we can bundle a distinct proxy with each service so that will describe in § 4 and § 5, respectively. We note multiple proxies may co-exist on a single node, have that while maintaining labels has reasonable memory one proxy per node responsible for multiple applica- cost even for large numbers of nodes, the effects of tions, or deploy proxies at specific edge nodes. The centrally updating this state via the K8s API at scale ASRB proxy is responsible for the following tasks: require further study. 3

Figure 1: System architecture and operatations.

4

Request Routing Logic

for pod p, and Ap is a static model-specific accuracy value. Latency is transformed into a normalized 4.1 Scoring function design latency score in [0, 1]. The cap at 1000 ms bounds the influence of outliers. Notably, it is up to serASRB routes requests based on a weighted scoring vice providers to configure λ and thus drive request mechanism that combines latency and model accuracy routing towards the appropriate trade-off point, capinto a unified QoS metric. For each pod, the proxy turing service provider QoS priorities that may vary collects the most recent latency measurement and for different applications. For instance, in interacretrieves the model’s accuracy annotation.3 The baltive user-facing applications, a higher λ prioritizes ance between these two factors is determined by the responsiveness, while in classification-critical systems, latency importance weight parameter λ ∈ [0, 1], which a lower λ increases the weight of model performance is supplied with each request. The scoring function is in the selection process. ASRB does not prescribe therefore defined as: how λ is selected. For example, service providers can    observe Service Level Objective (SLO) fulfillment perLp Sp = λ · 1 − min ,1 + (1 − λ) · Ap (1) formance, both historically and online, and use it as 1000 feedback to learn to dynamically adjust λ. We also where λ controls how strongly routing favors low la- remark that other quality objectives are straightfortency, Lp is the most recent latency measurement ward to support. For example, tracking failed requests per instance can be used to maintain a respective in3 ML model accuracy information is assumed to be known a stance availability metric, and weighted decisions that priori and available externally (e.g., by the application provider) as a metadata attribute, i.e., it is not measured at runtime by also factor in service availability could be made in an ASRB, contrary to latency. However, if the use case allows, identical way. the service provider may implement mechanisms to monitor accuracy and update these metadata at runtime.

4

4.2

Score-Based Routing Algorithm

Algorithm 1 ASRB Score-Based Request Routing Algorithm Input: Incoming request r, latency weight λ Output: Selected pod p∗ (1) Service Discovery: Retrieve all Ready pods for service r. (2) Pod Categorization: a) collect recent latency Lp from cache, b) collect score Sp from cache, c) classify pods into: bestPods: healthy and not overloaded, and overloadedPods: healthy but above resource limits. (3) Latency Refresh (optional): If too few pods have valid latency data, the proxy triggers an asynchronous latency-refresh routine to update stale or missing latency information, as summarized in Algorithm 2. (4) Score Computation: For each candidate pod p: Sp = λ · (1 − min(Lp /1000, 1)) + (1 − λ) · Ap where Ap is the model accuracy from annotations. (5) Ranking: Select the highest-scoring pod within: a) bestPods b) else overloadedPods c) else random from all healthy pods (6) Forwarding: Forward request r to the selected pod p∗ . (7) Feedback Update: a) measure RTT, b) update latency estimate and pod score, c) update node QoS label (fit/unfit). Note: The overloaded label is not set in this feedback step; it is maintained separately by the resource-monitoring component based on CPU and memory thresholds.

A proxy executes a lightweight routing algorithm for each incoming request, which selects the most suitable service instance. The routing process is summarized in Algorithm 1. Specifically, ASRB retrieves all pods for the requested service from the K8s API and retains only those in the Ready state. For each ready pod, it determines the label of its hosting node (fit, overloaded, or unfit), and obtains its most recent latency estimate (from the local cache or via latency approximation if the value is stale), as well as its model accuracy annotation. It proceeds by computing the QoS score for each pod using Eq. (1), extracting the value of λ from the request. At this point, ASRB selects pods running on nodes labeled fit and discards those whose score falls below a configured minimum threshold. If no suitable fit-node pods are available,4 it evaluates pods on overloaded nodes. If neither fit nor overloaded nodes provide a pod with an acceptable score, ASRB may optionally fall back to pods on unfit nodes or revert to a simple strategy such as random selection among healthy pods. Finally, it forwards the request to the highest-scoring pod from the final candidate set. Algorithm 1 performs a fixed number of constanttime operations per pod hosting the requested service (cache lookup and score computation take O(1) time and ranking evaluates each pod at most once), and therefore routes a request in time linear in the number of pods discovered. After a response is received, ASRB updates the latency measurement for the selected pod’s host, recomputes the best pod score per node, and uses this score and resource usage to update the selected node’s Algorithm 2 Asynchronous latency refresh routine label for future routing and monitoring decisions. Input: Set of candidate pods P Output: Updated latency cache entries and recom5 Adaptive Monitoring puted scores for all p ∈ P do 5.1 Design Principles if latency for p is missing or outdated then ASRB needs to keep an up-to-date view of the infrassend lightweight request to p at /echo tructure and service state, to make quality-informed measure round-trip time Lp routing decisions. Aggressive monitoring helps mainupdate latency cache for p with Lp tain accurate state, which is important for high-quality recompute score Sp using Lp & cached accurouting decisions, but the overhead can be significant, racy Ap store updated Sp in the score cache 4 It is also possible that a high-ranking pod fails to meet a end if constraint. This can happen if, e.g., ASRB is configured to end for prioritize for accuracy, but at the same time there is a stringent response time constraint that should be met.

5

particularly for large-scale deployments widely distributed over the continuum. To address this issue, our monitoring mechanism design is driven by the following intuition: On the one hand, static, fixedinterval monitoring may waste resources by probing nodes unnecessarily when conditions are stable, or miss important changes when the update frequency is too low under volatile conditions. On the other hand, there are nodes that, due to physical distance or the fact that they host low-quality ML models, may fail to meet the latency or accuracy QoS thresholds put in place by the service provider; such nodes are likely to be filtered out in routing decisions and become less relevant for intense monitoring. The system can therefore allocate more observation effort to uncertain or critical nodes, and reduce monitoring for consistently stable or obviously unsuitable ones. This boils down to the following monitoring principles:

changes. Latency approximation is regulated by a cooldown mechanism. After an approximation cycle is triggered for a service, further approximation attempts are suppressed for a configurable interval to avoid excessive probing. During this interval, cached latency estimates are reused. Resource usage metrics, including indicators such as CPU utilization, memory usage, and pod health status, are exposed by K8s and are obtained through the Metrics API or via Prometheus. They are refreshed periodically and used to label nodes as overloaded when predefined thresholds are exceeded. These labels are used in two ways: (i) they steer the monitoring process by refreshing overloaded nodes more frequently than fit nodes, which are in turn refreshed more frequently than unfit ones, and (ii) they support routing by treating pods on overloaded nodes as a fallback option when no suitable pods on fit nodes are available.

• Fit or overloaded nodes receive higher monitoring frequency, as they are more likely to be selected for routing.

6

Evaluation

• Monitoring frequency is increased for nodes with 6.1 Experimental Setup volatile scores or borderline performance (e.g., latency spikes, occasional failures). We evaluate our scheme against mechanisms from the state of the art on a k3s-based cluster, serving an • Nodes with consistent performance over a defined image classification task with a number of pods hostwindow are monitored less frequently to conserve ing either the MobileNet-V2 (faster, lower accuracy) resources. or the RestNet-50 model (slower, higher accuracy). We generate a stream of 1200 service requests to the • Nodes labeled unfit are still refreshed periodically, but at the lowest frequency, allowing the exposed API endpoint, and the experiment consists system to detect recovery without generating un- of six phases (T1–T6) where the setup undergoes controlled topology changes, overload events, and renecessary overhead. covery periods. Our cluster forms a representative (simulated) far edge (IoT device space) – near edge 5.2 Adaptive Active-Passive Monitoring (telco/MEC data center) – cloud topology. The folStrategy lowing compute nodes, each running as a separate ASRB monitors latency through two mechanisms: pas- Ubuntu 22.04 virtual machine in a local data center, sive, where feedback updates are collected after each are deployed over these three tiers: (Master node) request, and active, which initiates on-demand approx- A cloud-tier controller running the K8s control plane imation when cached values become stale. In this case, and hosting inference pods executing both MobileNet the proxy issues a lightweight request to the pod’s and ResNet models, reachable from the IoT device /echo endpoint to estimate the current round-trip (request source) with a 150 ms RTT; (Worker 1 and time. These active probes are triggered adaptively Worker 2) stable near-edge K8s nodes reachable at and only when required, rather than on every re- a 50 ms RTT from the request source, hosting both quest. When a real latency measurement becomes MobileNet and ResNet inference pods; (Worker 3) available after a period of approximation, the cached on-premise far-edge K8s node, reachable at a 5 ms latency value is updated to maintain an exponentially RTT; initially without a pod (T1), receives a Moweighted moving average. This prevents abrupt score bileNet pod in T4, and becomes overloaded in T5; 6

(Worker 4) far-edge K8s node reachable at a 10 ms RTT, hosting only a MobileNet model; added in T2, rebooted in T3, and overloaded in T5. Latencies are controlled using the Linux tc utility, and each node runs a request router instance capable of executing our candidate routing strategies. The six experiment phases are as follows: • T1 – Initial topology: Workers 1–3 active; Worker 3 has no model pod. • T2 – Node arrival: Worker 4 joins and becomes Figure 2: Latency distribution across all routing algorithms. available. • T3 – Reboot and recovery: Worker 4 becomes unavailable, then recovers. or recovery gating. Consequently, it reacts only to instantaneous latency values and does not handle • T4 – Pod deployment: Worker 3 receives a transient degradation or pod life-cycle events. QEP MobileNet pod. incorporates latency feedback and resource checks, • T5 – Overload: Workers 3–4 are overloaded: but applies them in an instantaneous manner. Overa temporary pod is deployed on each, executing load decisions are based on momentary CPU or memCPU and memory load using the stress util- ory thresholds and are not persisted through explicit ity (stress --cpu 4 --vm 2 --vm-bytes 1G). node-state labels or cooldown periods. In addition, This pushes node CPU utilization above 85%, pods may be treated as routable before networking information (e.g., PodIP, HostIP) is fully initialized, triggering the overload condition. which leads to failed requests during node joins or • T6 – Recovery: Overload removed. reboots. ASRB (λ = 1) preserves a latency-centric objective but stabilizes routing through explicit node The following routing strategies are evaluated: (i) labeling, cooldown logic, and stricter pod admission ASRB with λ ∈ {1, 0.5, 0}, and under two differ- semantics. Under dynamic conditions, these mechaent monitoring configurations; (ii) QEdgeProxy nisms reduce latency variance and extreme outliers (QEP) [13] with its latency-only routing and static by preventing routing to temporarily degraded or notmonitoring configuration; (iii) two different configu- yet-ready nodes. While the overall latency differences rations of proxy-mity [14], corresponding to pure between the latency-oriented approaches remain small, proximity-based routing (α = 1) and a version that they are statistically significant (see Table 1). trades latency for more even load balancing (α = 0.8); At the same time, ASRB allows to flexibly address (iv) Round Robin. latency-accuracy trade-offs: For example, when con6.2 Response Time vs. ML Accuracy Per- figured with λ = 0.5, while showing a noticeably higher median latency (440 ms) and increased variabilformance ity compared to a pure latency-oriented configuration, Figure 2 presents response time statistics for the dif- it routes 85.4% of the requests to high-accuracy inferent request routing schemes over all phases of the stances (vs. 7.5% for λ = 1, 7.1% for proxy-mity, experiment. ASRB with λ = 1 achieves not only the and 2.6% for QEP), selecting ResNet50 preferentially, lowest median latency but also a small interquartile while still routing to MobileNet during latency spikes range, indicating a highly stable latency profile. or transient overloads. At the other end of the conAlthough ASRB (λ = 1), QEP, and proxy-mity are figuration space, ASRB (λ = 0) exhibits a markedly all latency-oriented, they differ in how latency-based higher median latency and an expanded upper tail, decisions are maintained under dynamic conditions. reflecting the slower inference time of the ResNet50 Proxy-mity relies on network latency ranking with- model, which it selects almost exclusively for higher out explicit node-state tracking, overload persistence, accuracy. 7

Table 1: Mean pairwise response time difference (∆) vs. baselines in latency-oriented configurations. Negative values indicate lower response time than the baseline. The null hypothesis H0 : ∆ = 0 (no latency difference) is rejected at the 95% level: Phase-stratified bootstrap CIs (computed using the BCa method [24] and including 100 000 resamples) exclude 0; we also report the results of one-sided t-tests for the directional alternative hypothesis H1 : ∆ < 0. Baseline ∆ (ms) 95% Bootstrap (BCa) CI t p proxy-mity -9.79 [-18.76,-0.83] -2.07 .0193 QEP -18.64 [-26.62,-9.98] -4.22 <.0001 observed for proxy-mity, while Round-Robin shows high variance even in stable phases due to its static nature. It should finally be noted that during the whole duration of the experiment, ASRB experienced only 2 transient failures (vs. 36 for QEP).

6.4

Monitoring Cost Reduction

We compare the monitoring cost of ASRB (λ = 1) and QEP using the following metrics: (i) number of Figure 3: Latency evolution across T1–T6 for ASRB monitoring API calls to the K8s control plane, and (λ = 1). (ii) volume of the associated monitoring traffic. QEP polls for information about all nodes and instances 6.3 Responsiveness to Environment Dy- at fixed intervals, whereas ASRB controls polling frequency adaptively, resulting in fewer such calls. On namics the flip side, ASRB needs to send PATCH requests Figure 3 shows the request-by-request latency evo- to the K8s control plane to update pod scores and lution over the entire 1200-request experiment for node labels, which add to its monitoring cost. We ASRB (λ = 1). Compared with the baselines (fig- evaluate different monitoring configurations, paramures omitted in the interest of space), this ASRB eterized by the following: (i) node state monitoring configuration achieves the most stable profile across interval (node metrics cache time), and (ii) node label all phases. During T1, it consistently chooses the update interval. For ASRB, (i) represents the monfastest nodes (Worker 3 or Worker 1), resulting in low itoring period for overloaded nodes; fit/unfit nodes latency with only minor spikes caused by transient are monitored at 2× / 3× that interval. For QEP, all network fluctuations. When Worker 4 joins in T2, nodes are monitored at this interval. Note that (ii) ASRB incorporates it almost immediately, reflected is only relevant for ASRB, and a PATCH API call by a small dip in average latency. The reboot event to update a node label takes place only if the label in T3 is also handled quickly: the latency spikes are would change. Despite a potential increase in API short, and the system reroutes within a few requests. calls, ASRB significantly reduces the volume of moniIn T5, ASRB detects the overload on Workers 3–4 toring traffic, as Table 2 shows. This is because the early and shifts traffic toward stable nodes, preventing monitoring overhead of the baseline is dominated by the long clusters of high latency experienced by the the much larger payload size per call: QEP repeatedly queries all workers regardless of their relevance baselines. T6 returns to a clean and stable profile. QEP shows similar trends in stable phases but is or state. slower to detect the rebooted node in T3, while it conWhile the monitoring overhead in absolute terms tinues sending requests to degraded nodes for notice- is limited due to the small size of the testbed, the ably longer period than ASRB in T5. This confirms achieved savings (29%-73% less monitoring traffic than that ASRB’s adaptive monitoring is more suitable for QEP for the same metric update interval) will be dynamic edge environments. Similar behavior was significant in large-scale deployments. At the same 8

time, such savings come with only a modest latency penalty in some configurations. For example, with metric and label update intervals of 15 s and 60 s, respectively, ASRB saved 73% of monitoring traffic at the expense of only 5.3% higher mean response time than QEP, while in its most aggressive monitoring configuration, ASRB’s mean response time was 6.7% lower than that of QEP, still with 69% less monitoring traffic.

7

autonomous driving systems: A survey. IEEE Access, 10:14076–14119, 2022. [3] Weisi Chen, Zoran Milosevic, Fethi A. Rabhi, and Andrew Berry. Real-time analytics: Concepts, architectures, and ML/AI considerations. IEEE Access, 11:71634–71657, 2023. [4] Schahram Dustdar, Vı́ctor Casamayor-Pujol, and Praveen Kumar Donta. On distributed computing continuum systems. IEEE Trans. Knowl. Data Eng., 35(4):4092–4105, 2023.

Conclusion

[5] Steven A Hicks, Inga Strümke, Vajira Thambawita, Malek Hammou, Michael A Riegler, Pål Halvorsen, and Sravanthi Parasa. On evaluation metrics for medical applications of artificial intelligence. Scientific reports, 12(1):5979, 2022.

We presented ASRB, a decentralized request routing scheme tailored to inference serving, capable of balancing workloads while considering multi-dimensional QoS objectives—particularly latency and prediction quality—in a flexible way. ASRB integrates natively with Kubernetes ecosystems and enables IoT service providers to prioritize conflicting quality criteria depending on the specific requirements of applications that rely on ML service instances across the computing continuum. Its decentralized design and reduced monitoring overhead make it suitable for dispatching workloads in large-scale distributed deployments. More sophisticated monitoring adaptations, studying the effects of node state updates in very large deployments, and applying ideas from reinforcement learning to aspects such as dynamically adjusting QoS priorities, are directions for future work.

[6] Jon Pérez-Cerrolaza, Jaume Abella, Markus Borg, Carlo Donzella, Jesús Cerquides, Francisco J. Cazorla, Cristofer Englund, Markus Tauber, George Nikolakopoulos, and Jose Luis Flores. Artificial intelligence for safety-critical systems in industrial and transportation domains: A survey. ACM Comput. Surv., 56(7):176:1– 176:40, 2024. [7] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. IEEE CVPR, 2016. [8] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In Proc. ICLR, 2021.

Acknowledgment This work has been supported in part by the European Union’s Horizon Europe research and innovation programme under grant agreement No. 101079214 (AIoTwin) and by the European Regional Development Fund (grant No. KK.01.1.1.04.0108, project IoT-Field).

[9] Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proc. IEEE CVPR, 2018.

References

[1] Ayman Younis, Brian Qiu, and Dario Pompili. Latency-aware hybrid edge cloud framework for [10] Song Han, Huizi Mao, and William J. Dally. Deep compression: Compressing deep neural networks mobile augmented reality applications. In Proc. with pruning, trained quantization and huffman IEEE SECON, 2020. coding. In Proc. ICLR, 2016. [2] Tolga Turay and Tanya Vladimirova. Toward performing image classification and object de- [11] Cloud Native Computing Foundation. Kubertection with convolutional neural networks in netes. https://kubernetes.io, 2023. 9

Table 2: Monitoring overhead for different configurations. We also report percentage traffic savings and the mean pairwise response time difference (∆) vs. the respective QEP configuration. For configuration 15 s/60 s, the null hypothesis of ASRB being inferior (slower) than QEP by a margin higher than 10% is rejected at the 95% level: The phase-stratified bootstrap CI ([−23.29, −2.80]) excludes 0 and a one-sided t-test yields t = −2.38, p = .0086. Similarly, the non-inferiority margin for configuration 60 s/60 s is 3%. Across the board, ASRB reduces monitoring traffic by 29%-73%. Update intervals API calls Traffic (MB) ∆ (ms) Metrics Labels 15 s – 136 3.29 QEP 60 s – 40 1.07 15 s 5s 85 1.01 (-69%) -18.64 (-6.7%) 15 s 60 s 73 0.89 (-73%) +14.94 (+5.3%) ASRB 60 s 5s 81 0.76 (-29%) -16.62 (-5.6%) 60 s 60 s 72 0.63 (-41%) -1.26 (-0.4%) [12] Xiangfeng Zhu, Guozhen She, Bowen Xue, [18] Ming Tang and Vincent WS Wong. Deep reYu Zhang, Yongsu Zhang, Xuan Kelvin Zou, inforcement learning for task offloading in moXiongchun Duan, Peng He, Arvind Krishnabile edge computing systems. IEEE Trans. Mob. murthy, Matthew Lentz, Danyang Zhuo, and Comput., 21(6):1985–1997, 2020. Ratul Mahajan. Dissecting overheads of service [19] Ivan Cilic, Ivana Podnar Zarko, Pantelis A. Franmesh sidecars. In Proc. ACM SoCC, 2023. goudis, and Schahram Dustdar. Qos-aware load balancing in the computing continuum via multi[13] Ivan Cilic, Valentin Jukanovic, Ivana Podnar player bandits. IEEE Trans. Serv. Comput., 2026. Zarko, Pantelis A. Frangoudis, and Schahram In press. Dustdar. Qedgeproxy: QoS-aware load balancing for IoT services in the computing continuum. In [20] Samuel S. Ogden and Tian Guo. MDINFERProc. IEEE EDGE, 2024. ENCE: balancing inference accuracy and latency for mobile applications. In Proc. IEEE IC2E, [14] Ali J. Fahs and Guillaume Pierre. Proximity2020. aware traffic routing in distributed fog computing platforms. In Proc. IEEE/ACM CCGrid, 2019. [21] Daniel Mendoza, Francisco Romero, and Caroline Trippel. Model selection for latency-critical [15] Qingkun Wang, Yi Ren, Saqing Yang, Jianbo inference serving. In Proc. ACM EuroSys, 2024. Guan, Bao Li, Jianfeng Zhang, and Yusong Tan. Proxydwrr: a dynamic load balancing approach [22] Vinod Nigade, Pablo Bauszat, Henri E. Bal, and Lin Wang. Jellyfish: Timely inference serving for for heterogeneous-cpu kubernetes clusters. In dynamic edge networks. In Proc. IEEE RTSS, Proc. IEEE JCC, 2022. 2022. [16] José Santos, Tim Wauters, Filip De Turck, and Peter Steenkiste. Towards optimal load balanc- [23] Björn Rabenstein and Julius Volz. Prometheus: A Next-Generation monitoring system (talk). In ing in multi-zone kubernetes clusters via reinUSENIX SREcon15 Europe, 2015. forcement learning. In Proc. 33rd International Conference on Computer Communications and [24] Bradley Efron and Trevor Hastie. Computer Networks (ICCCN), 2024. Age Statistical Inference: Algorithms, Evidence, [17] Vasileios Karagiannis, Pantelis A. Frangoudis, Schahram Dustdar, and Stefan Schulte. Contextaware routing in fog computing systems. IEEE Trans. Cloud Comput., 11(1):532–549, 2023. 10

and Data Science. Cambridge University Press, Cambridge, 2016.

Record · ID 978389 · SHA-256 c7c1eacd1b43f0b6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.