Conceptio › Archive › arXiv CS
arXiv CSopen access

Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

arXiv:2609.04513v1 [cs.DC] 3 Sep 2026

Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters Milos Gravara

Andrija Stanisic

Stefan Nastic

Distributed Systems Group TU Wien [email protected]

Distributed Systems Group TU Wien [email protected]

Distributed Systems Group TU Wien [email protected]

selects the model variant, runtime parameters and hardware placement of each workflow stage to satisfy given Service Level Objectives (SLOs), which typically include latency, throughput, and cost constraints [13]–[16]. As production workflows grow in the number of stages, these choices create a combinatorial space of execution plans [17], [18]. Each plan occupies a different point in the accuracy-performance-cost space, which often compete [19]–[23]. This means that, for example, selecting larger AI models per stage may improve overall accuracy, but typically increases latency and cost. On the other hand, plans including cheaper or faster AI models may satisfy SLOs, but often at the expense of workflow accuracy. Therefore, deployment optimization amounts to jointly selecting model variants per stage and mapping them onto heterogeneous hardware, managing the resulting trade-offs to find plans that maximize accuracy among those feasible under the given SLOs. A common formulation of this optimization problem requires estimating accuracy and system behavior for each candidate execution plan before selection [13], [14], [23], [24]. Latency, throughput, and cost are mainly tractable in this setting because they can often be profiled for individual stage variants on target hardware and composed according to the workflow topology [13], [14], [22], [24]. Yet, such process cannot be applied to estimate workflow accuracy. As output of one stage becomes the input to downstream stages, errors and information loss introduced upstream can change the accuracy distribution of later stages [3], [17], [18]. Workflow-level accuracy therefore depends on how quality propagates through the execution plan, inducing a challenge in estimating plan accuracy before I. I NTRODUCTION deployment. The field of Artificial Intelligence (AI) is shifting from deExisting approaches commonly estimate workflow-level ploying monolithic AI models towards Compound AI systems. accuracy in one of two ways. End-to-end profiling evaluates Compound AI represents a distributed intelligence approach complete execution plans directly and provides faithful meacombining multiple specialized AI models with software com- surements for plan selection [3], [14], [22], [23]. This captures ponents into workflows, where each stage represents a single interactions between stages, but each added variant, parameter, model or component invocation, orchestrated to solve various or stage requires additional complete workflow evaluations, AI tasks [1]–[6]. This approach offers practical advantages making exhaustive profiling intractable as workflows grow in for reliability, scalability, and efficiency, enabling control over depth and variant count. To reduce profiling cost, recent work model outputs, component-specific adjustments, and adaptation has constructed surrogate accuracy models by composing perto changing conditions [7]–[12]. stage accuracy estimates, often as products of individual stage These advantages are not obtained by the workflow structure accuracies [13], [18]. Such surrogate models are efficient, but alone. A Compound AI workflow must be instantiated as an they treat stage contributions as largely independent and thereexecution plan before deployment [3], [13], [14]. Such a plan fore do not capture how upstream errors or information loss

Abstract—Compound AI workflows are increasingly used to serve complex AI tasks by coordinating multiple AI models and software components. This approach enables deployment flexibility, as each workflow stage can expose different model variants and resource requirements, but it also expands the deployment choices. A deployment must choose an execution plan that selects AI models for each compound AI workflow stage and places them on a heterogeneous cluster in order to satisfy SLOs. Deployment optimizers therefore need estimates to compare many candidate plans and identify feasible ones. System metrics can often be profiled per stage and composed according to workflow topology, but accuracy cannot, as errors and information loss at upstream stages affect the accuracy of downstream stages. Existing approaches either profile complete configurations end to end, which scales poorly, or use product-based accuracy surrogates that treat stages as independent and can misrank candidate plans. We introduce Atlas, a framework for optimizing compound AI deployments under SLO constraints. Atlas uses MAP, a Markovian Accuracy Predictor, to estimate configuration accuracy from local conditional accuracy transitions between adjacent workflow stages. MAP discretizes intermediate outputs into accuracy buckets and composes transition profiles according to workflow topology, giving the optimizer an accuracy estimate without exhaustive end-to-end profiling. Atlas formulates execution-plan selection as a mixedinteger linear program that maximizes predicted accuracy subject to SLOs. Across four compound AI workflows, MAP achieves Spearman correlation up to 0.947 while reducing profiling cost by up to 2.6× relative to exhaustive end-to-end profiling. Guided by MAP, the Atlas optimizer selects execution plans within 0.03 of oracle accuracy while reducing deployment cost by up to 42% through heterogeneous placement. Index Terms—Compound AI, Model Selection, Deployment Optimization, Distributed Inference

affect downstream behavior. As a result, existing methods either The remainder of this paper is organized as follows. Secpreserve interaction fidelity at high profiling cost or reduce cost tion II motivates the accuracy-estimation problem. Section III through assumptions that weaken accuracy estimation across presents the Atlas framework and system model. Section IV multi-stage workflows. introduces MAP. Section V formulates the MILP. Section VI In this paper, we introduce Atlas, a framework that optimizes evaluates Atlas, Section VII reviews related work, Section VIII model selection and hardware mapping for Compound AI discusses design choices and Section IX concludes. workflows under SLO constraints. Atlas selects execution plans II. M OTIVATION that maximize predicted task accuracy while satisfying latency, throughput, memory, and cost requirements. To estimate This section motivates the accuracy-estimation problem in workflow accuracy without exhaustive end-to-end profiling, Compound AI deployment optimization. We first illustrate Atlas profiles how model choices at one workflow stage how workflow configurations create a large space of execution affect the output quality of the next. The Markovian Accuracy plans where each plan induces different accuracy, latency, and Predictor (MAP) then discretizes these quality signals into cost trade-offs. We then examine two existing approaches to buckets and composes the resulting transition profiles according estimating configuration accuracy and their limitations. to the workflow topology, predicting end-to-end accuracy for any candidate configuration. The Atlas optimizer combines A. Compound AI Deployment Optimization these accuracy predictions with system performance profiles in To illustrate the deployment optimization problem, we a mixed-integer linear program (MILP), producing an execution consider three compound AI workflows of increasing complan that specifies model selection and hardware mapping for plexity, shown in Figure 1. The RAG pipeline (a) retrieves each workflow stage. documents, reranks them, and feeds the result to a language The main contributions of this work are: model that generates an answer [25]–[27]. Each stage exposes model variants and hyperparameters: three retrieval depths, two • Atlas - A novel framework for deployment optimization of Compound AI workflows on heterogeneous clusters. reranker models at two reranking depths, and six generator Atlas takes a workflow specification, candidate model models. A workflow configuration assigns one model variant to variants, a calibration dataset, and SLOs as input, and each stage, producing 72 configurations. Not all queries require produces an execution plan specifying model variant a large generator, so a router (b) can dispatch inputs to a small selection and hardware placement per workflow stage. or large generator based on estimated difficulty, expanding the It separates accuracy estimation from system profiling, space to 288 configurations [8], [28]–[32]. As smaller models keeping both tractable, and supports linear, routed, loop, typically produce lower-quality answers, a feedback loop (c) can trigger critic-driven revision, expanding the space to 1728 and composed workflow topologies. configurations [18], [33], [34]. Each added stage multiplies • MAP - A Markovian Accuracy Predictor that estimates configuration accuracy from local conditional quality tran- the number of configurations that the optimizer must evaluate, sitions between adjacent workflow stages. MAP discretizes and different configurations induce different accuracy, latency, intermediate outputs into quality buckets and composes and cost trade-offs. To deploy a selected configuration, an transition profiles according to workflow topology using execution plan maps each selected model onto a cluster tier, three operators: linear pipelines, routed workflows, and adding hardware placement to the variant assignment. For each configuration, the optimizer must estimate how the feedback loops. Across four evaluated workflows, MAP achieves the strongest ranking correlation among evaluated corresponding execution plan will perform before deployment. predictors, with Spearman ρ = 0.947 on RAG, ρ = 0.921 on System metrics are tractable. End-to-end latency, for example, RAG with routing, ρ = 0.783 on RAG with self-refinement, can be composed from per-stage profiles: a configuration with and ρ = 0.882 on the full composed workflow, while MiniLM retriever (12 ms), MS-MARCO reranker (38 ms), and reducing profiling cost by 2.6× on our largest measured Llama-3.1 8B generator (385 ms) yields approximately 435 ms on an RTX 4090. Swapping the generator to Llama-3.2 1B workflow, growing to over 80× in synthetic projection. reduces it to roughly 200 ms. Per-stage latencies can be profiled • Atlas Plan Optimizer - An execution plan optimizer that formulates plan selection as a Mixed-Integer Linear Pro- once per (variant, tier) pair and reused across configurations. gram (MILP) over model variant selection and hardware Cost, throughput, and memory behave similarly. placement, jointly optimizing predicted accuracy subject However, estimating accuracy across configurations requires to latency, throughput, memory, and cost constraints. The a different approach. For example, two RAG configurations predicted accuracy objective is supplied by MAP and differing only in the reranker (MS-MARCO vs. BGE) can evaluated over all candidate configurations before the surface different passages as input for the downstream generator. MILP is invoked, keeping the objective linear. On a The same Llama-3.1 8B generator can answer correctly given homogeneous cluster, Atlas selects execution plans within strong evidence from one reranker and fail given weaker 0.03 of oracle accuracy across all evaluated loads and evidence from the other, even though no generator parameter SLOs, while baseline approaches fall by up to 0.45. On a changed. This means that accuracy at one stage depends not heterogeneous cluster, Atlas matches oracle accuracy at up only on the model selected there but also on upstream selections. to 42% lower deployment cost than single-tier strategies. Workflow accuracy is thus a property of the full configuration

Llama-3 1B

LLM 1

LLM 1 Qwen-2.5 1.5B

Llama-3 1B

Retrieval

Rerank

LLM

Retrieval

Rerank

MiniLM

MS-M, BGE

k = {3, 5, 10}

k = {3, 5}

Router

Qwen-2.5 1.5B

Retrieval

Rerank

Llama3, Gemma

MiniLM

3B, 8B, 12B

k = {3, 5, 10}

MS-M, BGE

DistilBERT

LLM 2

k = {3, 5}

Rout. Threshold

MiniLM k = {3, 5, 10}

MS-M, BGE k = {3, 5}

Refine

Router DistilBERT

Llama-3 8B Gemma-3 12B

a) Retrieval Augmented Generation (RAG)

b) RAG + Model Routing

LLM 2

Critic

Llama-3 8B Gemma-3 12B

Llama-3 8B Gemma-3 12B

Rout. Threshold

c) RAG + Model Routing + Iterative Refinement

Fig. 1. Representative compound AI workflows with illustrative configuration counts.

and cannot be trivially composed from isolated workflow stages. Still, the optimizer must estimate workflow accuracy in order to rank candidate plans and select the best one for deployment. B. End-to-end accuracy profiling The most direct way to estimate configuration accuracy is to profile complete configurations end to end. Each candidate configuration is executed on a representative evaluation dataset and the final workflow output is scored. This captures interactions between stages as they occur in the deployed workflow and provides a faithful reference for plan selection.

Configurations evaluated

W1: RAG

10M 1M

W2: W1 + Router

W3: W2 + Refine

TABLE I PAS- NAIVE RANKING AGAINST MEASURED END - TO - END ACCURACY. ρ

τ

Top-5

Regret

0.295 0.186 0.068

0.203 0.171 0.023

0.00 0.00 0.00

+0.322 +0.720 +0.657

Workflow RAG RAG + Router RAG + Router + Loop

representative of this approach. For a configuration c, PAS assigns one standalone accuracy value to each selected stage variant and combines them multiplicatively, PAS(c) =

Y

a(s, vs ),

(1)

s∈S

where a(s, vs ) denotes the standalone accuracy of variant vs at stage s. Each stage can be measured independently and scores 10K can be reused across configurations, making PAS cheap to 1K compute. However, PAS treats stage accuracy contributions as 100 independent scalar factors, ignoring the dependencies between 10 stages, which could lead to poor accuracy estimates. 2 4 6 8 10 12 Variants per stage To test this assumption, we apply PAS to the three workflows in Figure 1 and compare resulting configuration rankings Fig. 2. Simulated end-to-end profiling cost for the workflows in Figure 1. against measured end-to-end accuracy. Table I reports the Exhaustive end-to-end profiling captures stage interactions results. The results show that ranking quality degrades consisdirectly but requires evaluating every candidate configuration. tently with workflow complexity. Spearman correlation falls Figure 2 shows a simulated profiling cost for the three from 0.295 on the RAG pipeline to 0.068 on the full composed workflows described in Figure 1, measuring the number of workflow, with zero Top-5 overlap across all three workflows. A deployment optimizer relying on PAS can therefore select complete configuration evaluations required as the number of a plan that satisfies constraints but delivers substantially lower variants per stage increases. Even for the simplest linear RAG accuracy than alternatives, as reflected by the regret values in pipeline, the number of configurations that must be evaluated Table I. The question is whether configuration accuracy can grows into the thousands with only a handful of variants per be estimated with enough fidelity to preserve configuration stage. For the composed workflow with a router and feedback rankings, without requiring full end-to-end measurement. loop, the count exceeds one million. Adding a single variant at any stage creates new combinations with every existing variant III. ATLAS F RAMEWORK OVERVIEW in the rest of the workflow, and adding a new stage multiplies Atlas is a framework for deployment optimization of the configuration count entirely. The profiling cost of exhaustive end-to-end measurement therefore becomes intractable as the compound AI workflows on heterogeneous clusters. It takes a configuration space grows, even for workflows of moderate workflow specification, candidate model variants, a calibration dataset, and SLOs as input, and produces an execution plan depth. that assigns one model variant and one hardware placement to C. Product Based Accuracy Estimation each workflow stage. The framework operates in three phases: A natural alternative to exhaustive end-to-end profiling is to profiling, optimization, and execution, as shown in Figure 3. estimate configuration accuracy from per-stage measurements. The candidate model variants are registered in the model The pipeline accuracy score (PAS), used by IPA [13], is store, which represents the ecosystem of models available 100K

Fig. 3. Atlas framework overview: profiling, optimization, and execution phases producing an accuracy-optimized Compound AI execution plan.

for optimization across workflow stages. The SLOs define the for each stage s, the hardware class ts on which the selected operating constraints under which Atlas selects execution plans, variant c(s) runs: covering latency, throughput, memory, and cost requirements.  p = c, (ts )s∈S , ts ∈ T. (2) In the profiling phase, two independent profilers operate in parallel. The accuracy profiler runs each pair of adjacent The optimizer therefore chooses both the workflow configustages on the calibration dataset, measuring how the output ration and the resources used to serve it. quality of an upstream stage shifts the quality distribution For each stage, variant, and hardware class tuple (s, v, t), of its downstream neighbor. This produces per-pair accuracy the system profiler provides tail execution latency ℓ(s, v, t), profiles that keep profiling cost proportional to the number sustained service capacity θ(s, v, t), and memory footprint of stage pairs rather than the number of full configurations. µ(s, v, t). Each hardware class t ∈ T has aggregate memory The performance profiler draws candidate variants from the capacity M and hourly worker cost rt . These profiles define t model store and measures the latency, throughput, and memory the system cost of deploying a selected workflow configuration. footprint of each variant on each hardware tier. Both sets of Under this model, variant choices determine task accuracy, profiles are stored in the profile catalog. while placement determines serving behavior. Atlas therefore In the optimization phase, the plan estimator consumes the treats workflow accuracy as a function of the workflow per-pair accuracy profiles and applies Markovian Accuracy Predictor (MAP) to estimate end-to-end accuracy for every configuration. As the Atlas optimizer relies on the accuracy escandidate configuration. The Atlas optimizer selects the model timation to rank candidate execution plans, we denote predicted variant assignment and hardware placement per stage that accuracy for an execution plan p extending configuration c as maximizes predicted accuracy subject to the SLO constraints. Ê(p) = Ê(c). In Atlas, this is computed through MAP from The resulting execution plan is submitted to the workflow the local accuracy profiles, further described in Section IV. Given an execution plan p, Atlas estimates serving feasibility executor, which deploys the selected variants on their assigned from the system profiles. As workflows can incur different hardware tiers and serves inference requests until a new plan topologies, latency is evaluated over feasible request paths. The replaces it. workflow specification and the selected control-flow parameters in configuration c define a finite set of paths P(c). A path A. System Model q ∈ P(c) is a sequence of stage invocations. This means that A compound AI workflow is represented as a directed graph routed topologies appear as different paths per branch and G = (S, E), where each stage s ∈ S is one AI model or iteration topologies appear as repeated stage invocations. component invocation, and each edge (si , sj ) ∈ E denotes the Atlas performs offline execution plan selection for a single data dependence between two stages. Each stage exposes a site cluster. Additionally, we consider that all workers communifinite set of variants Vs . A variant specifies the AI model used cate through the same cluster network fabric. Thus, we assume at that stage and the model-specific hyperparameters exposed that intra-cluster communication incurs negligible overhead to optimization. A workflow configuration c represents one relative to total execution time. Consequently, the system model variant setting for every stage of the workflow. does not represent inter-stage communication explicitly and Atlas deploys the workflow on a heterogeneous cluster attributes end-to-end latency only to stage executions. With located within one physical site. The cluster contains a finite this in mind, we model the estimated workflow latency as: set of hardware classes T , such as CPU workers and GPU   workers with different accelerators. |q| X An execution plan extends a workflow configuration with L(p) = max  ℓ(si , c(si ), tsi ) . (3) q∈P(c) deployment decisions. Given a configuration c, the plan selects, i=1

Let λ denote the operator-specified workflow throughput target. Atlas uses conservative per-stage provisioning, requiring every deployed stage to sustain the full workflow rate. Since the selected deployment for stage s provides capacity θ(s, c(s), ts ), throughput feasibility requires: θ(s, c(s), ts ) ≥ λ,

∀s ∈ S.

(4)

Memory feasibility is enforced per hardware class. Let Mt denote the aggregate memory capacity available on workers of class t. The total memory footprint of all stages placed on that class must fit within this capacity: X

µ(s, c(s), ts ) ≤ Mt ,

∀t ∈ T.

(5)

A. Workflow and Quality States MAP assigns every stage output to a quality bucket. Quality refers to the intermediate per-stage signal, such as retriever hit rate or answer F1, that propagates through the workflow and determines workflow accuracy at the terminal stage. MAP discretizes this signal into B buckets whose boundaries are computed per stage from calibration data pooled over variants, making buckets comparable across variants at the same stage. MAP assumes that the bucketed quality state preserves the information needed to predict downstream quality. Specifically, that it retains the signal relevant for ranking configurations, not all semantic properties of an intermediate output. A bucket trajectory is the sequence of quality buckets produced as an input moves through the workflow,

s∈S:ts =t

b = (b1 , . . . , bS ), Let rt denote the hourly vendor price of one worker on hardware class t. The hourly deployment cost of execution plan p is C(p) =

X

rts .

(6)

s∈S

Equations (3)–(6) define the system-level constraints under which workflow accuracy is optimized.

where bs ∈ {1, . . . , B} is the bucket realized at stage s. For a configuration c, MAP induces a trajectory distribution P (b | c). The predicted end-to-end accuracy is Ê(c) = Eb∼P (b|c) [ϕ(b, c)],

(8)

where ϕ maps the terminal quality bucket to a scalar accuracy value. B. Local Transition Model

B. Problem Definition Given a workflow graph G = (S, E), per-stage variant sets {Vs }s∈S , hardware classes T , system profiles (ℓ, θ, µ), hardware capacities and costs (M, r), and operator constraints (Lmax , Cmax ), Atlas selects an execution plan that maximizes predicted task accuracy under latency and cost constraints. The optimization problem is

MAP models quality propagation as a first-order Markov chain over quality buckets, estimating transition probabilities from per-pair quality profiles of adjacent stages. For adjacent stages s − 1 and s, the transition Gs (b′ | b, v, v ′ ),

v ∈ Vs , v ′ ∈ Vs−1 ,

(9)

gives the probability that stage s, using variant v, produces quality bucket b′ given that stage s − 1, using variant v ′ , produced quality bucket b. Intuitively, Gs captures how the max Ê(p) p quality produced by an upstream AI model affects the quality (7) distribution of its downstream neighbor. For example, a retriever s.t. L(p) ≤ Lmax , that retrieves low-quality documents shifts the downstream genC(p) ≤ Cmax erator’s output distribution towards lower accuracy, regardless The optimization is additionally subject to the throughput of which generator variant is selected. and memory feasibility constraints in Equations (4) and (5). Each row of Gs fixes one upstream quality bucket b and The remaining question is how to construct Ê(c) so that one variant pair (v, v ′ ) and gives a categorical distribution over it preserves the ordering of candidate configurations well downstream quality buckets, estimated from calibration samples enough to support optimization. The next section addresses this with smoothing applied to avoid zero-probability transitions in question through an accuracy predictor over stage interactions. low-support cells. Conditioning only on the immediate upstream bucket and selected variants is a deliberate design choice. IV. MAP: M ARKOVIAN ACCURACY P REDICTOR It keeps profiling cost local to adjacent stage pairs while still capturing the direct quality interactions that determine Atlas uses the Markovian Accuracy Predictor (MAP) to esconfiguration rankings. timate end-to-end configuration accuracy from per-pair quality As the first workflow stage doesn’t have an upstream bucket, profiles, avoiding exhaustive profiling. MAP tracks intermediate MAP estimates the initial distribution output quality as a discrete state across workflow stages and composes local conditional quality transitions according to the π1,b (v), v ∈ V1 , (10) workflow topology. The Plan Estimator invokes MAP to score every candidate configuration before the optimizer selects an the probability that stage s = 1 produces quality bucket b execution plan. when using variant v, estimated by running each candidate

Stage 2

Stage 1

Stage 3

Stage 4

Gj (b′ | z, vR , vj ), Example Bucket Trajectory Initial Bucket Distribution

Fig. 4. MAP’s local transition model for a linear workflow. Stage outputs are mapped to quality buckets, adjacent stages define conditional transition matrices, and the highlighted path shows one possible bucket trajectory.

variant on the calibration inputs and recording the bucket frequency. Figure 4 shows the local transition model. An initial quality bucket distribution at stage s = 1 and a conditional transition matrix Gs for each subsequent stage pair, together covering every candidate variant assignment in the workflow.

vj ∈ Vj ,

(14)

gives the probability that branch j, using variant vj , produces quality bucket b′ given routing state z and router variant vR . Intuitively, Rj captures which branch an input is sent to, while Gj captures what quality that branch produces once selected. Let µin (z) denote the distribution over routing states at the router input, estimated from calibration data for a standalone router or supplied by the preceding segment in a composed workflow. The joint trajectory probability is P (z, j, b′ | c) = µin (z) Rj (z | vR ) Gj (b′ | z, vR , vj ), and the predicted accuracy is X Ê(c) = P (z, j, b′ | c) qj,b′ ,

(15)

(16)

z,j,b′

where qj,b′ is the average accuracy of calibration inputs that were sent to branch j and produced quality bucket b′ . 3) Feedback Loop: A feedback loop alternates between a producer stage, which generates an output, and an evaluator C. Topology Operators stage, which scores it and triggers revision if the quality is MAP instantiates Equation 8 for three main topology primi- insufficient (e.g., a generator and critic in a self-refinement tives that cover the most common compound AI workflows: workflow). The producer’s initial output follows the bucket sequential pipelines, routed workflows, and loops. Each prim- distribution π0,b (vp ) defined in Section IV-B. The feedback itive defines how quality bucket distributions are propagated transition through its stages and how the terminal distribution is mapped T (b′ | b, vp , ve ), vp ∈ Vp , ve ∈ Ve , (17) to a predicted accuracy value. Composed workflows are handled by chaining these primitives, as described in Section IV-D. gives the probability that one feedback step moves the output 1) Linear Pipelines: A linear pipeline executes stages from quality bucket b to quality bucket b′ under producer v p sequentially, where the output of stage s − 1 becomes the input and evaluator v . Intuitively, T captures how effectively an e to stage s. Under the local transition model, the trajectory evaluator variant drives quality improvements in the producer distribution factorizes as output across feedback steps. P (b | c) = π1,b1 (v1 )

S Y

Gs (bs | bs−1 , vs , vs−1 ),

For a fixed budget K, the quality distribution after K feedback steps is (11) µK = π0 (vp ) T (vp , ve )K ,

s=2

and the predicted accuracy is the expected quality of the terminal bucket, X Ê(c) = P (b | c) qS,bS , (12)

and the predicted accuracy is Ê(c) =

B X

µK,b qb ,

(19)

b=1

b

where qS,bS is the average task accuracy of calibration inputs that produced quality bucket bS at the terminal stage, obtained from the same calibration data used to estimate the transition matrices. 2) Routed Workflows: A routed workflow sends each input to one of several downstream branches based on a discretized routing signal z, such as an upstream quality bucket. MAP estimates two quantities for each candidate router variant vR ∈ VR and branch variant vj ∈ Vj . The routing probability is Rj (z | vR ),

(18)

vR ∈ VR ,

(13)

For each fixed routing state z, Rj (z | vR ) gives the probability that router variant vR sends the input to branch j, P with j∈J Rj (z | vR ) = 1. The branch transition

where qb is the average task accuracy of calibration inputs that produced quality bucket b after K feedback steps. The fixed budget model assumes every input undergoes exactly K feedback steps. When the workflow stops adaptively based on evaluator score, MAP augments T with an absorbing stop state,   T pstop T̃ = cont , (20) 0 1 where Tcont captures transitions that continue the feedback loop and pstop gives the stopping probability from each quality bucket. The augmented chain produces two outputs: the final quality distribution, used to compute Ê(c), and the expected iteration count per input, which Atlas uses to estimate latency and throughput for the execution plan.

V. P LAN O PTIMIZER The problem in Equation (7) selects a workflow configuration and deployment tier per stage to maximize accuracy under SLO constraints on a heterogeneous cluster, as defined in Section III-A. We cast it as a MILP using two structural properties. Predicted accuracy {Ê(c)} from MAP is precomputed per Fm (µin (21) configuration and enters the objective as a constant coefficient, m , cm ), while the system constraints are linear in the deployment which maps an input quality bucket distribution µin m and segQ ment configuration cm to an output quality bucket distribution indicators. Decision variables. Let C = s∈S Vs denote the conµout . Each segment applies its own topology operator from m figuration space. The MILP uses the configuration selector P Section IV-C. ξc ∈ {0, 1} with ξ = 1; and deployment indicators c c P Passing quality information across segment boundaries ωs,v,t ∈ {0, 1} with ωs,v,t = 1 for each s, denoting v,t requires one additional profiled quantity. Each segment defines that stage s runs variantPv on tier t.PVariant selection is its own bucket boundaries from calibration data pooled over its the projection xs,v := out t ωs,v,t = c : c(s)=v ξc , linking own stages, so the output bucket distribution µm of segment configuration selection to deployment. m is not directly comparable to the input bucket space of Latency. Each request path through the workflow must segment m + 1. Thus, the boundary transition satisfy the latency SLO: XX (22) Gentry (b′ | b, ventry ), ventry ∈ Ventry , ℓ(s, v, t) ωs,v,t ≤ Lmax , ∀q ∈ P. (25) D. Workflow Operator Composition A composed workflow chains multiple topology primitives into a single workflow. For example, a sequential RAG pipeline may feed into a feedback loop backend. MAP treats each segment m as an operator

gives the probability that the entry stage of segment m + 1 produces quality bucket b′ given that segment m ended in quality bucket b. The entry distribution of segment m + 1 is µm+1,0 (b′ ) =

B X

′ µout m (b) Gentry (b | b, ventry ),

s∈q v,t

Throughput. Each deployed stage must serve its effective request rate: X θ(s, v, t) ωs,v,t ≥ Ks λ ∀s ∈ S, (26) v,t

(23)

where Ks is the iteration bound of the feedback loop enclosing s, and Ks = 1 outside loops. Thus, a stage inside a loop after which segment m+1 applies its own topology operator, serves each request up to Ks times. This mirrors the latency constraint, where loop iterations are unrolled into the request µout (24) paths P. m+1 = Fm+1 (µm+1,0 , cm+1 ). Memory. The aggregate footprint hosted on each tier must This preserves cross-segment quality dependence without conditioning on the full upstream trajectory, keeping composi- fit its capacity: X tion profiling cost additive across segment boundaries. µ(s, v, t) ωs,v,t ≤ Mt ∀t ∈ T . (27) s,v E. Complexity Analysis Cost. Hourly deployment cost sums each stage’s tier rate: 1) Profiling cost: MAP profiles only adjacent stage pairs. A X 2 linear segment of S stages requires O(SV B) conditional rows, rt ωs,v,t ≤ Cmax . (28) where V is the maximum variants per stage and B theP number s,v,t of quality buckets. A routed segment adds O(VR Bz j |Vj |) Objective. Predicted end-to-end accuracy enters as a prerows for the router and branches, while a feedback loop adds computed coefficient on ξc : X O(Vp Ve B) for the feedback transition. Each segment boundary max Ê(c) ξc . (29) adds O(Ventry B). Overall, profiling cost grows additively ξ, ω c∈C S across primitives and boundaries, compared with O(V ) for The values {Ê(c)}c∈C are predicted by MAP from the local exhaustive end-to-end profiling. MAP’s local structure reduces the cost of workflow evolution. pairwise profiles and stored before the solver is invoked, keepAdding one variant to a stage requires O(V · B) new pairwise ing the program linear inPall decision variables. A small costtransitions instead of V S−1 new configurations for end-to-end minimizing tiebreaker ε s,v,t rt ωs,v,t with ε = 10−4 /Cmax profiling. Inserting a new stage requires profiling only its two is subtracted from the objective to prefer cheaper placements among configurations with equal predicted accuracy. adjacent stages instead of complete re-profiling. 2) Prediction cost: For each topology, MAP operates on Complexity. All constraints and the objective are linear, so bucket distributions directly. A linear segment requires one the LP relaxation is convex. Integrality of ξ and ω makes MILP matrix-vector multiplication over B buckets per stage, giving solving NP-hard in general. MAP reduces profiling cost, while a prediction cost of O(SB 2 ) per configuration. For routed the MILP still represents the configuration space through ξc . segments, MAP sums over branch assignments. For feedback Because MAP scores configurations before solving, those with loops, it sums over iteration counts. In both cases prediction lower predicted accuracy at equal resource footprint can be cost remains polynomial in B and the number of stages. pruned to keep the problem tractable. b=1

TABLE II W ORKFLOWS USED IN THE EVALUATION . ID

Topology

Workflow

Configuration space

W1

Linear

RAG

kr ∈ 5, 10, 20; reranker ∈ MS-MARCO, BGE-base; krr ∈ 3, 5; generator ∈ Llama3.2 1B/3B, Llama-3.1 8B, Gemma-3 1B/4B/12B, Phi-3 3.8B, Qwen2.5 1.5B [35]–[39]

96

W2

Linear + Routing

RAG + Router

W1 with fixed kr and krr ; reranker ∈ MS-MARCO, BGE-base; DistilBERT threshold t ∈ 0.3, 0.4, 0.5, 0.6; local generator ∈ Llama-3.2 1B/3B, Gemma-3 1B, Qwen2.5 1.5B; remote generator ∈ Llama-3.1 8B, Gemma-3 4B/12B, Phi-3 3.8B.

128

W3

Linear + Feedback Loop

RAG + Refine

Five selected W1 configurations; four producer and evaluator pairs drawn from Llama3.2 1B/3B and Gemma-3 4B producers and Llama-3.2 1B and Gemma-3 4B evaluators; feedback budget K ∈ {1, 3, 5}.

60

W4

Linear + Routing + Feedback Loop RAG + Router + Refine

W1 with fixed kr and krr ; reranker ∈ MS-MARCO, BGE-base; DistilBERT threshold t ∈ 0.3, 0.4, 0.6; local producer ∈ Llama-3.2 1B/3B, Gemma-3 4B; evaluator ∈ Llama-3.2 1B, Gemma-3 4B; remote generator ∈ Gemma-3 4B/12B, Llama-3.1 8B, Phi-3 3.8B, Mistral 7B, Qwen2.5 7B; K = 5.

216

VI. E VALUATION This section presents a series of experiments as means to evaluate Atlas. Section VI-A details the carried-out experiments, experimental frameworks, and evaluation objectives, while Sections VI-B through VI-F present the results. A. Experimental Setup Atlas is implemented in Python and published as an opensource framework within the Polaris project1 . For MILP optimization, it uses the CBC solver through the python-mip library [40]. We evaluate Atlas by measuring how well MAP predicts accuracy across different compound AI workflows and whether these predictions translate into effective optimizer decisions. We first assess whether local quality profiles preserve the decision-relevant ordering of configurations across topologies presented in Section IV. We then test the execution plans produced by the Atlas optimizer against baselines under both homogeneous and heterogeneous cluster settings. 1) Workflows and Datasets: We evaluate Atlas across four workflows that cover the topology patterns modeled by the framework. The linear RAG workflow [25] (W1) tests conditional quality propagation across sequential stages. W2 extends W1 with a learned router [8], [28] that dispatches inputs to different generator branches. W3 extends W1 with a generator-critic feedback loop [18], [34], testing iterative quality evolution. W4 combines routing and refinement on top of W1, testing whether MAP can pass quality distributions across topology boundaries. Table II summarizes the configuration space for each workflow, including model variants and hyperparameters. W2, W3, and W4 each build on a subset of W1 configurations, extending them with routing, refinement, or both. Each configuration count is the product of the listed factor cardinalities, for example 3 × 2 × 2 × 8 = 96 for W1. All workflows are evaluated on the SQuAD [41] dataset using answer F1 as the end-to-end accuracy metric. Quality labels at intermediate stages describe whether retrieval and reranking preserve the evidence needed by the generator, the router’s dispatch decision and downstream branch quality, and the quality trajectory across refinement steps. MAP uses B = 4 1 https://github.com/polaris-slo-cloud/Atlas

Configs.

quality buckets. Bucket boundaries are computed per stage from 200 calibration samples pooled over all variants at that stage. 100 held-out samples are used for evaluation. 2) Baselines: We compare MAP against three accuracy estimation baselines. PAS-naive multiplies standalone benchmark accuracies per stage (Eq. 1), following the Pipeline Accuracy Score used by IPA [13]. PAS-fair is the strongest PASstyle baseline we can construct per topology. It retains scalar composition but replaces standalone scores with topologyaware terms drawn from the same calibration data available to MAP, such as conditioned stage accuracies for linear workflows and routing-weighted branch accuracies for routed workflows. The oracle reference profiles every configuration end to end on the evaluation dataset, following the exhaustive profiling approach used by [3], [14]. For optimizer evaluation, we compare execution plans produced by Atlas against IPA and Loki [14] on a homogeneous cluster, restricting Loki to singlevariant-per-stage placement for direct plan-level comparability. 3) Infrastructure: The evaluation uses three hardware tiers: an NVIDIA RTX 4090 (24 GB VRAM, $0.40/h), an NVIDIA RTX A4000 (16 GB VRAM, $0.20/h), and 32-core x86 CPU tier (32 GB RAM, $0.10/h). Hourly tier costs are set to approximate public cloud rates for comparable hardware. Experiments in Sections VI-B through VI-D run on a single RTX 4090 server. The heterogeneous evaluation VI-E deploys all three tiers on a K3s cluster spanning three physical nodes. 4) Metrics: We use two groups of evaluation metrics. For the accuracy model, we report Spearman and Kendall rank correlation against the oracle ranking, top-5 overlap, and top-1 regret, defined as the accuracy difference between the oracle’s best configuration and the one ranked first by the predictor. Rank correlation and top-k overlap evaluate whether the predictor preserves the ordering needed for optimization, while regret measures the practical cost of prediction errors. For the optimizer, we report the measured end-to-end accuracy of the selected execution plan, its p99 latency, and hourly deployment cost under varying SLOs and offered loads. The evaluation targets four objectives. First, it assesses whether MAP preserves the measured ranking of workflow configurations across all topology classes (VI-B). Second, it measures the profiling cost at which MAP achieves a useful

TABLE III MAP PREDICTION QUALITY ACROSS WORKFLOWS COMPARED TO PAS AND PAS- FAIR .

Workflow

Predictor

n

Spearman ρ

Kendall τ

MAE

Top-5 overlap

Top-1 regret

W1 W1 W1

PAS-naive PAS-fair MAP

96 96 96

0.295 0.555 0.947

0.203 0.378 0.805

0.279 0.085 0.036

0.00 0.00 0.80

+0.322 +0.240 +0.000

W2 W2 W2

PAS-naive PAS-fair MAP

128 128 128

0.186 0.425 0.921

0.171 0.329 0.757

0.451 0.161 0.098

0.00 0.20 0.80

+0.720 +0.699 +0.046

W3 W3 W3

PAS-naive PAS-fair MAP

60 60 60

−0.372 0.619 0.783

−0.250 0.487 0.595

0.267 0.126 0.069

0.20 0.40 0.60

+0.000 +0.000 +0.127

W4 W4 W4

PAS-naive PAS-fair MAP

216 216 216

0.068 0.294 0.882

0.023 0.234 0.697

0.642 0.152 0.076

0.00 0.20 0.60

+0.657 +0.632 +0.094

correlation and lowest regret in the table. Figure 5 confirms this visually. W1 and W2 concentrate close to the diagonal, indicating that MAP is both well ranked and calibrated on linear and routed workflows. W3 and W4 show wider spread because refinement introduces additional variance, but the point clouds still preserve the ordering needed for plan selection. The wider spread on W3 follows from the first-order Markov assumption. MAP applies one profiled feedback transition B. MAP Prediction Quality T (vp , ve ) K times, and repeated application mixes toward its We first evaluate whether MAP preserves the configura- stationary distribution, gradually erasing the bucket distribution tion ranking induced by end-to-end measurements. This is passed from the upstream W1 segment. Predicted accuracy the accuracy model’s primary requirement for optimization, across configurations that differ only in their W1 segment because the optimizer uses predicted accuracy to compare collapses from a spread of 0.29 at K = 1 to at most 0.03 at candidate configurations under serving constraints. Table III K ≥ 3, while the measured spread remains up to 0.36, and reports ranking, calibration, and selection metrics for four MAE grows from 0.036 to 0.095. Nevertheless, the ordering workflows, and Figure 5 shows MAP’s predicted accuracy within each W1 segment survives, which keeps MAP the best ranking predictor on W3. Conditioning the feedback transition against measured accuracy for each configuration. on the quality bucket at loop entry would restore the upstream dependence at a profiling cost multiplied by B. W1 W2 The gap between MAP and PAS-style baselines grows with 0.8 =0.95 0.8 =0.92 n=96 n=128 compositional complexity. On W1, MAP improves Spearman 0.6 0.6 correlation over PAS-fair by 0.39; on W4, the gap increases to 0.59. This pattern reflects that scalar composition becomes less 0.4 0.4 reliable as workflow behavior depends on cross-stage accuracy propagation. PAS-fair can correct some standalone effects, but 0.4 0.6 0.8 0.4 0.6 0.8 it does not represent how retrieval accuracy changes router W3 W4 =0.78 =0.88 decisions or how routed outputs affect refinement. PAS-naive 0.6 n=60 0.8 n=216 is anti-correlated on W3 (ρ = −0.37), meaning benchmarkproduct composition can invert the true configuration ordering. 0.4 0.6 Top-5 overlap and regret metric reinforce this result. MAP 0.2 0.4 recovers 60-80% of the true top-5 configurations across all workflows, while PAS-naive achieves zero overlap on W1, 0.2 0.4 0.6 0.4 0.6 0.8 Predicted quality Predicted quality W2, and W4. Since the optimizer selects from the top-ranked feasible configurations, accurate ranking at the top of the list Fig. 5. MAP predicted versus measured accuracy across four workflows. directly determines the quality of the chosen execution plan.

Measured quality

Measured quality

ranking signal relative to exhaustive end-to-end measurement (VI-C). Third, it evaluates whether the Atlas optimizer, guided by MAP, produces execution plans that match oracle-quality plan selection compared to existing approaches (VI-D). Fourth, it tests whether Atlas’s joint optimization over variant selection and heterogeneous tier placement produces better execution plans than single-tier deployment strategies (VI-E).

MAP gives the strongest ranking signal on all four workflows, with Spearman correlation ranging from ρ = 0.78 on W3 to ρ = 0.947 on W1, while also achieving the best Kendall

C. Profiling Cost We next evaluate whether MAP reduces the measurement cost needed to obtain a useful ranking signal. Figure 6 compares

MAP with PAS baselines and exhaustive end-to-end profiling.

Profiled units

Spearman

Oracle: V S

0.75 0.50 0.25 102

103

Unit-sample measurements (a) Profiling budget

Atlas-MAP

IPA (b) W2, 500 ms

0.8

106

1.00

0.00

Measured accuracy

PAS: S V MAP: (S 1) B V 2

Oracle E2E MAP

0.6

104

0.4

102 2 4

8

16

Variants per stage | |

32

(b) Synthetic scaling

Fig. 6. Profiling cost of MAP compared to PAS and exhaustive profiling.

Measured accuracy

PAS-fair PAS-naive

Oracle / Loki (a) W1, 500 ms

(c) W1, 800 ms

(d) W2, 800 ms

0.8 0.6 0.4

Figure 6(a) shows Spearman rho as the number of per-pair measurements increases. W1 contains 23 such terms, so using K calibration samples per term requires 23K measurements. PAS-fair and PAS-naive appear as horizontal lines because their scalar estimates do not improve with additional workflowspecific measurements. End-to-end profiling gives the reference ranking but requires measuring all 96 configurations on 200 samples. MAP surpasses PAS-fair with 115 measurements and reaches ρ = 0.93 with 1,840 measurements, recovering most of the reference ranking signal at 2.6× lower cost. Figure 6(b) shows how this gap grows with the number of variants per stage in a synthetic scaling scenario. End-to-end profiling scales with the full configuration space, whereas MAP scales with local pairwise terms. At small variant counts (|V| ≤ 4), end-to-end profiling remains practical and may require fewer measurements than MAP’s pairwise terms. However, the gap inverts quickly: at |V| = 8, MAP requires 5.3× fewer profiling units, and at |V| = 32, the gap widens to over 80×.

1

2

5

10

Load (req/s)

20

1

2

5

10

Load (req/s)

20

Fig. 7. Execution plan accuracy comparison with baselines over varying SLOs.

reaching 0.40 on W1 and 0.35 on W2 at λ = 20. The accuracy gap between Atlas and IPA grows with workflow complexity, exceeding 0.30 on every W2 load with a median gap of 0.36. This confirms findings from Section VI-B, where PAS achieves weak Spearman correlation on routed workflows. Weak accuracy estimation leads to misordered configurations, which in turn leads the optimizer to select worse execution plans. Since all three systems share the same optimizer, the observed differences are attributable solely to the accuracy model. E. Plan Selection on a Heterogeneous Cluster

We evaluate whether heterogeneous tier placement produces better execution plans than single-tier strategies. We hold the D. Plan Selection on a Homogeneous Cluster accuracy predictor fixed to MAP so that observed differences We evaluate the execution plans produced by Atlas against are attributable solely to placement choice. Under four latency those of IPA and Loki on a homogeneous single-tier cluster SLOs (400–1000 ms), we compare four strategies on W1: All(RTX 4090, 24 GB memory). We restrict placement to a CPU, All-A4000, All-RTX4090, and Atlas-Het, which uses single hardware tier to match the conditions under which IPA the full MILP with unrestricted tier assignment per stage. and Loki are designed, and evaluate heterogeneous placement Figure 8 reports measured accuracy, p99 latency, and plan separately in VI-E. Each system is given the same workflow, cost per SLO. All-CPU is infeasible because LLM p99 on same candidate variants, and same latency SLO. We report the CPU exceeds 1.6 s. All-A4000 reaches measured accuracy of measured accuracy of the chosen plan. 0.73 at the 400 ms SLO, while All-RTX4090 and Atlas-Het To isolate the contribution of the accuracy model, all three reach 0.79, because A4000 latency forces the optimizer to systems are run through Atlas’s MILP with identical constraints. select a smaller LLM variant. At 600 ms and above, all GPUUnder single-variant-per-stage assignment, Loki reduces to based strategies converge on 0.79. Atlas-Het achieves the same a direct end-to-end accuracy lookup and is plotted together accuracy as All-RTX4090 at 42% lower plan cost at 400 ms with the oracle reference. Loki’s full formulation supports and 33% lower at 600 and 800 ms, by placing the embedder concurrent multi-variant deployment per stage, but we restrict and reranker on cpu-edge and the LLM on A4000. This to single-variant assignment for direct plan-level comparability. cost reduction follows from the tiebreaker term in the MILP We evaluate on W1 and W2 at two latency SLOs (500 ms, 800 objective, which selects the cheapest placement among equallyms) and five offered loads (λ ∈ {1, 2, 5, 10, 20} req/s). W3 accurate configurations. Heterogeneous placement therefore and W4 are excluded as neither IPA nor Loki express a system matches single-tier accuracy at lower cost across feasible SLOs. model that supports feedback loop topologies. We next test whether MAP retains oracle-quality plan Figure 7 reports measured accuracy of the selected execution selection on the heterogeneous cluster for W4. We fix placement plan as a function of offered load under varying SLOs. Atlas to Atlas-Het and vary only the accuracy predictor across stays within 0.03 of Oracle/Loki across all loads and SLO Atlas-MAP, Atlas-PAS, and Atlas-Oracle. We sweep three cost tiers on both workflows, while IPA falls sharply with load, budgets ($0.90, $1.20, $5.00) under a loose SLO of 8000 ms to

×

×

400

×

600

×

800

Latency SLO (ms)

1000

All-A4000

All-RTX4090

Atlas-Het

Plan cost ($/h)

0.8 0.6 0.4 0.2 0.0

p99 latency (ms)

Measured accuracy

All-CPU 400 200 0

×

400

(a) Measured accuracy vs SLO

×

600

×

800

×

Latency SLO (ms)

1000

1.0 0.5 0.0

×

×

400

×

600

×

800

Latency SLO (ms)

(b) Measured p99 latency vs SLO

1000

(c) Cost vs SLO

Fig. 8. Homogeneous versus heterogeneous placement on W1: measured accuracy, latency, and plan cost across latency SLOs.

(a) Acc. vs Cost Budget

1.00 0.75 0.50 0.25 0.00

3000 2000 1000

0.5 $0.90 $1.20 $5.00

Cost budget ($/h)

SLO headroom (%)

Plan cost ($/h)

1.0

0.0

(b) Lat. vs Cost Budget

$0.90 $1.20 $5.00

(c) Plan Cost vs Budget

Atlas-Oracle

0

Spearman

Atlas-PAS

p99 latency (ms)

Measured accuracy

Atlas-MAP

1.0 0.8 0.6 0.4 0.2 0.0

(d) SLO vs Cost Budget

100

W2

3

4

B=1 (PAS)

1

$0.90 $1.20 $5.00

W1

2

W3

W4

B=4 (Default)

5

6

Number of accuracy buckets B

7

8

Fig. 10. MAP ranking quality as a function of bucket count B across four workflows

75 50 25 0

$0.90 $1.20 $5.00

Cost budget ($/h)

Fig. 9. Accuracy predictor comparison on W4 with heterogeneous placement

ensure that the cost budget, not latency, is the binding constraint in this experiment. Figure 9 reports measured accuracy, p99 latency, and plan cost across cost budgets. Atlas-MAP matches Atlas-Oracle at every budget, reaching 0.58 at $0.90 and 0.90 at $1.20 and $5.00. Atlas-PAS lags by 0.22 to 0.35 in measured accuracy, its p99 latencies are approximately 450 ms higher, and exceed Atlas-MAP cost by $0.20/h at the loose budget. Since placement and optimizer are held fixed, these differences are attributable to the accuracy model. PAS selects bge-base as the reranker because its standalone benchmark score is higher than msmarco, but bge-base interacts worse with downstream stages. Additionally, at 494 ms p99 on cpu-edge, bge-base cannot be placed on the cheap tier, forcing the MILP onto A4000 and increasing cost. MAP conditions the reranker’s contribution on the quality it passes downstream, correctly selects ms-marco, and recovers oracle-quality plans across the full budget range. F. Bucket Sensitivity Analysis We analyze MAP’s sensitivity to the number of quality buckets B, which sets the granularity of intermediate output discretization. Figure 10 shows MAP’s ranking quality as a function of the number of B for all four workflows. At B = 1, each stage is described by a single scalar accuracy value and

MAP reduces to PAS-style composition. Increasing B to 2 produces the largest improvement across all workflows, as even a binary partition allows MAP to condition downstream behavior on upstream quality. All workflows plateau by B = 4, with marginal gains beyond that point. The dip at B = 3 on W3 occurs because W3’s F1 distribution is heavily concentrated at 0 and 1. This causes the equal-frequency quantile boundaries to land on thresholds that merge failures with mediocre outputs into a single bucket, losing the discriminative split that B = 2 and B = 4 preserve. We use B = 4 as the default throughout the evaluation, as it provides stable ranking quality across all topologies. Automatic bucket selection from calibration data is a natural extension of this sweep. VII. R ELATED W ORK Prior work relevant to Atlas spans three areas: 1) inference serving, 2) accuracy estimation for multi-stage workflows, and 3) optimization formulations for inference deployment. A. Inference Serving Systems Inference serving systems such as Clipper, Clockwork, and Nexus establish the substrate for low-latency ML inference through batching, model management, GPU scheduling, and predictable execution under latency SLOs [42]–[44]. A second line of work makes accuracy a first-class objective by selecting among model variants under latency, throughput, or cost constraints. ModelSwitching switches to cheaper models under load spikes, INFaaS automates model and hardware selection, Cocktail optimizes ensembles, RAMSIS selects models using inter-arrival-aware scheduling, and Proteus performs accuracy scaling for high-throughput serving [19]–[22], [45]–[48].

These systems laid the foundation for accuracy-aware inference serving but target a single task or model pool at one endpoint. Atlas targets compound AI workflows where downstream accuracy depends on upstream output quality.

process. This works when the upstream quality bucket captures the information most relevant to downstream behavior, as shown by the strong ranking results across the evaluated workflows. This assumption may be less accurate when later stages depend on outputs several hops earlier. Atlas supports sequential B. Accuracy Estimation for AI Workflows pipelines, routed workflows, loops, and their composition, Two lines of approaches estimate the accuracy of AI covering common patterns such as RAG. Other workflow workflows. Product-based surrogates compose standalone stage structures, such as fan-out aggregation workflows, would accuracies multiplicatively, of which IPA’s Pipeline Accuracy require additional operators. MAP also assumes that each stage Score (PAS) is representative [13]. This approach is com- exposes an intermediate quality signal that can be discretized putationally efficient since per-stage measurements can be into buckets. For stages without such a signal, task-specific reused across configurations, but the independence assumption instrumentation may be needed before transitions can be fails when upstream output quality affects downstream stage profiled. Finally, MAP reduces profiling cost substantially behavior. Several approaches rely on end-to-end accuracy compared to exhaustive profiling, but it does not eliminate profiling on a representative dataset. This faithfully captures profiling entirely. For small configuration spaces, exhaustive inter-stage interactions but scales with the full configuration profiling may remain simpler, while Atlas pays off for larger space [3], [14], [23], [49], [50]. As the number of stages, AI workflows where exhaustive profiling becomes impractical and model variants, and hyperparameters, grows, exhaustive end-to- product surrogates fail to preserve configuration rankings. end profiling becomes intractable for deployment optimization. In Atlas, MAP sits between these two extremes. It profiles IX. C ONCLUSION local conditional quality transitions between adjacent workflow Deploying Compound AI workflows requires selecting an stages, preserving the cross-stage accuracy dependencies that product-based surrogates discard, while keeping profiling cost execution plan that maximizes accuracy under SLO constraints. proportional to the number of stage pairs rather than the full Accurate estimation of workflow accuracy is the central obstacle in Compound AI deployment optimization. Atlas addresses configuration space that end-to-end approaches must cover. this by separating accuracy estimation from system profiling C. Deployment Optimization for AI Workflows and keeping both tractable. MAP profiles conditional quality Several systems formulate AI workflow deployment as transitions between adjacent stages, discretizes intermediate constrained optimization. IPA optimizes variant selection, batch outputs into quality buckets, and composes the transitions sizes, replicas, and resource allocation for linear inference according to the workflow topology. The Atlas optimizer selects pipelines as an integer program, using PAS as the accuracy ob- execution plans with a MILP over the predicted accuracy and jective [13]. Loki combines hardware and accuracy scaling for the SLO constraints. Across four Compound AI workflows, tree-structured pipelines using MILP-based resource allocation MAP achieves Spearman correlation up to 0.947 and reduces and runtime routing to reduce SLO violations [14]. Both target profiling cost by 2.6× relative to exhaustive profiling. Critically, homogeneous clusters and linear or tree-structured pipelines, MAP’s ranking advantage over product based approaches grows modeling neither routed, feedback, nor composed compound AI with workflow complexity, where deployment optimization topologies. Other systems optimize deployment without treating matters most. Guided by MAP, the Atlas optimizer selects accuracy as a joint objective. InferLine provisions and scales execution plans within 0.03 of oracle accuracy across all prediction pipelines under latency constraints with accuracy evaluated SLOs, while heterogeneous tier placement matches fixed externally [24]. JellyBean deploys ML workflows across oracle accuracy at up to 42% lower deployment cost than heterogeneous edge-to-cloud tiers, minimizing cost subject to homogeneous placement strategies. throughput and accuracy SLOs [23]. Future work will extend MAP with operators supporting Atlas differs along two dimensions. First, it formulates various topologies, such as fan-out. Additionally, MAP accuracy plan selection for compound AI workflows where variant modeling will be explored to support dynamic adaptation selection, hardware placement, and topology jointly determine mechanisms, such as runtime model selection under varyfeasibility and predicted accuracy. Second, it uses MAP to ing workloads. Finally, Atlas will incorporate dynamic reestimate configuration-level accuracy from local conditional optimization, enabling execution plans to be revised as cluster quality transitions, whereas existing systems either fix accuracy, conditions change without requiring full re-profiling. compose standalone stage scores independently, or rely on endto-end profiling. ACKNOWLEDGMENT VIII. D ISCUSSION Atlas provides an efficient middle ground between exhaustive profiling and product-based accuracy surrogates. However, it relies on several design assumptions. MAP models quality propagation through adjacent stage pairs as a first-order Markov

This work was partly funded by the European Union under the Horizon Europe programme through the SNS JU (Grant Agreement No. 101192912, NexaSphere). Views expressed are those of the authors and do not necessarily reflect those of the EU or the SNS JU.

R EFERENCES [1] M. Zaharia, O. Khattab, L. Chen et al., “The Shift from Models to Compound AI Systems,” https://bair.berkeley.edu/blog/2024/02/18/ compound-ai-systems, 2024. [2] M. Gravara, A. Stanisic, and S. Nastic, “A novel compound AI model for 6G networks in 3D continuum,” in Proceedings of the 2025 Joint European Conference on Networks and Communications & 6G Summit (EuCNC/6G Summit). Poznan, Poland: IEEE, 2025, arXiv:2505.15821. [3] M. Gravara, J. L. Herrera, and S. Nastic, “Compass: Optimizing Compound AI Workflows for Dynamic Adaptation,” in 2026 IEEE 26th International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 2026, pp. 84–93. [4] O. Khattab, A. Singhvi, P. Maheshwari, Z. Zhang, K. Santhanam, S. Vardhamanan, S. Haq, A. Sharma, T. T. Joshi, H. Moazam, H. Miller, M. Zaharia, and C. Potts, “DSPy: Compiling declarative language model calls into self-improving pipelines,” in Proceedings of the International Conference on Learning Representations (ICLR), 2024. [5] Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang, “AutoGen: Enabling next-gen LLM applications via multiagent conversations,” in Proceedings of the First Conference on Language Modeling (COLM), 2024. [6] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” in Proceedings of the International Conference on Learning Representations (ICLR), 2023. [Online]. Available: https://arxiv.org/abs/2210.03629 [7] L. Chen, M. Zaharia, and J. Zou, “FrugalGPT: How to use large language models while reducing cost and improving performance,” Transactions on Machine Learning Research, 2024. [8] D. Ding, A. Mallick, C. Wang, R. Sim, S. Mukherjee, V. Ruhle, L. V. S. Lakshmanan, and A. Awadallah, “Hybrid LLM: Cost-efficient and qualityaware query routing,” in Proceedings of the International Conference on Learning Representations (ICLR), 2024. [9] S. Han, Z. Hu, A. D. Shah, H. Jin, Y. Yao, D. Stripelis, Z. Xu, and C. He, “Torchopera: A compound ai system for llm safety,” 2024. [Online]. Available: https://arxiv.org/abs/2406.10847 [10] M. C. Kaya, T. W. Pusztai, A. Stanisic, and S. Nastic, “Currus - A Compound AI Approach to Distributed Vehicle Trajectory Reconstruction in the Edge-Cloud,” in Proceedings of the 15th International Conference on the Internet of Things, ser. IOT ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 254–262. [Online]. Available: https://doi.org/10.1145/3770501.3770531 [11] M. Gravara, C. Marcelino, A. Stanisic, and S. Nastic, “PLAIground: SLO-driven runtime model selection for compound AI systems in the edge-cloud-space continuum,” in 2026 IEEE International Conference on Smart Computing Workshops and Other Affiliated events (SmartComp Companion), 2026, pp. 261–266. [12] M. Gravara, A. Stanisic, and S. Nastic, “Design Methodology and Performance Trade-offs Management for Distributed and Compound AI Systems,” in 2026 IEEE 19th International Conference on Cloud Computing (CLOUD). IEEE, 2026, pp. 376–387. [13] S. Ghafouri, K. Razavi, M. Salmani, A. Sanaee, T. Lorido-Botran, L. Wang, J. Doyle, and P. Jamshidi, “IPA: Inference pipeline adaptation to achieve high accuracy and cost-efficiency,” Journal of Systems Research, vol. 4, no. 1, 2024. [14] S. Ahmad, H. Guan, and R. K. Sitaraman, “Loki: A system for serving ml inference pipelines with hardware and accuracy scaling,” in Proceedings of the 33rd International Symposium on High-Performance Parallel and Distributed Computing (HPDC). ACM, 2024, pp. 267–280. [15] J. Jiang, G. Ananthanarayanan, P. Bodik, S. Sen, and I. Stoica, “Chameleon: Scalable adaptation of video analytics,” in Proceedings of the ACM SIGCOMM Conference. ACM, 2018, pp. 253–266. [16] Nigade et al., “Jellyfish: Timely Inference Serving for Dynamic Edge Networks,” in 2022 IEEE Real-Time Systems Symposium (RTSS). Houston, TX, USA: IEEE, Dec. 2022, pp. 277–290. [17] S. Wu, P. Sarthi, S. Zhao, A. Lee, H. Shandilya, A. Mladenic Grobelnik, N. Choudhary, E. W. Huang, K. Subbian, L. Zhang, D. Yang, J. Zou, and J. Leskovec, “Optimas: Optimizing compound AI systems with globally aligned local rewards,” in Proceedings of the International Conference on Learning Representations (ICLR), 2026. [18] L. Chen, J. Q. Davis, B. Hanin, P. Bailis, M. Zaharia, J. Zou, and I. Stoica, “Optimizing model selection for compound ai systems,” arXiv preprint arXiv:2502.14815, 2025.

[19] J. Zhang, S. Elnikety, S. Zarar, A. Gupta, and S. Garg, “Model-switching: Dealing with fluctuating workloads in machine-learning-as-a-service systems,” in Proceedings of the 12th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud). USENIX Association, 2020. [20] D. Mendoza, F. Romero, and C. Trippel, “Model selection for latencycritical inference serving,” in Proceedings of the Nineteenth European Conference on Computer Systems (EuroSys). ACM, 2024, pp. 1016– 1038. [21] F. Romero, Q. Li, N. J. Yadwadkar, and C. Kozyrakis, “Infaas: Automated model-less inference serving,” in Proceedings of the 2021 USENIX Annual Technical Conference (USENIX ATC). USENIX Association, 2021, pp. 397–411. [22] S. Ahmad, H. Guan, B. D. Friedman, T. Williams, R. K. Sitaraman, and T. Woo, “Proteus: A high-throughput inference-serving system with accuracy scaling,” in Proceedings of the Nineteenth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), Apr. 2024, pp. 318–334. [23] Y. Wu, M. Lentz, D. Zhuo, and Y. Lu, “Serving and optimizing machine learning workflows on heterogeneous infrastructures,” Proceedings of the VLDB Endowment, vol. 16, no. 3, pp. 406–419, 2022. [24] D. Crankshaw, G.-E. Sela, X. Mo, C. Zumar, I. Stoica, J. E. Gonzalez, and A. Tumanov, “Inferline: Latency-aware provisioning and scaling for prediction serving pipelines,” in Proceedings of the 11th ACM Symposium on Cloud Computing (SoCC). ACM, 2020, pp. 477–491. [25] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, “Retrieval-augmented generation for knowledge-intensive nlp tasks,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 9459–9474. [Online]. Available: https://arxiv.org/abs/2005.11401 [26] W. Jiang, S. Subramanian, C. Graves, T. Kraska, and G. Alonso, “RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving,” in Proceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA). ACM, 2025, pp. 974–989. [27] S. Ray, R. Pan, Z. Gu, K. Du, S. Feng, G. Ananthanarayanan, R. Netravali, and J. Jiang, “METIS: Fast quality-aware RAG systems with configuration adaptation,” in Proceedings of the 31st ACM Symposium on Operating Systems Principles (SOSP). ACM, 2025, pp. 606–622. [28] I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs with preference data,” in Proceedings of the International Conference on Learning Representations (ICLR), 2025. [29] R. Bao, N. Xue, Y. Sun, and Z. Chen, “Dynamic quality-latency aware routing for LLM inference in wireless edge-device networks,” in Proceedings of the 2025 IEEE/CIC International Conference on Communications in China (ICCC Workshops). IEEE, 2025, pp. 1– 6. [30] Q. J. Hu, J. Bieker, X. Li, N. Jiang, B. Keigwin, G. Ranganath, K. Keutzer, and S. K. Upadhyay, “RouterBench: A Benchmark for Multi-LLM Routing System,” 2024. [Online]. Available: https: //arxiv.org/abs/2403.12031 [31] S. Somerstep, F. M. Polo, A. F. M. de Oliveira, P. Mangal, M. Silva, O. Bhardwaj, M. Yurochkin, and S. Maity, “CARROT: A Cost Aware Rate Optimal Router for Multi-LLM Serving,” 2025. [Online]. Available: https://arxiv.org/abs/2502.03261 [32] D. Sikeridis, D. Ramdass, and P. Pareek, “PickLLM: Context-aware RL-assisted large language model routing,” in AI for Research and Scalable, Efficient Systems: AI4Research 2025 and SEAS 2025, Held in Conjunction with AAAI 2025, ser. Communications in Computer and Information Science. Springer, 2025, pp. 227–239. [33] L. Chen, J. Q. Davis, B. Hanin, P. Bailis, I. Stoica, M. Zaharia, and J. Zou, “Are more LLM calls all you need? Towards the scaling properties of compound AI systems,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 37, 2024. [34] A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark, “Selfrefine: Iterative refinement with self-feedback,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 36, 2023. [35] W. Wang, B. Bi, M. Yan, C. Wu, L. Bao, L. Peng, J. Si, and S. Wang, “MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers,” in Advances in Neural Information Processing Systems (NeurIPS), 2020.

[36] T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng, “MS MARCO: A human generated MAchine Reading COmprehension dataset,” in Proceedings of the Workshop on Cognitive Computation: Integrating Neural and Symbolic Approaches (CoCo@NIPS), ser. CEUR Workshop Proceedings, vol. 1773, 2016. [37] Gemma Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière et al., “Gemma 3 technical report,” arXiv preprint arXiv:2503.19786, 2025. [38] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle et al., “The Llama 3 Herd of Models,” arXiv preprint arXiv:2407.21783, 2024. [39] J. Johnson, M. Douze, and H. Jégou, “Billion-scale similarity search with GPUs,” IEEE Transactions on Big Data, vol. 7, no. 3, pp. 535–547, 2021. [40] J. Forrest, T. Ralphs, S. Vigerske et al., “COIN-OR branch-and-cut solver (CBC),” https://github.com/coin-or/Cbc, 2024. [41] P. Rajpurkar, R. Jia, and P. Liang, “Know What You Don’t Know: Unanswerable Questions for SQuAD,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), 2018. [42] D. Crankshaw, X. Wang, G. Zhou, M. J. Franklin, J. E. Gonzalez, and I. Stoica, “Clipper: A low-latency online prediction serving system,” in Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI). USENIX Association, 2017, pp. 613–627. [43] A. Gujarati, R. Karimi, S. Alzayat, W. Hao, A. Kaufmann, Y. Vigfusson, and J. Mace, “Serving DNNs Like Clockwork: Performance Predictability from the Bottom Up,” in Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI). USENIX Association, 2020, pp. 443–462. [44] H. Shen, L. Chen, Y. Jin, L. Zhao, B. Kong, M. Philipose, A. Kr-

ishnamurthy, and R. Sundaram, “Nexus: A GPU Cluster Engine for Accelerating DNN-Based Video Analysis,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP). ACM, 2019, pp. 322–337. [45] J. R. Gunasekaran, C. S. Mishra, P. Thinakaran, B. Sharma, M. T. Kandemir, and C. R. Das, “Cocktail: A multidimensional optimization for model serving in cloud,” in Proceedings of the 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI). USENIX Association, 2022, pp. 1041–1057. [46] Z. Zhao, Y. Hu, G. Yang, Z. Gong, C. Shen, L. Zhao, W. Li, X. Liu, and W. Qu, “SLOpt: Serving real-time inference pipeline with strict latency constraint,” IEEE Transactions on Computers, vol. 74, no. 4, pp. 1431–1445, 2025. [47] B. Hu, L. Xu, J. Moon, N. J. Yadwadkar, and A. Akella, “Mosel: Inference serving using dynamic modality selection,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. [48] C. Chang, E. Lo, and C. Ye, “Biathlon: Harnessing model resilience for accelerating ML inference pipelines,” Proceedings of the VLDB Endowment, vol. 17, no. 10, pp. 2631–2640, 2024. [49] Y. Zhang, X. Zhang, G. Ananthanarayanan, A. Iyer, Y. Shu, V. Bahl, Z. M. Mao, and M. Chowdhury, “Vulcan: Automatic query planning for live ml analytics,” in Proceedings of the 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI). USENIX Association, 2024. [50] F. Romero, M. Zhao, N. J. Yadwadkar, and C. Kozyrakis, “Llama: A heterogeneous and serverless framework for auto-tuning video analytics pipelines,” in Proceedings of the ACM Symposium on Cloud Computing (SoCC), 2021.

Record · ID 660789 · SHA-256 9269d848ba04785e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.