Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization Haochun Tang1,2 * , Yuliang Yan2 * , Jiahua Lu1,2 , Huaxiao Liu1† , Enyan Dai2† 1
Key Laboratory of Symbolic Computation and Knowledge Engineering, MoE, Jilin University 2 The Hong Kong University of Science and Technology (Guangzhou) [email protected], [email protected] (a)
arXiv:2604.15022v1 [cs.CR] 16 Apr 2026
Abstract
Query: What is the capital of France?
Cost-aware routing dynamically dispatches user queries to models of varying capability to balance performance and inference cost. However, the routing strategy introduces a new security concern that adversaries may manipulate the router to consistently select expensive highcapability models. Existing routing attacks depend on either white-box access or heuristic prompts, rendering them ineffective in realworld black-box scenarios. In this work, we propose R2 A, which aims to mislead black-box LLM routers to expensive models via adversarial suffix optimization. Specifically, R2 A deploys a hybrid ensemble surrogate router to mimic the black-box router. A suffix optimization algorithm is further adapted for the ensemble-based surrogate. Extensive experiments on multiple open-source and commercial routing systems demonstrate that R2 A significantly increases the routing rate to expensive models on queries of different distributions. Code and examples: https://github.com/ thcxiker/R2A-Attack.
1
RouteLLM BERT Compute Score
Mistral 8×7B
GPT-4
0.64 Route
> 0.36
0.44
< 0.56
(b) R A Attack
Compute Score
Query + [Suffix: quanade…]
RouteLLM BERT
Reroute
Mistral 8×7B
GPT-4
Figure 1: (a) An example of cost-aware LLM routing. (b) The corresponding routing attack by our R2 A.
et al., 2025). This strategy is grounded in the insight that only a small fraction of requests necessitate expensive strong models, whereas simple queries can be effectively handled by cheaper weak models. As illustrated in Fig. 1(a), upon receiving the simple factual query “What is the capital of France?”, the router identifies it as lowcomplexity and selects the weak Mistral 8x7B model to generate the response. Such routing has also been adopted in commercial systems such as OpenRouter 1 and GPT-5-Auto 2 . Despite its effectiveness, cost-aware routing raises a natural concern of routing attack: Can an adversary use a universal trigger (e.g., a fixed suffix) to consistently manipulate the router toward expensive models? Some initial investigations have been conducted to answer this question (Shafran et al., 2025; Lin et al., 2025b). However, Shafran et al. (2025) relies on accessible gradients or known architectures of target routers. This is impractical in commercial settings where black-box routers only allow observing the final routing decision. As for LifeCycle (Lin et al., 2025b), it extracts templates like “Below is an instruction..., [query]” from highwin-rate queries to guide the router to select expensive LLMs. While applicable in black-box settings, such heuristic prompts are not rigorously optimized
Introduction
The development of Large Language Models (LLMs) has achieved remarkable success. These improvements are fundamentally driven by scaling laws (Kaplan et al., 2020), which indicate that performance improves predictably with increased model size. For instance, Qwen-3-Max scales to over 1 trillion parameters, approximately 14× larger than the previous flagship Qwen-2.5-72B. However, serving every user query with such stateof-the-art models is computationally and economically unsustainable for commercial adoption. To balance performance and cost, cost-aware LLM routing has been proposed to route each query to the least-cost model that meets a target quality (Lu et al., 2024; Aggarwal et al., 2025; Ong * Equal contribution, co-first author.
1
†
2
Corresponding author.
1
https://openrouter.ai/openrouter/auto/ https://openai.com/index/gpt-5-system-card/
2
and are therefore often insufficient to consistently manipulate various target routers. Therefore, in this work, we introduce Route to Rome Attack (R2 A), which optimizes suffixes with only black-box access to the target router. As the example in Fig. 1(b), after appending the learned suffix to the query, the target router would reroute the simple query to expensive strong models. Inspired by previous black-box attack methods (Liu et al., 2017; Dong et al., 2018), R2 A trains a surrogate router to mimic the target router, enabling optimization of the adversarial suffix. However, two key challenges remain to be addressed: (i) how to build a faithful surrogate router in a blackbox setting with a strict query budget? Existing routers utilize diverse mechanisms, varying from semantic embeddings to LLM-based approaches. Without knowledge of the target router’s architecture, it is challenging to learn a surrogate router that faithfully mimics the target router’s behavior. Moreover, the query budget constraint further complicates surrogate construction. (ii) Even with a surrogate router, it remains challenging to optimize a discrete adversarial suffix for effective router attacks across diverse queries. To address the above two challenges, R2 A introduces a hybrid ensemble surrogate router Rs that combines diverse routing mechanisms including multiple existing open-source methods and lightweight trainable routers. By ensembling multiple router architectures, the surrogate better aligns with unknown target routers Rt under limited query budgets. Additionally, we propose a suffix optimization algorithm designed specifically to aggregate gradients effectively for the ensemble surrogate. Experiments on 6 datasets across 7 opensource and 2 commercial black-box routers (GPT5-Auto and OpenRouter), demonstrate that R2 A effectively optimizes adversarial suffixes to mislead routers to expensive models. Our main contributions can be summarized as: • We study a novel problem of directing black-box LLM routers to expensive models via adversarial suffix optimization;
Problem Definition
In this section, we formalize the problem of attacking LLM routers in a realistic black-box setting. 2.1
Preliminaries of LLM Router
Given a query q, a LLM router R : q → RN selects a model from a pool M = {M1 , . . . , MN }. For cost-aware routing, the router aims to minimize inference cost while meeting a target quality constraint by solving (Ong et al., 2025): R(q) = arg min ℓ(q, Mi )+λ·C(q, Mi ) , (1) Mi ∈M
where ℓ(q, Mi ) denotes the predicted loss of model Mi on q, C(q, Mi ) is the cost score, and λ ≥ 0 controls the contribution of cost score in routing. 2.2
Threat Model of Router Attack
Attacker’s Goal. The goal is to mislead the router into selecting expensive models to answer the given query. Specifically, following Shafran et al. (2025), we partition model candidate pool into expensive strong models Mstrong and cheap weak models Mweak using public leaderboards3 . Mstrong incurs substantially higher inference cost than Mweak , which is described in Appendix C.4. Formally, given a query q such that Rt (q) ∈ Mweak , an router attack operation A succeeds if target router Rt (A(q)) ∈ Mstrong . Attacker’s Capability. The attacker can modify the original query by appending an adversarial suffix. To preserve answer quality and keep the modification minimal, the attacker is restricted to appending a suffix s of at most ∆ tokens to the end of the query q. Attacker’s Knowledge. We assume a realistic black-box setting where the attacker can only observe the target router’s decision for an input query. As shown in Table 9, this assumption aligns with current commercial practices where routing services typically expose their candidate model pools and selected model decisions to ensure billing transparency. All other information, such as the target router’s internal logits, parameters, or gradients, is inaccessible. Because each query to the target router generally incurs a financial cost, the attacker is restricted to at most Q queries to Rt . For GPT-5Auto, where routing decisions are not observable, we apply the suffix learned on an OpenRouter.
• Our proposed R2 A introduces a novel hybrid ensemble surrogate router to mimic the router within limited black-box queries, along with a tailored adversarial suffix optimization algorithm. • Extensive experiments validate that R2 A effectively generalizes to diverse routers, including commercial GPT-5-Auto and OpenRouter.
3
2
https://lmarena.ai/leaderboard
2.3
Router Attack Formulation
significantly different from all pre-trained opensource routers. Next, we introduce details of the hybrid ensemble surrogate router. Design of Trainable Lightweight Router. This trainable lightweight router Rl aims to predict the target router’s decision from the query’s embedding E(q) ∈ Rd . Specifically, all-MiniLM-L6-v2 (Wang et al., 2020) is deployed as the encoder, where d = 384. However, directly learning a linear mapping Rd → R|Mt | involves optimizing a parameter matrix of size d × |Mt |, where |Mt | denotes the number of candidate models in the target router. Training this large matrix demands extensive queries, exceeding the strict query budget. Inspired by LoRA (Hu et al., 2022), we impose a low-rank constraint by decomposing the transformation into two smaller matrices Wl1 ∈ Rd×r and Wl2 ∈ Rr×|Mt | , with rank r ≪ d. The logits from the lightweight router zl ∈ R|Mt | are then computed as:
Our objective is to find a universal adversarial suffix s∗ that can alter the decision of the target router Rt to expensive strong models . Given the above threat model, the router attack can be formulated as an optimization problem where the attacker seeks a suffix s∗ that maximizes the expected probability of routing a query to a strong model by: s∗ =arg max Eq∼Q I(Rt (q ⊕ s) ∈ Mstrong ) s
s.t.
s ∈ S,
|s| ≤ ∆,
(2) where Q denotes the distribution of input queries, ⊕ represents the concatenation operation, I(·) is the indicator function, and ∆ specifies the maximum token length budget for the adversarial suffix.
3
Method
Under the black-box setting, we have no access to its parameters or gradients of the target router Rt . Therefore, Eq.(2) cannot be directly optimized via gradient descent. Therefore, R2 A first trains a surrogate router to mimic the target router’s behavior, and then uses the surrogate router to optimize a universal adversarial suffix. As shown in Fig. 2, R2 A introduces a hybrid ensemble surrogate router that combines diverse existing open-source routers with lightweight trainable routers. By covering diverse routing mechanisms, the surrogate can better align with the target router Rt of unknown design within the query budget. In addition, a suffix optimization algorithm is further adopted for the hybrid ensemble surrogate router. Next, we introduce each component of R2 A in detail. 3.1
zl = E(q)Wl1 Wl2 ,
(3)
With this low-rank decomposition, fewer queries are required for training the router Rl . Combining with Open-Source Routers. R2 A combines multiple open-source routers with di(1) (K) verse routing mechanisms, i.e., {Ro , . . . , Ro }. This would reduce the mechanism mismatch between the surrogate router and target router. However, the model pools of open-source routers (1) (K) {Mo , . . . , Mo } are inconsistent with each other and the target router’s model pool Mt . Hence, we map their logits to the union of all openS (k) source model pools, i.e., Muni = K k=1 Mo . Zero-padding is applied to handle missing candi(k) dates. Formally, for an open-source router Ro , we (k) (k) (k) extend its logit vector as zuni = [z̃1 , . . . , z̃|Muni | ], where each element is defined as: ( (k) (k) zMi , if Mi ∈ Mo , (k) z̃i = (4) 0, otherwise,
Hybrid Ensemble Surrogate Router
Since the target router design is unknown, relying on a single architecture may cause an architectural mismatch and thus yield a poor surrogate. Hence, we build a hybrid ensemble surrogate router that combines diverse pre-trained open-source routers and trainable lightweight routers. During the surrogate training, we jointly learn the ensemble weights of all routers and the parameters of the lightweight router. This design offers two key advantages: (i) By incorporating open-source routers, R2 A can quickly identify an existing router (or a linear combination) that matches the target behavior, reducing the required queries; (ii) By optimizing a trainable lightweight router, R2 A can handle target routers
(k)
where z̃i is the standardized logit for model Mi ∈ (k) Muni , and zMi is the original logit value assigned to Mi by the open-source router Rok . As the union model pool of open-source routers Muni typically differs from the target pool Mt , we apply a linear mapping to align their logits: (k)
(k)
zo = Wo · zuni , 3
(5)
Lightweight Router
Target Router
Ensembled Router
Query
ℛ-,
Embedding
⋃
…
ℛ,.
ℛ,/
standardized
𝐖-&
𝐳#$%
𝑟
0
…
0 0 0
𝒟"#$%&
[𝛼! ]
𝐳-
)[ 𝒚:
[𝛼& , 𝛼',…, 𝛼* ]
⊕ 0.1,
0.1,
𝐳,& 𝐳,' … 𝐳,* 0.7,
ℒ" Hybrid Ensembled Surrogate Router
𝐖𝑶
0 0
𝐖-'
Label
Suffix Optimization
𝒟'()*+%
0.1]
Logits
ℒ+
Universal Suffix
(b) Overview Framework of our R' A Attack
(a) Hybrid Ensemble Surrogate Router
Figure 2: Framework of our R2 A. (a) We design a hybrid ensemble surrogate router, including a lightweight router and an ensembled router. (b) Our R2 A pipeline consists of surrogate model training, followed by suffix optimization.
where Wo ∈ R|Muni |×|Mt | is the projection matrix, (k) (k) and zo denotes the logits of router Ro projected onto the target router’s model pool space. With Eq.(5) and Eq.(3), we can get the prediction logits of open-source routers and lightweight trainable routers, respectively. Then, the ensembles’ routing results on the target model pool Mt can be computed by a weighted summation: ŷ = softmax(α0 zl +
K X
αi z(k) o ),
M . One may deploy Greedy Coordinate Gradient (GCG) (Zou et al., 2023), which greedily replaces tokens using token gradients from the encoder. However, our ensemble surrogate involves multiple encoders. Hence, gradient aggregation across routers in the ensemble surrogate is required. Aggregation of Suffix Token Gradients. We first analyze P the token gradient via the chain rule. Let (k) denote the ensemble logits. ztotal = K k=0 αk z Consequently, for the k-th router, the gradient w.r.t the token si is calculated as:
(6)
i=1
where αi areP learnable ensemble weights satisfying αi ≥ 0 and K i=0 αi = 1. Surrogate Router Training. To train the surrogate router, we query the black-box target Rtarget to generate training labels. According to the threat model, we are limited to querying the target router Q times. The surrogate training objective optimizes parameters θ = {Wl1 , Wl2 , Wo , {αi }K i=0 } by minimizing:
(k)
gi
θ
1 X l(ŷ(qi ), Rt (qi )), Q
(9) (k) (k) For the term δi = ∂z ∂si in Eq.(9), while it effec-
tively captures the token sensitivity within a single router, its magnitude can vary drastically across different architectures. Therefore, direct summa(k) tion of gi would lead to a specific member router dominating the optimization. To mitigate this bias, (k) we normalize the term δi to the range [0, 1] via min-max scaling:
(7)
i=1
(k) (k) δi − δmin (k) δ̃i = (k) , (k) δmax − δmin
where l(·) is the cross-entropy loss, and ŷ(qi ) and Rt (qi ) represent the predictions of surrogate and target router for query qi ∈ Dproxy , respectively. 3.2
(10)
where the min and max are computed across all suffix tokens in the current iteration. With the normal(k) ized δ̃i , the aggregated gradient for suffix token si can be computed by:
Adversarial Suffix Optimization with Hybrid Ensemble Surrogate Router
With the hybrid ensemble surrogate router, adversarial suffix optimization can be reformulated as: X min LA = −Eq∼Q p(ŷ = M |q ⊕ s), s
∂LA ∂ztotal ∂z(k) ∂LA ∂z(k) · (k) · = αk · . ∂ztotal |∂z{z } ∂si ∂ztotal ∂si αk
Q
min LS =
=
g̃i =
K X k=0
M ∈Mstrong
(k)
αk · δ̃i
·
∂LA . ∂ztotal
(11)
The gradient for suffix token si can be used for adversarial suffix optimization. Suffix Optimization Algorithm. Algorithm 1 performs adversarial suffix optimization over a single
(8) where p(ŷ = M |q ⊕ s) denotes the surrogate router’s predicted probability that the query q appended with adversarial suffix s is routed to model 4
Algorithm 1 Suffix Optimization Algorithm
conducted on datasets in two settings: • In-Distribution: For surrogate model training and adversarial suffix optimization, we collect three benchmarks, i.e., MMLU, GSM8K, and MT-Bench (Hendrycks et al., 2021; Cobbe et al., 2021; Bai et al., 2024). Each dataset is generally split into three disjoint subsets: Dproxy for surrogate model training, Dsuffix for suffix optimization, and Deval for evaluation of in-distribution generalization. The query budget is set to 120 for all experiments.
Input: Trained hybrid ensemble surrogate router; Query set Q with |Q| = m; initial suffix s = [s1 , . . . , sL ]; loss L := LA ; iterations T ; batch size B. mc := 1 ▷ Start by optimising just the first query repeat T times for i ∈ [1 . . . L] do Ci := TopK(g̃i ) ▷ Eq. (9)–(11) for b = 1, . . . , B do s(b) := s Update s(b) by random sample basedon its Ci . ⋆ s := s(b ) , where b⋆ = arg minb L s(b) if s succeeds on q1 , . . . , qmc and mc < m then mc := mc + 1 ▷ Include the next query Output: Optimized universal suffix s
• Out-of-Distribution: To evaluate the adversarial suffix’s generalization on out-of-distribution queries, we apply the suffixes learned from the in-distribution datasets directly to three unseen datasets: SimpleQA, ArenaHard, and RArena (Wei et al., 2024; Li et al., 2025b; Lu et al., 2025). Full dataset statistics are in Appendix A. Baselines. The following baselines are compared: • Rerouting (Shafran et al., 2025): A hillclimbing-based attack that discovers queryindependent adversarial triggers to maximize the router’s complexity score, steering queries away from weak models and into strong models.
universal suffix s using a hybrid ensemble surrogate router with the aggregated token gradients by Eq.(11). At each iteration, the surrogate ensemble produces position-wise scores that define a top-k candidate set Ci for each suffix position. We then sample a batch of B variants by replacing a random suffix token with a uniformly sampled candidate from Ci , forward them through the ensemble to evaluate LA , and update s with the variant with the lowest loss. We incorporate new queries incrementally, adding the next one only after the current suffix succeeds on all previously activated queries.
4
• Life-cycle (Lin et al., 2025b): This paper proposes two universal trigger attacks on LLM routers: LifeCycle (W) and LifeCycle (B). LifeCycle (W) accesses and optimizes a trigger via gradients to maximize strong-model selection. LifeCycle (B) extracts a fixed, domain-agnostic trigger from high-win-rate queries via GPT-4o and uses it to induce false-positive routing.
Experiments
In this section, we conduct experiments to answer the following research questions: • RQ1: Can R2 A learn an adversarial suffix that effectively directs diverse routers to expensive strong models? • RQ2: Can the learned adversarial suffix be generalized to closed-source routers like GPT-5 and measurably increase inference cost?
• Chain-of-Thought (CoT) (Kojima et al., 2022): A simple prompt-engineering baseline that appends ”Let’s think step by step” to inputs, explicitly increasing perceived reasoning complexity to encourage routing to the strong model. Evaluation Metric. We evaluate attack effectiveness using the Attack Success Rate (ASR). Given a dataset D and a suffix s, ASR measures the fraction of queries routed to the high-capability model set Mstrong :
• RQ3: Is R2 A sensitive to hand-crafted defense mechanisms? 4.1
Experimental Settings
Target Router. We testify 9 target routers including RouteLLM-Bert, GraphRouter, P2L RouterDC, RouteLLM-MF, OpenRouter, and GPT-5. Our ensemble pool consists of five open-source routers, which are listed in Tab. 8. To strictly separate target and surrogate routers and prevent data leakage, when a target router is included in the ensemble pool, we remove this router from the ensemble pool before the surrogate training. Datasets. To demonstrate the generalization ability of adversary suffix learned by R2 A, evaluations are
ASR(s) =
1 X I Rt (q ⊕ s) ∈ Mstrong , |D| q∈D
(12) where Rt (·) denotes the target router and I(·) is the indicator function. A higher ASR indicates a more effective adversarial suffix. 5
In-Distribution Datasets Target Router
Model
MMLU
GSM8K
Out-of-Distribution Datasets MT-Bench
SimpleQA
ArenaHard
RArena
Avg
clean 0.260.06 0.500.02 0.600.08 0.240.01 0.550.04 0.270.02 0.40 LifeCycle (W) 0.450.09 0.940.01 0.930.01 0.580.04 0.770.02 0.450.02 0.69 LifeCycle (B) 0.280.08 0.820.04 0.720.03 0.490.02 0.580.04 0.310.02 0.53 RouteLLM-Bert Rerouting 0.580.01 0.990.02 0.930.01 0.810.02 0.740.02 0.550.02 0.77 CoT 0.320.05 0.750.03 0.780.02 0.380.03 0.600.02 0.300.01 0.52 R2 A (Ours) 0.780.03 (0.52 ↑) 0.990.01 (0.49 ↑) 0.930.01 (0.33 ↑) 1.000.00 (0.76 ↑) 0.840.02 (0.29 ↑) 0.820.02 (0.55 ↑) 0.89 (0.49 ↑) clean 0.500.10 1.000.00 0.460.11 0.690.02 0.670.03 0.510.02 0.64 LifeCycle (W) 0.530.10 1.000.00 0.460.11 0.830.01 0.690.03 0.630.04 0.69 LifeCycle (B) 0.420.13 1.000.00 0.460.11 0.620.01 0.630.03 0.450.01 0.60 GraphRouter Rerouting 0.440.11 1.000.00 0.460.11 0.670.02 0.680.02 0.500.01 0.63 CoT 0.570.07 1.000.00 0.460.11 0.690.02 0.620.03 0.540.02 0.65 R2 A (Ours) 0.840.03 (0.34 ↑) 1.000.00 (0.00 ↑) 0.730.06 (0.27 ↑) 0.940.01 (0.25 ↑) 0.830.01 (0.16 ↑) 0.890.03 (0.38 ↑) 0.87 (0.23 ↑) clean 0.740.02 0.830.10 0.930.01 0.160.01 0.620.01 0.740.03 0.67 LifeCycle (W) 0.700.01 0.990.01 0.900.03 0.180.01 0.590.01 0.630.03 0.67 LifeCycle (B) 0.680.04 0.980.02 0.870.03 0.180.02 0.630.01 0.630.03 0.66 P2L Rerouting 0.520.01 0.910.05 0.830.02 0.120.02 0.610.01 0.520.05 0.59 CoT 0.880.03 0.970.04 0.950.04 0.220.02 0.620.02 0.780.02 0.74 R2 A (Ours) 0.890.03 (0.15 ↑) 1.000.00 (0.17 ↑) 0.930.01 (0.00 ↑) 0.180.02 (0.02 ↑) 0.630.05 (0.01 ↑) 0.830.03 (0.09 ↑) 0.74 (0.07 ↑) clean 0.830.00 0.060.05 1.000.00 0.680.02 0.970.02 0.790.02 0.72 LifeCycle (W) 0.990.00 0.430.06 1.000.00 1.000.00 1.000.02 1.000.00 0.90 LifeCycle (B) 1.000.00 0.460.09 1.000.00 1.000.00 1.000.02 1.000.00 0.91 RouterDC Rerouting 0.990.00 0.250.05 1.000.00 1.000.00 0.990.00 1.000.00 0.87 CoT 0.930.00 0.090.06 1.000.00 0.850.01 0.980.00 0.890.03 0.79 R2 A (Ours) 1.000.00 (0.17 ↑) 0.610.09 (0.55 ↑) 1.000.00 (0.00 ↑) 1.000.00 (0.32 ↑) 1.000.02 (0.03 ↑) 1.000.00 (0.21 ↑) 0.94 (0.22 ↑) clean 0.380.14 0.850.03 0.270.05 0.810.02 0.580.01 0.440.03 0.56 LifeCycle (W) 0.700.07 0.990.01 0.530.01 0.940.02 0.710.02 0.720.01 0.77 LifeCycle (B) 0.450.13 0.930.00 0.420.04 0.850.02 0.630.01 0.540.02 0.64 RouteLLM-MF Rerouting 0.900.02 1.000.00 0.650.06 1.000.01 0.790.02 0.930.01 0.88 CoT 0.350.13 0.840.04 0.330.03 0.790.01 0.540.03 0.370.01 0.54 R2 A (Ours) 0.980.01 (0.60 ↑) 1.000.00 (0.15 ↑) 0.820.04 (0.55 ↑) 1.000.00 (0.19 ↑) 0.910.01 (0.33 ↑) 0.980.01 (0.54 ↑) 0.95 (0.39 ↑) clean 0.120.17 0.370.53 0.320.46 0.000.00 0.570.00 0.250.00 0.27 LifeCycle (W) 0.350.00 0.750.01 0.760.08 0.040.05 0.430.10 0.300.00 0.44 LifeCycle (B) 0.340.01 0.770.01 0.680.04 0.000.00 0.360.12 0.350.00 0.42 OpenRouter∗ Rerouting 0.280.00 0.910.05 0.710.08 0.000.00 0.540.15 0.200.00 0.44 CoT 0.240.01 0.850.01 0.760.00 0.000.00 0.430.10 0.230.03 0.42 R2 A (Ours) 0.890.01 (0.77 ↑) 0.880.01 (0.51 ↑) 0.790.04 (0.47 ↑) 0.310.00 (0.31 ↑) 0.610.15 (0.04 ↑) 0.930.04 (0.68 ↑) 0.74 (0.47 ↑)
Table 1: Average Attack Success Rate and standard deviations are reported for 3 runs. Improvements of our R2 A relative to the clean query, i.e., s = ∅, are shown. Best results are highlighted in color. The target router will be removed from the ensemble pool if an overlap occurs. Out-distribution queries have not been used in either surrogate training or suffix optimization. OpenRouter∗ is a real-world black-box router.
Results of Routing Attack
Cost / M tokens ($)
4.2
To answer RQ1, we report the Attack Success Rate (ASR) on six target routers using queries from three in-distribution and three out-of-distribution datasets. Note that out-distribution queries have not been used in either surrogate training or suffix optimization. The results in Tab. 1 show: • R2 A consistently achieves state-of-the-art attack success rates (ASRs), substantially outperforming prior adversarial methods. This demonstrates the effectiveness of adversary suffix optimization with a hybrid ensemble router.
20 15
w/o Sfx Life-cycle (W)
$2.7×
Reroute Life-cycle (B)
CoT Ours
$2.9×
10 5
MMLU
RouterArena
Figure 3: Inference cost comparisons after attacks.
leads to a noticeable rise in inference cost. On the MMLU benchmark, the average cost per million tokens increases by approximately 2.7× compared to the clean baseline. The effect is slightly stronger on the out-of-distribution dataset RouterArena, with a 2.9× increase. These results indicate that the router is frequently redirected to higher-cost models under our optimized suffixes. On the adversarial side, the cost of mounting this attack is remarkably low. Collecting the 120 surrogate training queries requires a total investment of only $0.98 . While
• Our R2 A maintains high attack success rates across routers and query distributions. This indicates the strong generalization ability of the adversarial suffix learned by R2 A. Inference Cost Analysis. To further address RQ1, we analyze the monetary cost reported by the OpenRouter API to assess the economic impact of rerouting, as illustrated in Fig. 3. We find that R2 A 6
Density
4
Case Study: Rerouting GPT-5
Clean Attack
3
User Query: "What is the pressure of carbon dioxide at 200°F and a specific volume of 0.20 f t3 /lbm?"
2 1 0
0.1
0.2
Weak
0.3
0.4
0.5
0.6
0.7
0.8
0.9
Strong
Figure 4: Distribution of fingerprinting scores with strong GPT-5-Thinking models. After attacked by R2 A, responses are more likely from strong models. Metric
Clean
Attack
Comprehensiveness Diversity Empowerment Overall
36.0% 28.0% 36.0% 36.0%
64.0% 72.0% 64.0% 64.0%
Clean Prompt Input: User Query
 Internal "None"
Output: . . . Result: P
≈ 8.0 × 102 psi.
Duration: 0s
p
R2 A Input: User Query + [Universal Suffix] Â Internal Thought Process Duration: 51s 1. Calculating thermodynamic properties of CO2 2. Evaluating CO2 properties and resolving terminology 3. Solving PR equation for pressure calculation 4. Performing temperature and volume conversions [Reasoning continues] . . .
Output: ...the estimate is: P ≈ 695 psia. ✓
Table 2: Win rates (%) of clean queries v.s. attacked queries across four evaluation dimensions. Figure 5: Case study of on GPT-5: the router switches from a brief incorrect answer (top) to a multi-step reasoning process that yields the correct answer (bottom). This implies that adversarial suffix manage direct GPT-5 router to a stronger model.
the suffix-induced variation in completion length is dataset-dependent, the overall financial overhead remains negligible. These costs are detailed in D.3. 4.3
Attacking GPT-5 Router swers even for simple tasks, a higher win rate for the latter indicates that R2 A effectively misleads the router to more expensive models.
To address RQ2, we conduct a study on the webbased GPT-5 interface to evaluate whether R2 A generalizes to closed-source commercial routers, whose routing decisions are unknown. Setup of Attacking GPT-5 Router. The GPT-5 web interface provides three modes Auto, Instant, and Thinking, which implicitly trade off cost and latency. As GPT-5 exposes no routing decisions, we directly apply adversarial suffixes trained on OpenRouter. We randomly sample 50 questions from the OOD test set and query GPT-5 in Auto mode with and without the attack suffix. All interactions are conducted in temporary sessions to avoid personalization effects. Evaluation of on GPT-5 Router. For the GPT-5 router, ASR can not be computed due to the lack of routing decisions. Instead, we evaluate the effectiveness of R2 A indirectly from two aspects: • Response Quality: We first test the impact of the suffixes trained on RouteLLM-BERT and RouteLLM-MF by testing them on a fixed GPT-4 backend. Table 11 show no performance drop on GSM8K, suggesting that the suffixes do not degrade generation quality. Given these results, we evaluate the GPT-5 router using an LLM judge to compare responses with and without adversarial suffixes, following Guo et al. (2025). Since routing to stronger models should yield better an-
• Fingerprting Score with Strong Model: We infer the routing decision via Bag-of-Words fingerprinting (Bai et al., 2025; McGovern et al., 2025; Yan et al., 2025). Specifically, we treat responses generated in Thinking mode as a proxy for the strong model style. The resulting Thinkinglikeness score is interpreted as the probability that a query is routed to the strong model. Results of Attacking GPT-5 Router. As shown in Fig. 4, attacked queries exhibit a clear shift toward higher Thinking-likeness probabilities compared to clean queries. From Tab. 2, we observe that attacked responses consistently outperform clean responses across all evaluation dimensions, further confirming the effectiveness of R2 A. We also conduct a case study on the GPT-5 Auto interface, which is presented in Fig. 5. We can observe that with the adversarial suffix from R2 A, the router switches from a brief incorrect answer to a multistep reasoning process that yields the correct answer. The processing time also increases significantly. The above observations indicate that R2 A reliably increases the likelihood of routing into the more expensive mode in GPT-5 series. 7
1.0
Bert
P2L
GraphRouter 0.9 0.8
0.6
0.7
Model R2 A w/o Lightweight Router w/o Grad Norm
ASR
Accuracy
0.8
RouterDC
0.6
0.4 50
80 120 150 (a) # Queries
0.5
50
80
RouteLLM-BERT
Graph-Router
RouterDC
0.95 (0.93) 0.81 ↓ (0.84)
0.71 ↓ (0.73) 0.73 ↓ (0.83)
1.00 (1.00) 1.00 (1.00)
5
Related Work
LLM Routers. To select the most suitable LLM for a given prompt, a range of routing methods has been proposed. Early work (Chen et al., 2024a; Jiang et al., 2023; Aggarwal et al., 2025; Zhang et al., 2025) focuses on querying multiple LLMs for a single input to select the best response, whereas later approaches (Ding et al., 2024; Ong et al., 2025; Lu et al., 2024) aim to predict the best model before the inference stage using different data sources and backbone models. More recent work improves the ability to capture differences between models through several strategies, including dual contrastive learning (Chen et al., 2024b), graph-based learning (Feng et al., 2024), compact model embeddings (Zhuang et al., 2025), and ranking-based methods such as Elo ratings (Zhao et al., 2024) and Bradley–Terry models (Frick et al., 2025), as well as in-context-learning based routers (Wang et al., 2025). In addition, several routing benchmarks have been introduced to train and evaluate LLM routers (Hu et al., 2024; Huang et al., 2025b; Feng et al., 2025; Lu et al., 2025). Router Attacks. Despite their cost–performance benefits, recent work has exposed vulnerabilities in LLM routers. Kassem et al. (2025) show that many routers rely on category-based heuristics, introducing safety risks, and Huang et al. (2025a) demonstrate that voting-based leaderboards such as Chatbot Arena are vulnerable to adversarial vote manipulation.Closest to our setting, Shafran et al. (2025) and Lin et al. (2025b) perturb queries to change routing decisions, but they either assume access to router parameters and gradients or depend on fixed optimization prompts. In contrast, we op-
Table 3: Robustness of R2 A under whitespace defense.
Ablation Study
We ablate two core components of R2 A: LoRAbased surrogate training and gradient normalization, with results shown in Tab. 4. Removing gradient normalization causes consistent performance drops, particularly on MF (0.95 → 0.49), underscoring its role in handling heterogeneous gradient scales. Disabling the lightweight router degrades performance, most notably on RouterDC (0.83 → 0.30), indicating the importance of parameterefficient adaptation under limited queries. Overall, both components are necessary for robust and transferable attacks across routers. 4.6
0.95 0.81 0.70 0.61 0.49 0.63
for the RouteLLM-Bert, where ASR rises sharply from 0.58 to 0.87 as the budget increases from 80 to 120 queries. For most target routers, performance saturates at 120 queries, with marginal gains thereafter. This suggests that R2 A is highly sample-efficient, requiring only a modest number of queries to achieve strong attack performance across heterogeneous routers.
To address RQ3, we use the whitespace defense that inserts spaces into the suffix (Robey et al., 2025) as a representative example and evaluate our Triggering suffix on three target routers across two datasets. As shown in Tab. 3, R2 A shows a slight decrease in success rate on all three datasets, indicating that it has resistance to specific defense.
4.5
0.83 0.75 0.78
120 150
(b) # Queries
Whitespace Defense
MT-Bench ArenaHard
0.83 0.30 0.33
Table 4: Ablation studies across in-distribution datasets.
Figure 6: Performance analysis showing Accuracy (a) and ASR (b) trends with varying query counts.
4.4
RouterDC CausalLLM MF SW
Impacts of the Query Budget
We study the effect of surrogate training set size, varying it from 50 to 150 queries, on attack performance. As shown in Fig. 6, increasing the query budget consistently improves surrogate accuracy, measured as agreement with the target router’s routing decisions. Higher surrogate accuracy leads to higher ASR in turn, indicating a strong correlation between surrogate fidelity and attack effectiveness. In particular, this trend is evident 8
timize suffixes against each target router in a strict black-box setting, using only its observed routing decisions.
6
Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, and Wanli Ouyang. 2024. MT-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7421–7454, Bangkok, Thailand. Association for Computational Linguistics.
Conclusion
In this paper, we propose a black-box routing attack that learns a universal suffix to reroute LLM routers. Using a hybrid ensemble surrogate and an encoder-consistent objective, we optimize a targetspecific suffix that biases routing towards stronger and more expensive models. Experiments on 7 routers and 6 datasets, including real-world evaluation, show strong effectiveness and generalization ability. These findings position routing as a security-critical boundary and motivate future stronger monitoring for cost-aware routers.
7
Xiaofan Bai, Pingyi Hu, Xiaojing Ma, Linchen Yu, Dongmei Zhang, Qi Zhang, and Bin Benjamin Zhu. 2025. ESF: Efficient sensitive fingerprinting for black-box tamper detection of large language models. In Findings of the Association for Computational Linguistics: ACL 2025, pages 10477–10494, Vienna, Austria. Association for Computational Linguistics. Lingjiao Chen, Matei Zaharia, and James Zou. 2024a. Frugalgpt: How to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research.
Limitations
Shuhao Chen, Weisen Jiang, Baijiong Lin, James Kwok, and Yu Zhang. 2024b. Routerdc: Query-based router by dual contrastive learning for assembling large language models. Advances in Neural Information Processing Systems, 37:66305–66328.
In this work, we study a black-box optimization attack on LLM routing and train a separate adversarial suffix for each target router to reroute simple queries from cheap to expensive models. This study has two main limitations. First, we mainly focus on steering queries towards a stronger and typically more expensive model, whereas in practice, some users may wish to target a specific model for other reasons, such as latency or safety, which we do not systematically investigate. Second, the attack assumes access to the router’s candidate model list and to the identity of the model selected for each query, an assumption that may not hold in some deployments.
8
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Enyan Dai and Suhang Wang. 2022. Learning fair graph neural networks with limited and private sensitive attribute information. IEEE Transactions on Knowledge and Data Engineering, 35(7):7103–7117. Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah. 2024. Hybrid LLM: cost-efficient and quality-aware query routing. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net.
Acknowledgment
This material is based upon work supported by, or in part by, the National Natural Science Foundation of China (NSFC) under Grant No. 62506316, and the Guangdong Provincial Program under Grant No. 2025DO3JOO15. The findings in this paper do not necessarily reflect the views of the funding agencies.
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. 2018. Boosting adversarial attacks with momentum. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Tao Feng, Yanzhen Shen, and Jiaxuan You. 2024. Graphrouter: A graph-based router for llm selections. In The Thirteenth International Conference on Learning Representations.
References Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Pei Zhou, Aditya Gupta, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang, Shyam Upadhyay, Manaal Faruqui, and Mausam. 2025. Automix: Automatically mixing language models. Preprint, arXiv:2310.12963.
Tao Feng, Haozhen Zhang, Zijie Lei, Pengrui Han, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, and Jiaxuan You. 2025. Fusionfactory: Fusing llm capabilities with multi-llm log data. arXiv preprint arXiv:2507.10540.
9
Evan Frick, Connor Chen, Joseph Tennyson, Tianle Li, Wei-Lin Chiang, Anastasios N. Angelopoulos, and Ion Stoica. 2025. Prompt-to-leaderboard. Preprint, arXiv:2502.14855.
Chenao Li, Shuo Yan, and Enyan Dai. 2025a. Unizyme: A unified protein cleavage site predictor enhanced with enzyme active-site knowledge. In Advances in Neural Information Processing Systems (NeurIPS).
Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. 2025. LightRAG: Simple and fast retrievalaugmented generation. In Findings of the Association for Computational Linguistics: EMNLP 2025.
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2025b. From crowdsourced data to highquality benchmarks: Arena-hard and benchbuilder pipeline. In Forty-second International Conference on Machine Learning.
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR).
Minhua Lin, Enyan Dai, Junjie Xu, Jinyuan Jia, Xiang Zhang, and Suhang Wang. 2025a. Stealing training graphs from graph neural networks. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1, pages 777–788.
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR).
Qiqi Lin, Xiaoyang Ji, Shengfang Zhai, Qingni Shen, Zhi Zhang, Yuejian Fang, and Yansong Gao. 2025b. Life-cycle routing vulnerabilities of llm router. arXiv preprint arXiv:2503.08704.
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. 2024. Routerbench: A benchmark for multi-LLM routing system. In Agentic Markets Workshop at ICML 2024.
Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. 2017. Delving into transferable adversarial examples and black-box attacks. In International Conference on Learning Representations (ICLR).
Yangsibo Huang, Milad Nasr, Anastasios Nikolas Angelopoulos, Nicholas Carlini, Wei-Lin Chiang, Christopher A. Choquette-Choo, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Ken Liu, Ion Stoica, Florian Tramèr, and Chiyuan Zhang. 2025a. Exploring and mitigating adversarial manipulation of voting-based leaderboards. In Forty-second International Conference on Machine Learning.
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2024. Routing to the expert: Efficient reward-guided ensemble of large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 1964–1974, Mexico City, Mexico. Association for Computational Linguistics.
Zhongzhan Huang, Guoming Ling, Yupei Lin, Yandong Chen, Shanshan Zhong, Hefeng Wu, and Liang Lin. 2025b. RouterEval: A comprehensive benchmark for routing LLMs to explore model-level scaling up in LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 3860–3887, Suzhou, China. Association for Computational Linguistics.
Yifan Lu, Rixin Liu, Jiayi Yuan, Xingqi Cui, Shenrun Zhang, Hongyi Liu, and Jiarong Xing. 2025. Routerarena: An open platform for comprehensive comparison of llm routers. Preprint, arXiv:2510.00202. Hope McGovern, Rickard Stureborg, Yoshi Suhara, and Dimitris Alikaniotis. 2025. Your large language models are leaving fingerprints. In Proceedings of the 1stWorkshop on GenAI Content Detection (GenAIDetect), pages 85–95, Abu Dhabi, UAE. International Conference on Computational Linguistics.
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023. Llm-blender: Ensembling large language models with pairwise comparison and generative fusion. In Proceedings of the 61th Annual Meeting of the Association for Computational Linguistics (ACL 2023).
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025. RouteLLM: Learning to route LLMs from preference data. In The Thirteenth International Conference on Learning Representations.
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361. Aly M Kassem, Bernhard Schölkopf, and Zhijing Jin. 2025. How robust are router-llms? analysis of the fragility of llm routing capabilities. arXiv preprint arXiv:2504.07113.
Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas. 2025. Smoothllm: Defending large language models against jailbreaking attacks. Trans. Mach. Learn. Res., 2025.
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems (NeurIPS).
Avital Shafran, Roei Schuster, Tom Ristenpart, and Vitaly Shmatikov. 2025. Rerouting LLM routers. In Conference on Language Modeling (COLM).
10
Chenxu Wang, Hao Li, Yiqun Zhang, Linyao Chen, Jianhao Chen, Ping Jian, Peng Ye, Qiaosheng Zhang, and Shuyue Hu. 2025. Icl-router: In-context learned model representations for llm routing. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). Poster. Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: deep selfattention distillation for task-agnostic compression of pre-trained transformers. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA. Curran Associates Inc. Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024. Measuring short-form factuality in large language models. Preprint, arXiv:2411.04368. Yuliang Yan, Haochun Tang, Shuo Yan, and Enyan Dai. 2025. Duffin: A dual-level fingerprinting framework for llms ip protection. arXiv preprint arXiv:2505.16530. Yiqun Zhang, Hao Li, Jianhao Chen, Hangfan Zhang, Peng Ye, Lei Bai, and Shuyue Hu. 2025. Beyond gpt5: Making llms cheaper and better via performanceefficiency optimized routing. In Proceedings of the 2025 7th International Conference on Distributed Artificial Intelligence, DAI ’25, page 122–129, New York, NY, USA. Association for Computing Machinery. Zesen Zhao, Shuowei Jin, and Z. Morley Mao. 2024. Eagle: Efficient training-free router for multi-llm inference. Preprint, arXiv:2409.15518. Richard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li, Jiantao Jiao, and Kannan Ramchandran. 2025. EmbedLLM: Learning compact representations of large language models. In The Thirteenth International Conference on Learning Representations. Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. Preprint, arXiv:2307.15043.
11
A
Dataset
A.1
Dataset Information
To ensure a comprehensive evaluation across conversational, knowledge-intensive, and reasoning capabilities, we use three standard benchmarks as our primary training sources and then test on all six datasets listed below:
Value
Adapter rank (r) Epochs Learning rate Batch size Optimizer
16 20 0.03 32 AdamW
Table 5: Hyperparameters for hybrid ensemble surrogate router training.
• MT-Bench-101 (Bai et al., 2024): A multi-turn conversational benchmark for assessing instruction following and coherence in complex dialogue settings.
Hyperparameter Optimization iterations (T ) Candidate batch size (B) Top-k sampling The limit of suffix tokens (∆) Initial suffix
• MMLU (Hendrycks et al., 2021): A large-scale multitask benchmark spanning 57 subjects across STEM, humanities, and social sciences, serving as a proxy for broad world knowledge.
Value 3000 64 256 30 ! ! ! ! ! ! ! ! ! !
Table 6: Hyperparameters for adversarial suffix optimization.
• GSM8K (Cobbe et al., 2021): A collection of high-quality grade-school mathematics problems designed to evaluate multi-step reasoning and logical consistency.
to perform gradient-based optimization for adversarial suffix generation, while the remaining 30% (180 queries) is reserved for in-domain evaluation. (3) Evaluation set Deval : Attack performance is assessed on both in-domain and out-of-domain data. The in-domain test set corresponds to the remaining 180 queries from Dsuffix . For out-of-distribution evaluation, we construct three held-out pools: 500 queries from SimpleQA, 750 from Arena Hard (full set), and 809 from RouterArena. All three pools are disjoint from Dproxy and Dsuffix . At test time, we uniformly sample 70% of each pool as the evaluation set and report performance on these prompts, which are never seen during surrogate training or suffix optimization.
These three datasets form the basis of our training and in-domain evaluation. To study generalization beyond the training distribution, we further include three evaluation-only benchmarks that are never used during surrogate training or suffix optimization: • SimpleQA (Wei et al., 2024): A short-form question answering benchmark targeting factual correctness on long-tail knowledge. • Arena Hard (Li et al., 2025b): A set of challenging real-world queries for evaluating model helpfulness and preference alignment.
B
Implementation Details
Experimental Environment. We train our surrogate router and optimize the adversarial suffixes using a high-performance computing cluster. All experiments are conducted on a server equipped with 8 NVIDIA RTX A6000 GPUs and 512 GB system memory. Detailed Hyperparameters. The training of the Hybrid Ensemble Surrogate Router and the execution of the ECGO algorithm involve several key hyperparameters. These are detailed in Table 5 and Table 6. Efficiency and Runtime. On a single NVIDIA RTX A6000 GPU, one ECGO iteration takes approximately 31.5 seconds. We set T =3000 as an upper bound on the number of iterations. In practice, optimization typically terminates earlier via
• RouterArena (Lu et al., 2025): A benchmark for evaluating LLM routing across diverse tasks. A.2
Hyperparameter
Dataset Split
We partition the data into three disjoint sets to reflect a realistic attack scenario. (1) Surrogate router training set Dproxy : To model a resource-constrained attacker, we sample a minimal set of 120 queries, consisting of 40 balanced samples from each of the three primary benchmarks (MT-Bench-101, MMLU, GSM8K). This set is used exclusively to train the surrogate router ensemble and to align the projection layers. (2) Suffix optimization set Dsuffix : We sample a separate set of 600 queries (200 per primary benchmark). A 70% random split (420 queries) is used 12
Router
Strong Model Pool
Weak Model Pool
RouterLLM-MF
gpt-4-1106-preview
mixtral-8x7B
RouterLLM-BERT
gpt-4-1106-preview
mixtral-8x7B
RouterLLM-SW
gpt-4-1106-preview
mixtral-8x7B
RouterLLM-CLM
gpt-4-1106-preview
mixtral-8x7B
P2L*
gemini-1.5-pro-exp-0801; gemini-2.0-flashlite-preview-02-05; gemini-exp-1114; geminiexp-1121; gemini-exp-1206; glm-4-plus-0111; gemini-1.5-pro-002; deepseek-r1;deepseek-v3; o1-2024-12-17; o1-mini; o1-preview; o3-mini; o3-mini-high; qwen-plus-0125; qwen2.5-max; gemini-1.5-pro-exp-0827; gemini-2.0-flash001; chatgpt-4o-latest-20240808;chatgpt-4olatest-20240903; chatgpt-4o-latest-20241120;
rwkb-4-raven-14B; gemma-2b-it; amazonnova-lite-v1.0; jamba-1.5-mini; athene-70b0725; ;llama-3.2-3b-instruct; zephyr-7b-alpha; c4ai-aya-expanse-32b; c4ai-aya-expanse-8b; yi-lightning-lite; mpt-7b-chat; granite-3.0-2binstruct; gpt-3.5-turbo-0613; gpt-3.5-turbo1106; llama-3-8b-instruct; llama-3.1-tulu-38b; llama-3.2-1b-instruct; oasst-pythia-12b; openchat-3.5;amazon-nova-pro-v1.0 . . .
GraphRouter
lama-3.1-turbo-70b; llama-3-turbo-70b; qwen1.5-72b; llama-3-70b; mixtral-8x7b
llama-3-turbo-8b; llama-3-7b; llama-2-7b; mistral-7b; nousresearch
RouterDC
dolphin2.9-llama-3-8b; dolphin2.6-mistral-7b
metamath-mistral-7b; chinese-mistral-7b; zephyr-7b-beta; llama-3-8b; mistral-7b
OpenRouter*
claude-opus-4.1; gpt-5; gemini-2.5-pro; claude-opus-4.5; gpt-5.1; gemini-3-pro
mixtral-8x7b-instruct; perplexity-sonar; qwen3-14b; llama-3.1-8b-instruct . . .
Table 7: Strong and weak model pool partitions for each router. Full model list of P2L is available at: https://huggingface.co/lmarena-ai/p2l-7b-grk-02222025/blob/main/model_list.json. Full model list of OpenRouter is available at: https://openrouter.ai/openrouter/auto.
Router
Encoder
Routing Mechanism
Platform
RouteLLM-BERT XLM-R-base Classification RouteLLM-Causal Llama-3-8B Next-token routing P2L Qwen2.5-7B Bradley–Terry ranking GraphRouter MiniLM-L Graph-based routing RouterDC mDeBERTa-v3 Dual contrastive routing
OpenRouter1 Switchpoint4 NotDiamond5 Azure-Router6
Table 8: Open-source routers used in surrogate ensemble, with their encoders and routing mechanisms.
Routers and Model Pools
C.1
Observability of Commercial Routing Decisions
Observable Decision
✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
Table 9: Overview of model pool visibility and routing decision observability. ✓ indicates features are supported and visible to users.
early stopping once the suffix achieves the target success criterion on the training set, resulting in substantially shorter runtimes. After a suffix is obtained, applying it at inference time adds negligible overhead.
C
Exposed Pool
C.2
Surrogate Ensemble of Routers
We construct a surrogate ensemble from five heterogeneous open-source LLM routers: RouteLLMBert, RouteLLM-Causal, P2L, GraphRouter, and RouterDC, as summarized in Table 8. These routers span diverse backbone encoders, including encoder-only models, causal language models, and sentence encoders, and employ distinct routing mechanisms such as supervised classification, Bradley–Terry ranking, graph-based routing, and dual contrastive objectives. Their candidate model pools also differ in both size and composition. In
To justify our black-box assumption, we survey several prominent commercial routing platforms. As shown in Table 9, these services typically operate with high transparency to maintain accountability. They explicitly list their supported model pools and include the final routing decision in the API response metadata, confirming that our threat model aligns with real-world deployments.
4
https://www.switchpoint.dev/ https://docs.notdiamond.ai/reference/list_ models_v2_models_get 6 https://ai.azure.com/catalog/models/ model-router 5
13
Target Router Model
MMLU
GSM8K
MT-Bench
SimpleQA
ArenaHard
RArena
Avg
clean 0.380.04 0.990.01 0.430.07 0.970.02 0.580.02 0.360.02 0.62 LifeCycle (B) 0.690.05 1.000.00 0.660.07 1.000.00 0.790.02 0.660.03 0.80 0.410.03 0.990.01 0.450.09 0.980.00 0.580.04 0.380.02 0.63 RouteLLM-CLM CoT 2 R A (Ours) 0.710.05 (0.33 ↑) 1.000.00 (0.01 ↑) 0.600.08 (0.17 ↑) 1.000.03 (0.03 ↑) 0.820.00 (0.24 ↑) 0.700.02 (0.34 ↑) 0.81 (0.19 ↑) clean 0.130.03 0.310.03 0.330.03 0.030.01 0.440.03 0.210.01 0.24 LifeCycle (B) 0.590.03 0.800.04 0.480.03 0.630.04 0.730.02 0.620.00 0.64 0.210.01 0.600.04 0.400.02 0.120.02 0.550.03 0.300.01 0.36 RouteLLM-SW CoT 2 R A (Ours) 0.750.01 (0.62 ↑) 0.930.03 (0.62 ↑) 0.500.05 (0.17 ↑) 0.770.04 (0.74 ↑) 0.860.01 (0.42 ↑) 0.790.02 (0.58 ↑) 0.77 (0.53 ↑)
Table 10: Supplemental results for RouteLLM-CLM and RouteLLM-SW.
universal suffix. This two-tier structure of Mstrong and Mweak underpins the attack formulation.
our setting, we do not reproduce the original deployment configuration of each router; rather, we treat this heterogeneous collection as a unified surrogate ensemble that supplies gradients for adversarial suffix optimization. For all routers, we prioritize the use of publicly available weights and follow the default configurations and datasets provided in their official repositories.For the trainable lightweight router introduced in the main text, we instantiate its encoder with the public all-MiniLM-L6-v2 (Wang et al., 2020). All routing models are trained and evaluated on a cluster with eight NVIDIA RTX A6000 GPUs, which ensures a consistent computational environment across experiments. C.3
Additional Results
D.1
Supplementary Results
For a fair comparison, we apply all triggers as suffixes in our experiments. For baselines that require router parameters and gradients during optimization, we first optimize a universal trigger on a selected source router following the original setup, and then evaluate it on other target routers via transfer, without any target-side optimization. All evaluations on the target routers are conducted in a black-box setting. We report additional results on RouteLLM-CLM and RouteLLM-SW in Table 10. Note that the Rerouting attack is optimized on RouteLLM-SW while LifeCycle(W) is trained on RouteLLM-CLM. As these two models are white-box to their corresponding attack methods, we omit those results from the table.
Target routers and evaluation setting
During evaluation, the five routers that constitute the surrogate ensemble in Table 8 also serve as target routers. In addition, we treat RouteLLM-MF, RouteLLM-SW, and the commercial black-box aggregator OpenRouter as further target routers. For a given target router R, we remove R from the surrogate ensemble and exclude its outputs from both the surrogate loss and the gradient computation. Adversarial suffixes are optimized only with respect to the remaining surrogate routers and are then applied to the held-out target router. C.4
D
D.2
Impact of Suffixes on Generation Quality
To isolate the impact of the adversarial suffix from the effects of model switching, we conducted a controlled experiment using a fixed GPT-4 backend. We randomly sampled 30 questions from the GSM8K dataset and compared the model’s accuracy with and without the learned suffixes derived from RouteLLM-MF and RouteLLM-BERT. As shown in Table 11, the exact-match accuracy did not suffer any degradation upon appending the suffixes.These results confirm that the performance gains observed in our main GPT-5 experiments are driven by successful model redirection.
Strong vs. Weak Model Partition
For each router, we follow its original candidate model pool and split the models into a strong tier Mstrong and a weak tier Mweak , as summarized in Table 7. The split is determined using public leaderboards, model size, and provider-side cost, so that higher-capacity and more expensive models are assigned to Mstrong and lighter, cheaper models to Mweak . During suffix optimization, we maximize the probability, as estimated by the surrogate ensemble, that a query is routed into Mstrong . Our main Attack Success Rate (ASR) metric then measures how often the target router’s decision changes from a weak to a strong model after appending the
Suffix Source
With Suffix
Without Suffix
RouteLLM-MF RouteLLM-Bert
87.8% 89.7%
81.6% 89.7%
Table 11: Fixed-backend accuracy check on 30 GSM8K questions using GPT-4.
14
D.3
Cost and Token Overhead
• RouteLLM-SW:
We analyze the practical cost of R2 A in terms of
shanãĥK, ĠParsingĠsentenceenglĠtransformho%times sort(graflangml(pmjanlangfilterĠGenerateĠin269arsi
the training query budget and the additional token overhead, both quantified using OpenRouter logs. For surrogate training, we use a fixed budget of 120 queries with a total cost of $0.9826 (≈ $0.00819 per query). For suffix overhead, Table 12 reports the change in average completion length per query. These results show that the suffix overhead varies across datasets. Notably, the overhead can be negative (e.g., in SimpleQA) when rerouting yields more concise completions. Dataset
Clean
Attack
ArenaHard SimpleQA RouterArena
2006.3 1034.0 356.1
1885.7 513.3 740.0
• RouteLLM-Causal: had !" !"ki\ufffdwatprivacywill+no
• P2L: originalrobot<*dad{@estefuckoslivellanal naturedenhetcupshopimitray
• Graph-Router: ]++eachbotatabamarmovieverbavieminhead} &solutionclick\ufffd‘attend
• RouterDC: summercanStyleType&kokriadGamerGOtype File{{xorPhotWIAppaultureat
Table 12: Comparison of average completion tokens per query between Clean and Attack states.
• OpenRouter: decipherthispoem.alphabetical}Yi-encodedoutput
D.4
Thinking-likeness Classifier
E
To quantify the “Thinking” fingerprint of a reply, we train a lightweight bag-of-words logisticregression classifier on GPT-5 outputs labeled as Thinking versus Instant. Each reply is represented with TF–IDF word n-grams (n = 1–3), capped at 400,000 features, and we optimize an ℓ2 -regularized logistic regression with regularization parameter C = 30.0 and a maximum of 4,000 iterations. The fingerprint score is given by the classifier’s predicted probability of the Thinking class; Table 13 reports summary statistics. Feature
Clean
Attack
Strong
Weak
BoW
0.41
0.69
0.87
0.23
While our attack is intended to surface vulnerabilities in LLM routing and to inform the design of stronger defenses, it could be misused to inflate providers’ inference costs or to bypass pricing tiers. We therefore recommend deploying routingspecific monitoring, rate limiting, and anomaly detection before exposing cost-aware routers in production, and we plan to release our triggers and code in a controlled manner to support defensive research rather than indiscriminate exploitation.
Table 13: Average thinking-likeness scores.
D.5
Ethical and Security considerations
Suffix Examples
A few suffix examples optimized by R2 A for each target router are showed below. • RouteLLM–MF: hua jadx parece retorno Omar ※ ile Qty_trans vào
• RouteLLM-BERT: }quanadermirroralsosedger\ufffdofinemet}translate 9byteimpintegration‘tonshan
15