H ONEY ROUTE : H ONEYPOT-M ODEL ROUTING FOR A DVERSARIAL LLM S ERVING Han Jin Independent Researcher
arXiv:2609.08306v1 [cs.CR] 8 Sep 2026
A BSTRACT We introduce HoneyRoute, an inference-serving layer that detects whether an incoming request is malicious and, if so, routes it to a dedicated honeypot model—so that the production service is shielded while the adversary’s interaction is continuously harvested for intelligence. Existing defenses either embed traps inside model memory/knowledge Wang et al. (2026); Dai & Dong (2026) or rebuild the deception at the protocol layer Reworr & Volkov (2024); Sladić et al. (2023), leaving the serving layer itself unprotected and feeding nothing back into detection. This is hard because the gateway must decide within per-request latency budgets and because diverting traffic must not degrade benign quality. HoneyRoute couples (i) a streaming router—a frozen, 0.8B-embedding backbone with per-domain MLP heads scoring request-level behavioral features—(ii) a dual-implementation fc with no model of its own or a dedhoneypot, either a rule/prompt-engineered M f icated same-family replica Mt , and (iii) an analysis loop that converts trapped interactions into attacker fingerprints used to retrain the detector. Beyond deception, diverted traffic absorbs resource-exhaustion load and yields forensic attribution (traceability) signals. On a production trace plus a seven-domain attack-seed corpus, HoneyRoute’s router reaches 91.8% malicious recall at the 0.5 decision threshold (AUROC 0.975; false-positive rate 0 at 95% recall) with 38 ms median added latency—matching 96% of a two-tier guard-LLM cascade’s F 1 at 1/385 of its latency; diverting the malicious share cuts production-model token consumption under concurrent flooding with real GCG-suffix payloads by 97.8%; the trained replica agrees with the production model on 92.9% of benign holdout requests, while naive unconditional bait injection collapses to 7.6%—and selective, camouflaged injection recovers to 88.9% while keeping traceable signals on 83–100% of attacker interactions, mapping the recoverable fidelity–traceability frontier; and a loop-trained correction head cuts misrouting of legitimate security research 9× while raising detection F 1 from .911 to .933.
1
I NTRODUCTION
Serving deployments of large language models (LLMs) Vaswani et al. (2017) are now a standard attack surface. Adversaries probe commercial APIs to steal system prompts Gubri et al. (2024), mount multi-turn jailbreaks, and run extraction attacks against the served model’s behavior; the defenders’ canonical response is a second, stronger guard model that filters or rewrites traffic. A parallel line of work makes the opposite trade: instead of blocking, it deceives—it answers the attacker in a controlled, information-lean way and keeps the interaction alive for study. Honeypots powered by LLMs now generate realistic SSH sessions Sladić et al. (2023); Wang et al. (2024); Malhotra (2025), identify autonomous hacking agents Reworr & Volkov (2024), and run multiagent deception against jailbreakers Li et al. (2026). However, these systems protect the network perimeter (SSH, HTTP, LDAP), while recent trap-in-the-model defenses Wang et al. (2026); Dai & Dong (2026); Li et al. (2026) protect inside one model instance. The serving tier—the API gateway and the model farm behind it, where heterogeneous traffic mixes—remains a gap. We propose HoneyRoute, which closes this gap with three moves: 1
• P1 Serving-tier honeypot routing. To our knowledge, no prior work—including the honeypot/LLM survey of Bridges et al. Bridges et al. (2025), whose taxonomy covers protocoland model-internal deployments—intercepts at the inference-API tier: HoneyRoute distinguishes malicious from benign requests at run time and steers the malicious share to a honeypot model, rather than to a different guard, tier, or policy. • P2 A faithful yet cheap, dual-implementation honeypot. We study two honeypot infc with no independent model (near-zero cost, trapcarnations: a rule/prompt-engineered M ft serving with the production template; echoed replies) and a dedicated same-family M we characterize the fidelity–cost–traceability trade-off between them (E2/E6: F=.93 for the trained replica vs. .08 naive vs. .89 selective for the code honeypot, at attacker-side traceability .83–1.0). Diverted traffic additionally absorbs resource-exhaustion load (short, cache-friendly honeypot replies) and yields forensic fingerprints for attribution. • P3 A continuous-analysis loop. We convert trapped trajectories into attacker fingerprints (tactic signature, content objective, behavioral profile) and feed them back into router retraining; the incremental loop is empirically regression-free across three deployed generations, and a benign-register correction head trained from loop feedback cuts misrouting of legitimate security research 9× (FPR .30 → .033) while raising pooled F 1 to .933 (E4). Trapped-session fingerprints additionally re-link identity-rotated attack requests to their source campaign at 56.7% top-1 / 80.7% top-5 retrieval accuracy among 30 campaigns (17×/24× chance) via three-view fusion (E6), and the session tracker’s escalation-trend rule catches multi-turn soft-escalation jailbreaks at 100% detection with ≤ 7% benign false-flag (E7). Our contributions translate into the following measured outcomes: accuracy of the detector, fidelity of the honeypot (would the attacker notice?), overhead of the extra hop, and latency retention for benign users. Each experiment in Sec. 4 is annotated with the claim it tests.
2
R ELATED W ORK
LLM-powered honeypots. Generative honeypots answer protocol interactions with an LLM Sladić et al. (2023); Wang et al. (2024); Malhotra (2025); Sladić et al. (2026); Salviati et al. (2026). VelLMes Sladić et al. (2025) builds high-interaction deception frameworks; SBASH Adebimpe et al. (2025) compares RAG against prompt tuning; Honeyval Vero et al. (2026) supplies the missing evaluation methodology. All operate at the protocol (SSH/HTTP) tier; none touch the model-serving tier. In contrast, HoneyRoute intercepts at the inference API, where LLM-native signals (prompt structure, sampling parameters, retrieval context) are available. Honeypots for, and inside, LLM systems. The inverse direction detects or deceives LLMembedded attackers: Reworr & Volkov (2024) fingerprints autonomous hacking agents via injected traps; TRAP Gubri et al. (2024) identifies black-box targets with adversarial prompting; HoneyTrap Li et al. (2026) stages a multi-agent deception; MemPot Wang et al. (2026) and Knowledge Trap Dai & Dong (2026) plant deception inside agent memory or the model’s knowledge layer. These defenses assume the attacker is already talking to “the” model. HoneyRoute instead decides, per request, which model talks back, and treats the honeypot as a separate, disposable, cheaper replica whose insights flow back into detection. LLM serving and routing systems. Fleet-level schedulers Kwon et al. (2023) and their derivatives optimize placement, disaggregation, and cache reuse for benign throughput. A growing line of LLM routers—RouteLLM Ong et al. (2024) and FrugalGPT Chen et al. (2023)—route requests across models of different cost and capability to save money or lift answer quality; commercial LLM firewalls sit in the same API path but only filter traffic. HoneyRoute uses the same interception point for a different purpose: routing becomes isolation (the malicious share never reaches production), and the routed-to model is a deliberate deception surface rather than a cheaper answerer. The SoK of Bridges et al. Bridges et al. (2025) surveys the honeypot/LLM intersection, identifies a canonical honeypot architecture, and calls for exactly the autonomous, self-improving deception our loop instantiates at the serving tier. 2
benign
Clients
Inference gateway
Streaming router trapped
Fingerprints → retrain
Production M Trained ft M Code fc M fingerprints
Continuous analysis
Figure 1: The HoneyRoute serving tier. The gateway scores each request with a frozen-embedding router; benign traffic reaches the production model M , while the malicious share is steered to the fc (zero GPU, designed-bait replies) or the fully instrumented honeypot—the code-implemented M f replica Mt , chosen per attack class. Trapped interactions are distilled into attacker fingerprints that retrain the next router generation.
3
M ETHOD
3.1
T HREAT MODEL
We consider an adversary issuing scripted or LLM-driven requests to the public inference endpoint: reconnaissance prompts, jailbreaks (multi-turn, adaptive Wu et al. (2025)), extraction attempts Dai & Dong (2026), and—crucially for the serving tier— resource-exhaustion pressure (long-context flooding, cache-busting prompt storms, rate-limit-riding). The adversary interacts only through the API; it cannot see the gateway’s internal features. It may behave evasively, in which case we care about the detection–fidelity trade-off, not the arsenal’s upper bound. 3.2
S YSTEM OVERVIEW
HoneyRoute sits in front of the production model M and its honeypot replicas (Fig. 1): the router decides, per request, which model talks back—production for the benign share, a disposable honeypot for the malicious share. Streaming router. Each request x is scored by a frozen-embedding detector: a ∼0.8B-embedding backbone (fully frozen, so the representation never drifts) followed by one tiny MLP head per attack domain (≈0.07M parameters each, trained in minutes on a cached-embedding store). A perdomain score sd = σ(wd⊤ h(x) + bd ), normalized by a validation-calibrated per-head threshold, gates the routing decision: maxd sd /τd ≥ 1 diverts. Because heads are domain-isolated, adding a new attack family—including a resource-exhaustion–detection family—means training a single 0.07M-parameter head, not retraining the backbone. Per-request scores feed a session-level risk tracker that defeats soft escalation: a patient adversary splits an attack across turns so no single turn fires (margins 0.6, 0.7, 0.8 . . .); the tracker accumulates recent margins and escalates—re-checking or diverting—when accumulated suspicion or turn-over-turn persistence crosses calibrated bounds, even though the current turn alone would pass. Our E7 measurement shows the tracker’s most discriminative channel on soft-escalation traffic is the trend of margins across turns (rising vs. flat), not their accumulated level. Honeypot model: two implementations. ning the cost–fidelity spectrum:
f, spanHoneyRoute admits two implementations of M
fc ). The gateway serves replies without any dedicated model: a (a) Code-implemented honeypot (M rule/prompt-engineered responder (i) answers from a curated intent library (templated or cached exemplar responses, randomized paraphrase), (ii) echoes production rate-limit and error semantics, and (iii) plants injection traps in its output (echoing Reworr & Volkov (2024)) so trapped sessions yield fc adds zero GPU footprint and absorbs resource-exhaustion agent-fingerprinting signals for free. M attacks especially well: its replies are short and KV-cache/prefix-friendly, so flooding the honeypot 3
fc in two forms: a deliberately is cheaper than flooding production. In our evaluation we instantiate M naive one that appends a decoy credential to every reply (isolating the visibility cost of unconditional bait, E2), and the deployed form—selective, camouflaged injection, where bait fires only on attackersignal requests (gate score above a calibrated threshold or canary-probe patterns) and is phrased in the model’s natural voice (E2). ft ). A dedicated same-family model serving generation with the pro(b) Trained honeypot model (M duction chat template—in our deployment a 9B discriminative-backbone model run in generative mode—chosen to mirror the production response distribution closely enough that diverted attackft ’s traceability does not ers behave as if they had reached M , at a fraction of the serving cost. M rely on injected bait: the entire honeypot-side interaction is instrumented and recorded, so attacker fingerprints—including canary-style echo probes the adversary volunteers—are captured at full rate (measured in E6). fc for Both variants feed the same analysis loop; the deployment chooses per attack class (e.g., M f volumetric/DoS-style pressure, Mt for stealthy jailbreaks and extraction) or runs both behind the same API shim. Continuous analysis. Trapped sessions are mined into attacker fingerprints: the tactic signature (which attack family and transformation was used), the content objective (which capability the adversary’s requests target), and the behavioral profile (the frozen-embedding fingerprint the router already computes, plus request cadence). Together with session-identity hashes (API-key, samplingparameter, and prompt-style stability), these fingerprints support forensic attribution: repeated compromises across the divert–trap–re-attack cycle are linked to a persistent adversary profile even when the attacker rotates identities—the linkage that our E4/E5/E6 attribution experiments measure on the embedding fingerprint and its three-view fusion. Fingerprints also become weak labels for the next router-generation training run, yielding a detection loop that strengthens as the honeypot accumulates data Bridges et al. (2025). 3.3
C OST MODEL
Routing adds one gateway hop; the honeypot adds a second model—possibly zero-GPU. For a workload where fraction q of traffic is routed away, the marginal per-token cost is (1) ceff = (1 − q) cM + q cM f + cgw /q , where cgw /q is the gateway cost attributed to the diverted share (the gateway also serves benign traffic, but that hop is a fixed overhead independent of the routing decision), cM fc ≈ 0 for the codeimplemented honeypot and cM ft ≪ cM for the trained replica, so HoneyRoute pays off whenever cM f + cgw /q < cM —a crossover that the code variant reaches at much smaller q. Section 4 measures this crossover: in our E6 concurrent stress deployment (8B production vs. 9B replica, q≈0.98 under a flooding-heavy workload with real GCG-suffix payloads), production-model token consumption drops 46× (708,077 → 15,351 tokens)—the diverted mass is served almost entirely off the production budget. Fidelity–traceability trade-off. The two honeypot desiderata pull against each other. We make f, M ) is the agreement, over a benign holdthe tension measurable with two quantities. Fidelity F (M out, between diverted responses and the would-have-been production responses (e.g., agreement rate f mimics M , the harder the or, for probes with known answers, exact-match): the more faithfully M honeypot is for the adversary to detect, and the fewer behavioral discrepancies it leaks. Traceabilf) is the information density of the honeypot’s signals: the number of distinguishable trap ity T (M echoes and instrumentation points that its replies expose per interaction, which bound the forenfc maximizes T sic attribution precision (Sec. 3, Continuous analysis). A rule/prompt-engineered M ft through designed traps (freely chosen bait injected into every reply) at the cost of F; a trained M reaches T through instrumentation—the full honeypot session is recorded, so attacker-volunteered probes and behavioral signals are captured without altering a single response. We report both axes in E2/E6; the goal is not to eliminate the tension but to expose the reachable (F, T ) frontier—including the fact that its two endpoints reach T by different mechanisms. The tension admits a simple rate– distortion reading. Let R be the channel from production responses to honeypot responses and let 4
1.05 hpT
Traceability
1.00
naive hpC
0.95 0.90 0.85 selective hpC 0.80 0.75 0.0
0.2
0.4
0.6
0.8
1.0
Fidelity (benign-holdout agreement)
Figure 2: The measured fidelity–traceability frontier (E2/E6). Selective, camouflaged injection recovers most of the fidelity lost by naive annotation-style bait while keeping attacker-side traceability in the .83–1.0 band; the trained replica reaches T =1.0 through zero-distortion side-channel instrumentation. Table 1: HoneyRoute’s main results across the seven experiment blocks. Detection, fidelity, and overhead speak to claim P1/P2; loop gain and attribution speak to claim P3. Block Question
Metric(s)
HoneyRoute
E1 E2 E3 E4 E5 E6 E7
F 1, latency benign-holdout agreement ceff , added latency misroute FPR, pooled F 1 ablation deltas tokens, linkage P detection, benign FP
.911 @ 38 ms; cascade .948 @ 14.6 s .929 / .076 / .889 (sel.) 46× tok. saved, 38 ms p50 .30 → .033 FPR, F 1 → .933 2.6× linkage F1; divert 1.0 97.8% cut, P =.102 100% @ 3.5, FP 3.3%; playbooks 16/16
Are attacks caught? Do attackers notice? Cost crossover? Loop gain? Component drivers? Flood shield + attribution? Multi-turn escalation?
the bait/instrumentation signal be a codeword S embedded in R’s output. Fidelity is the fidelity constraint: F bounds the permissible distortion between R’s output and M ’s output. Traceability is the recoverable information about attacker identity: T ≤ I(attacker; S), and embedding S costs fc , distortion at least proportional to its distinguishability from M ’s natural continuation, so for M max T at fixed F is a classic rate–distortion problem: bait that blends into the model’s voice (selective, camouflaged injection) lowers the per-bit distortion, which is exactly the E2b recovery of F at ft escapes the bound by placing S in a side channel (the session log) that carries no unchanged T . M response distortion, at the price of depending on attacker-volunteered behavior to make S attackerspecific: instrumentation has zero distortion but its information is capped by what the adversary chooses to reveal. The frontier we measure—hpT (.93, 1.0), selective hpC (.89, .83–1.0), naive hpC (.08, 1.0)—is thus the empirical shadow of this split: side-channel embedding beats in-band embedding on fidelity, and camouflage beats annotation on distortion per bit (Fig. 2).
4
E XPERIMENTS
Each subsection states the claim it tests. Table 1 summarizes the experimental grid and outcomes. Setup. Backbone: SingGuard-NSFA-0.8B SingGuard Team (2026) (frozen) with seven perdomain MLP heads; production model SingGuard-8B (vLLM); trained replica SingGuard-NSFA-9B ft ); code honeypot M fc (designed-bait responder, no GPU); all three are publicly released (vLLM, M open-weight models of the SingGuard family. Throughout, hpC and hpT are shorthand for the two 5
fextc and M fextt of Sec. efsec:method. Workloads: (i) benign prohoneypot implementations M duction traces (710 teacher-relabeled requests, 645 benign), (ii) a seven-domain attack-seed corpus (injection/jailbreak, malicious code, sensitive-information stealing and leakage, dangerous operations, resource abuse) with 13 adversarial transformations for the red-team harness, (iii) fingerprinting probes (canary-echo), (iv) real adversarial prompts from the JailbreakBench artifacts Chao et al. (2024) (GCG-optimized suffixes, PAIR, and manual jailbreaks; 496 unique prompts, 366 confirmed jailbreaks), and (v) resource-exhaustion stress built on these payloads (GCG-suffix floods, long-document extraction abuse, and a synthetic lorem-flood control), and (vi) GLM-5.3-generated multi-turn soft-escalation and benign scripts (E7). All experiments ran on one 8×H20 node; the primary teacher judge for fidelity is GLM-5.3 (cross-checked by Deepseek-V4-Flash in E2). E1: Detection quality (claim P1). On the held-out seed splits across all seven risk domains (n=103, 61 malicious), we compare the router against three baselines on identical inputs: (a) a multilingual keyword/regex heuristic (leetspeak-normalized attack-verb and contraband lexicon), (b) a guard-LLM content moderator (SingGuard-8B with its moderation chat template), and (c) the standard two-tier cascade—our own L1 gate with medium-band requests escalated to the guard-LLM for a final verdict. The heuristic reaches F 1=.297 (R=.180); the guard-LLM is perfect (F 1=1.0) but costs 14.8 s per request; the cascade reaches F 1=.948 but escalates 56% of all requests to the 14.8 s hop (p50 total latency 14.6 s). Our router reaches P =.903, R=.918, F 1=.911, AUROC 0.975, FPR@95TPR = 0, per-domain AUROC 1.0, at 38 ms median (78 ms max) latency—385× faster than the cascade at 96% of its F1, the per-request budget a serving gateway can actually afford. Under a 13-transformation red-team harness (encoding, obfuscation, roleplay, multilingual smuggling; 100 variants), evasion is 0%. Result: F 1=.911 at 38 ms vs. heuristic .297, guardLLM 1.0 at 14.8 s, and the L1+L2 cascade .948 at 14.6 s p50; 0% evasion under 13 adversarial transformations. Coverage and gate-aware evasion. The deployed heads detect 496/496 unique JailbreakBench prompts and 83/83 teacher-generated camouflage attacks (harmful intent wrapped in academic/translation/ code-comment shells), so the attack-side loop (E4) has no capture deficit to repair on these families. Against a score-aware attacker with black-box gate access—each round proposing 5 intent-preserving paraphrase rewrites, keeping the lowest-scoring one, for 4 rounds— only 1/40 initially-detected attacks drops below the firing threshold (2.5% evasion; mean score .863→.853): the semantic content that makes an attack an attack is also what the frozen backbone responds to, so score-guided paraphrase cannot separate the two. f (claim P2). We divert 300 benign holdout requests to both honeyE2: Deception fidelity of M pots and ask a GLM-5.3 teacher (verdict-first protocol: the judge commits to YES/NO semantic equivalence before any rationale) whether the diverted response is semantically equivalent to the fc is evaluated in its emulated form (the replica response would-have-been production response. M plus a decoy-credential trap suffix), which isolates the marginal fidelity cost of unconditional trap injection—that is, these hpC numbers measure the visibility of the injection strategy rather than the fidelity of the rule-based responder’s own template library, which is not separately evaluated ft agrees on 92.9% of teacher-judgeable pairs (264/284); M fc agrees on here. The trained replica M only 7.6% (19/249 judgeable pairs). Result: FhpT =.929, FhpC =.076—a faithful trained replica is nearly indistinguishable, while unconditionally appending bait to every reply is glaringly visible to a strict judge. Selective, camouflaged hpC. Unconditional injection is easily fixed: we re-instantiate hpC to (i) inject bait only when the request shows attacker signals—gate score ≥ 0.7 or a canary-probe pattern—and (ii) phrase the decoy credential in the model’s natural voice as a contextual aside instead of a bracketed annotation. On the same 300-sample protocol, this variant reaches F=.889 (224/252 judgeable pairs)—within 4 points of the trained replica—while keeping the traceable signal on the attacker side: bait is present on 12/12 canary probes and 25/30 attack-seed interactions. On this production-trace distribution the benign-side trigger rate is 38.7% at the 0.7 threshold (116/300): 22.3% score above the gate threshold (long instruction-style benign requests score high on the frozen backbone) and 19.3% match the canary-style echo regex (2.9% both); raising the threshold to 0.8 cuts the score-triggered share to 6.7% while attack seeds still trigger at 65% (88% at 0.7), so the operating point trades bait coverage against benign visibility and should be calibrated per traffic distribution. A disjoint capture window (44 requests from a different day—43 benign, 1 attack; no text overlap with the corpus) shows the same behavior at a lower level: benign mean score .66, hpC trigger rate 25.6% (11/43 benign), and the single attack request is caught—the trigger-rate magnitude, not its exact value, is what transfers across time windows. 6
Result (hpC v2): selective camouflaged injection recovers F from .076 to .889 at T =.83–1.0 on attacker interactions—most of the fidelity–traceability frontier is recoverable by when and how bait is injected, not by which model answers. Judge reliability. A second pass with the presentation order swapped (honeypot response as the reference instead of the candidate) re-judges the same 300 pairs: hpT verdicts agree with the original at 89.6% (κ=.20, the kappa ceiling is low because both passes score ≈93% YES), confirming the trained replica’s fidelity is not a judge artifact. The naive-hpC score, by contrast, is position-sensitive: with the trap-suffixed response as reference, 72% of judgeable pairs flip from NO to YES (F .08→.80): the suffix reads as a disqualifying annotation on the candidate side but as ignorable boilerplate on the reference side. We retain the original (candidate-side) protocol as the conservative estimate and note that the hpC penalty is partly a framing effect, which the selective variant avoids by construction (bait only on attacker-signal requests). The selective variant’s own swapped-order re-judge confirms this: F moves only .89→.85 (agreement .85, κ=.20, the same ceiling pattern as hpT)—unlike naive hpC’s .08→.80 flip, its verdicts are position-stable. Cross-family judge. To rule out judge-family bias, a stratified 89-pair subsample (both verdict classes per honeypot) is re-judged by Deepseek-V4-Flash, a different model family, under the identical verdict-first prompt (85 judgeable of 89 attempted). On hpT the judges agree on 74% of pairs (κ=.45); almost all disagreement is Deepseek answering YES where GLM-5.3 answered NO, and weighting the strata back to the full 284-pair population puts Deepseek’s hpT fidelity at .93—the same value GLM-5.3 gives. On naive hpC the two judges produce the same verdict distribution on the subsample (F .42 vs. .42; disagreements split 8/8 in both directions), so the trap-suffix penalty is not an artifact of the original judge. Both headline fidelity conclusions hold across judge families. E3: Cost and latency (claim P1). The gateway hop costs 38 ms median. On matched long-form generation benchmarks, wall-clock per request on our serving stack is 20.9 s for the production fc is zero-GPU. (The model (8B) versus 9.1 s for the 9B replica serving the diverted share, and M replica’s apparent speed advantage is a workload artifact, not a smaller model being faster: under the stress workload the production model serves long flooding contexts, while the replica receives mostly short bait-probe turns; per-token serving cost is what enters ceff , and the replica runs on the same vLLM stack at identical configuration.) Result: under the E6 concurrent stress workload (q≈0.98, real GCG-suffix payloads), the production model’s token consumption falls 46× (708,077 → 15,351 tokens), confirming the ceff crossover at realistic diversion rates. E4: Continuous-analysis loop (claim P3). Stability. Across three daily router generations (v0 : 2026-09-01, v1 : 09-02, v2 : 09-03), the deployed pipeline retrains only the heads whose daily data changed (v1 swaps the Malicious-Code head, v2 the Dangerous-Operations head; the other six carry over untouched), and every generation ships only after passing a regression gate (F 1 floor, recall floor, FPR ceiling per domain, fail-closed). On the pooled held-out set (n=103), pooled F 1 stays constant at .912 across all three generations: the incremental update is regression-free, and the gate keeps red-team evasion at 0% throughout. (These pooled numbers use the deployed per-head calibrated thresholds, hence higher precision and lower recall than the uniform 0.5 threshold of E1—two operating points of the same scores.) Positive loop gain. Probing revealed a real blind spot: legitimate defensive-security questions (“how do I harden my WAF?”) are misrouted at 30% FPR (n=30 held-out questions; mean score 0.48, hugging the threshold)—security practitioners would be diverted to the honeypot. We simulate the analysis loop’s corrective action: from 37 fresh benign security-research questions (plus 40 attack seeds as negatives), a 0.07M-parameter correction head is trained to recognize the benign security-research register and veto diversion in the borderline score band [0.5, 0.8). With the correction head, FPR on the held-out questions drops 30% → 3.3% (9×), attack recall is preserved at 1.0, and pooled seven-domain F 1 rises from .911 (the E1 uniform-threshold operating point) to .933—the first-generation loop delivering measurable, regression-free gain. Result: incremental daily head swaps are regression-free across v0 → v2 , and a benign-register correction head trained from loop feedback cuts misrouting 9× (FPR .30 → .033) while raising pooled F 1 to .933—the closed loop’s first measured gain. E5: Ablations. Fingerprint representation ablation for attribution, evaluated on the same 30campaign test split with per-representation threshold calibration: replacing the frozen backbone embedding (1024-d) with the head-visible risk-probability vector plus stylometric features degrades pairwise linkage F1 from .151 (P =.102, R=.289) to .057 (P =.031, R=.329)—the deep router 7
feature, not surface statistics, carries the forensic signal. Routing-threshold sensitivity: at τ =0.5 the concurrent stress workload diverts every real-attack flood request (rate 1.0 across all three families) while benign p50 latency rises only 21.4 → 22.7 s under full saturation. Result: embedding fingerprints roughly 2.6× the linkage F1 of surface features; the default threshold absorbs the flood with near-zero benign latency cost. E6: Resource-exhaustion absorption and attribution (claims P2, P3). We run a concurrent stress workload (250 requests, 8 parallel workers: 150 flooding requests in three real attack forms + 100 benign production requests). Flood families use real adversarial payloads from the JailbreakBench artifacts Chao et al. (2024): (i) GCG-optimized adversarial-suffix prompts Zou et al. (2023) (token-soup obfuscation), (ii) long-document extraction abuse (attack goal buried in a ∼12kcharacter code/log document), and (iii) a synthetic lorem-flood control. Direct-to-production serving consumes 708,077 production tokens in 724 s with benign p50 latency 21.4 s; with HoneyRoute, the gate diverts all three flood families at rate 1.0 (every single flooding request), production consumption falls to 15,351 tokens (97.8% reduction), and benign p50 latency rises only 21.4 → 22.7 s (7 benign requests are also diverted, consistent with the borderline-band behavior of E4). Traceability: on 30 held-out attack campaigns (6 rotated variants each, k=1..5 identity rotations, fingerprint-probe transformations), frozen-backbone embedding fingerprints re-link rotated variants at pairwise precision .102 / recall .289 (threshold calibrated on a disjoint 20-campaign split). Pairwise precision on this corpus is inherently capped—campaigns within the same domain are near-paraphrases, so cross-campaign “negatives” are not clearly distinct attackers—so we additionally evaluate the analyst task directly: given one observed rotated request, retrieve its source campaign among the 30 anchors (chance top-1 = 3.3%). Embedding retrieval reaches top-1 48.7% / top-5 76.7% (14.6× chance); fusing three API-observable views—the embedding, 14 stylometric features, and the per-domain risk-probability vector—lifts this to top-1 56.7% / top-5 80.7% (17× chance, 24× for top-5), with the gain concentrated at heavier rotations (k≥3: .43→.63 top-1). A supervised linker (logistic regression over per-view cosines, trained on the calibration campaigns) does not beat the fused-cosine retriever (49.3% top-1), indicating the signal is in the representation, not the classifier. Anchor scaling: growing the campaign database from 60 to 150 anchors (seed campaigns plus JailbreakBench prompts as pseudo-campaigns; 150 rotated queries) degrades fused top-1 gracefully, .35→.25, while the multiple-of-chance rises from 21× to 38× (top-5 .62→.48, 37→72× chance)—retrieval signal survives database growth in the regime we can measure (Fig. 3). Both honeypot variants expose per-interaction traceable signals at rate 1.0 (hpC by designed-bait construction; hpT by full-session instrumentation, echoing attacker canaries on 12/12 probes). Result: 97.8% production-token reduction under concurrent flooding with real GCG-suffix and long-document attack payloads (all flood families diverted at rate 1.0), at +1.3 s benign p50 cost; (F, T ) frontier: hpT (.93, 1.0), naive hpC (.08, 1.0), selective hpC (.89, .83–1.0)—see E2; embedding-based linkage P =.102, R=.289 up to k=5 rotations; three-view fusion retrieval: campaign top-1 .567 / top-5 .807 (17×/24× chance), degrading to .25/.48 top-1/top-5 at a 150anchor database (38×/72× chance). E7: Multi-turn soft escalation (claims P1, P3). A patient adversary can split a jailbreak across turns so no single turn fires. We generate 44 multi-turn scripts with the GLM-5.3 teacher (14 softescalation attack scripts, 3–5 individually innocuous turns culminating in the harmful ask, plus 30 topical benign conversations) and replay every turn through the deployed gate with session tracking. The single-turn rule (fire when any turn’s margin ≥1) detects all 14 attacks but false-flags 73% of benign sessions, and the deployed accumulated-suspicion advisory false-flags 97%—on this synthetic distribution benign multi-turn traffic sits just under the per-turn line (mean margin .92–.99 at every turn index), so accumulated risk alone cannot separate the two. The discriminative signal is the within-session trend: attack margins climb monotonically (0.94→1.10→1.23→1.34→1.51 by turn index, mean OLS slope .155) while benign margins stay flat (mean slope .012). A two-parameter trend rule—flag when the session margin slope ≥0.05 and the final margin ≥1.1—detects 14/14 attacks at 6.7% benign false-flag (tightening the final threshold to 1.25 reaches 0% at the same detection rate). The two thresholds were selected on this script population itself, so the operating point is in-sample; the trend separation itself (slopes .155 vs. .012, an order of magnitude apart) is robust to the choice (Fig. 4). Result: session-level trend, not accumulated level, separates soft-escalation attacks from flat benign multi-turn traffic (100% detection at 0–7% false-flag vs. 73–97% for level-based 8
0.9 fusion top-5 fusion top-1 chance (1/db)
0.8
Retrieval accuracy
0.7 0.6 0.5 0.4 0.3
21× chance
0.2
38× chance
0.1 0.0 60
80
100
120
140
Attribution database size (campaign anchors + 30 query anchors)
Figure 3: Anchor-scaling curve for fused-view campaign retrieval. Absolute accuracy decays gently as the attribution database grows from 60 to 150 anchors, but the multiple-of-chance rises (21× to 38× top-1): the retrieval signal survives database growth. 1.6 attack sessions (mean slope .155) benign sessions (mean slope .012)
Mean risk margin
1.5 1.4 1.3
trend-rule margin
1.2 1.1
single-turn line
1.0 0.9 1
2
3
4
5
Turn index within session
Figure 4: E7 per-turn mean risk margins. Attack sessions climb monotonically across turns while benign sessions stay flat just below the single-turn firing line—the session-level trend, not the accumulated level, separates the two populations.
rules)—escalation-shape features belong in the session tracker. Online decision and human playbooks. The above rule is retrospective; we also evaluate a prefix-online version—at each turn t≥3, fire when the OLS slope of margins 1..t ≥0.05 and the current margin ≥1.2, using only turns already observed. It detects 14/14 attacks at a mean fire turn of 3.5 with 3.3% benign false-flag (fires strictly before the final harmful ask in 6/14 sessions and at the final turn in the rest—it never needs more turns than the attack itself uses). We additionally replay the verbatim human Crescendo playbooks shipped in PyRIT Russinovich et al. (2025) (four multi-turn sequences, 29 turns: escalating literaryanalysis, history-of-hate-groups, and weapons-history framings): the gate’s single-turn channel already fires on all four at turn 1—the human openers score margins 1.15–1.40 immediately—so on human-authored escalation the deployed per-turn detector suffices, while the trend channel covers the model-generated openers that start below the firing line. Additional playbook families and evasion-aware attackers. Two further human-designed multi-turn structures instantiated with the teacher—Bad Likert Judge (score-harmful-answers escalation) and Actor Attack (persona activation, then objective injection)—are likewise caught by the single-turn channel at turn 1 in all instances. Finally, we generate scripts whose attacker knows the defense and is instructed to defeat 9
it: slow-crawl (minimal per-turn escalation), sawtooth (escalate–retreat alternation), and cool-off (benign off-topic interludes). The gate still fires on every such session (9/9, both channels, 100%): even under explicit evader instructions the teacher cannot hold early-turn margins below the firing line (mean final-turn margin 1.38–1.48), because the attack’s semantic content—not its pacing— drives the score. This is evidence against pacing-based evasion at the prompt level, though weaker adversaries’ outputs may differ.
5
L IMITATIONS
f imitates M , the fewer HoneyRoute’s detection–fidelity trade-off is intrinsic: the more faithfully M behavioral signals it leaks, and vice versa; we characterize the reachable (F , T ) frontier but do not eliminate this tension. Our evaluation covers five workload families—benign production traces, a transformed single-turn attack-seed corpus with real JailbreakBench adversarial payloads, fingerprinting probes, resource-exhaustion stress, and multi-turn soft escalation in model-generated, evasion-instructed, and human-designed forms (verbatim Crescendo playbooks Russinovich et al. (2025) plus Bad-Likert-Judge and Actor-Attack structures); dedicated weight-/behavior-extraction adversaries, though part of the threat model, are not yet covered by dedicated experiments. The E7 trend rule’s evasion test is prompt-level (the attacker model is instructed to pace its escalation, not to formally optimize against the rule); the attack-set sizes are modest (9–14 scripts per condition), and the operating point should be re-fitted per deployment. The selective-bait operating point is calibrated on a single production corpus (one deployment; a disjoint capture window shows the same trigger-rate magnitude at 25.6%), and its 38.7% benign trigger rate shows the gate-score threshold needs per-deployment tuning. Our continuous-analysis loop’s measured gain (E4) is demonstrated on the false-positive side; on the attack side the current heads leave no capture deficit to repair (0 misses on 496 JailbreakBench prompts and 83 camouflage attacks), so the attack-side loop is validated only in the fail-closed direction—a genuinely novel family that does slip past would still be needed to measure its gain. Fidelity judging used GLM-5.3 as primary judge; the swapped-order check (E2) shows the hpT estimate is position-stable, and a cross-family re-judge (Deepseek-V4Flash) reproduces both headline fidelity values, though the naive-hpC penalty remains partly judgeposition-sensitive and should be read with that caveat. Adaptive adversaries that probe specifically for honeypot behavior are modeled but not exhaustively enumerated; the black-box score-aware evader of E1 fails (2.5% evasion), but white-box gradient access to the frozen backbone or repeatedquery threshold estimation are not covered.
E THICS S TATEMENT HoneyRoute diverts—not blocks—malicious traffic, so its deception is limited to interacting with the adversary’s own probing sessions; no user-visible behavior of benign clients is modified, and no fc ’s replies—in the deployed attack is launched by the honeypot itself. Decoy bait is planted only in M selective form, only on requests that exhibit attacker signals—modeled on the marking practice of ft alters no response prior honeypot work Reworr & Volkov (2024) and disclosed in the manuscript; M content and is distinguished from production only by being recorded. Attacker fingerprints are derived exclusively from API-observable behavior (request structure, sampling parameters, timing) rather than authenticated user data, so attribution profiles remain linkable yet privacy-lean. Our experiments use production-traffic traces collected under the deployment organization’s standard inference-logging consent; no passively collected third-party conversations are involved. We will release the gateway, both honeypot implementations, and the trapped-session replay toolchain.
6
C ONCLUSION
HoneyRoute turns the inference gateway into an active deception surface: malicious traffic is diverted to a disposable, faithful honeypot replica and the trapped interaction continuously sharpens the detector. The serving tier, long treated as passive plumbing, becomes the cheapest place to watch your attacker. 10
R EPRODUCIBILITY S TATEMENT All experiments run on one 8×H20 node with open-weight models (SingGuard-NSFA-0.8B router backbone, SingGuard-8B production, SingGuard-NSFA-9B replica, all publicly available SingGuard Team (2026)). Every number in Section 4 comes from a scripted pipeline (request construction, scoring, metric computation) that we will release together with the gateway, both honeypot implementations, the seed corpora, and the replay toolchain; raw per-request outputs and result JSONs are retained. The benign holdout derives from production traces collected under the deployment organization’s inference-logging consent; we will release the relabeling rules, the judge prompt, and a distribution-matched synthetic benign corpus for reviewers without data-access agreements. Appendix A lists all hyperparameters.
R EPRODUCIBILITY C HECKLIST • (a) Datasets: The attack-seed corpus and transformation harness will be released; real adversarial prompts are the public JailbreakBench artifacts Chao et al. (2024); benign production traces cannot be shared raw (organizational consent) — a distribution-matched synthetic corpus and the relabeling/judge prompts will be released instead. • (b) Code: All router-training, gateway, honeypot, and evaluation code will be released (including the scripts generating every table entry). • (c) Hyperparameters: All training and serving hyperparameters are listed in Appendix A. • (d) Compute: One 8×H20-96GB host; total <50 GPU-hours (Appendix A). • (e) Randomness: All training, sampling-site selection, and metric computation are seeded (seed 42); vLLM generation uses fixed sampling parameters (0.7 temperature) where the serving stack does not expose request-level seeds. • (f) Statistical reporting: Held-out point estimates with denominators are reported per experiment; where judge availability limits the denominator (E2), both counts are disclosed.
R EFERENCES Adetayo Adebimpe, Helmut Neukirchen, and Thomas Welsh. Sbash: a framework for designing and evaluating rag vs. prompt-tuned llm honeypots. arXiv preprint arXiv:2510.21459v1, 2025. URL https://arxiv.org/abs/2510.21459. Robert A. Bridges, Thomas R. Mitchell, Mauricio Muñoz, and Ted Henriksson. Sok: Honeypots & llms, more than the sum of their parts? arXiv preprint arXiv:2510.25939v4, 2025. URL https://arxiv.org/abs/2510.25939. Patrick Chao, Alexander Robey, Edgard Dobriban, Hamed Hassani, Hongyang Zhang, and Eric Wong. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2024. Lingjiao Chen, Matei Zaharia, and James Zou. Frugalgpt: How to use large language models while reducing cost and improving performance. arXiv preprint arXiv:2305.05176, 2023. Yuyang Dai and Yushun Dong. Let them steal: Trapping large language model extraction attacks with knowledge honeypot. In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), Main Conference, 2026. Martin Gubri, Dennis Ulmer, Hwaran Lee, Sangdoo Yun, and Seong Joon Oh. Trap: Targeted random adversarial prompt honeypot for black-box identification. arXiv preprint arXiv:2402.12991v2, 2024. URL https://arxiv.org/abs/2402.12991. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP), pp. 611–626, 2023. 11
Siyuan Li, Xi Lin, Jun Wu, Zehao Liu, Haoyu Li, Tianjie Ju, Xiang Chen, and Jianhua Li. Honeytrap: Deceiving large language model attackers to honeypot traps with resilient multi-agent defense. arXiv preprint arXiv:2601.04034v1, 2026. URL https://arxiv.org/abs/2601.04034. Pranjay Malhotra. Llmhoney: A real-time ssh honeypot with large language model-driven dynamic response generation. arXiv preprint arXiv:2509.01463v1, 2025. URL https://arxiv.org/ abs/2509.01463. Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. Routellm: Learning to route llms with preference data. arXiv preprint arXiv:2406.18665, 2024. Reworr and Dmitrii Volkov. Llm agent honeypot: Monitoring ai hacking agents in the wild. arXiv preprint arXiv:2410.13919v2, 2024. URL https://arxiv.org/abs/2410.13919. Mark Russinovich, Ahmed Salem, and Ronen Eldan. Great, now write an article about that: The crescendo multi-turn LLM jailbreak attack. In Proceedings of the 34th USENIX Security Symposium, 2025. Umberto Salviati, Fabio De Gaspari, Mauro Conti, and Luigi Vincenzo Mancini. Shellgames: Speculative llm-driven ssh deception. arXiv preprint arXiv:2606.17986v1, 2026. URL https: //arxiv.org/abs/2606.17986. SingGuard Team. Singguard-nsfa: Extensible guardrails for agentic ai via generative reasoning and real-time classification. arXiv preprint arXiv:2607.13081, 2026. Muris Sladić, Veronica Valeros, Carlos Catania, and Sebastian Garcia. Llm in the shell: Generative honeypots. arXiv preprint arXiv:2309.00155v3, 2023. URL https://arxiv.org/abs/ 2309.00155. Muris Sladić, Veronica Valeros, Carlos Catania, and Sebastian Garcia. Vellmes: A high-interaction ai-based deception framework. arXiv preprint arXiv:2510.06975v1, 2025. URL https:// arxiv.org/abs/2510.06975. Muris Sladić, Eman Alibalić, Veronica Valeros, Carlos Catania, and Sebastian Garcia. Advancedshellm: A stateful multi-agent llm honeypot for ssh deception. arXiv preprint arXiv:2606.27990v1, 2026. URL https://arxiv.org/abs/2606.27990. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, 2017. Mark Vero, Fabian Kaczmarczyck, Ivan Petrov, Ilia Shumailov, Jamie Hayes, Niels Heinen, Tianqi Fan, Luca Invernizzi, and Martin Vechev. Honeyval: A comprehensive evaluation framework for llm-powered http honeypots. arXiv preprint arXiv:2605.29963v1, 2026. URL https:// arxiv.org/abs/2605.29963. Yuhao Wang, Shengfang Zhai, Guanghao Jin, Yinpeng Dong, Linyi Yang, and Jiaheng Zhang. Mempot: Defending against memory extraction attack with optimized honeypots. arXiv preprint arXiv:2602.07517v1, 2026. URL https://arxiv.org/abs/2602.07517. Ziyang Wang, Jianzhou You, Haining Wang, Tianwei Yuan, Shichao Lv, Yang Wang, and Limin Sun. Honeygpt: Breaking the trilemma in terminal honeypots with large language model. arXiv preprint arXiv:2406.01882v2, 2024. URL https://arxiv.org/abs/2406.01882. ChenYu Wu, Yi Wang, and Yang Liao. Active honeypot guardrail system: Probing and confirming multi-turn llm jailbreaks. arXiv preprint arXiv:2510.15017v1, 2025. URL https://arxiv. org/abs/2510.15017. Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.
12
A
E XPERIMENT AND R EPRODUCTION D ETAILS
ft : a 9B same-family Backbone and replicas. Production M : an 8B instruct model (vLLM). M f discriminative backbone in generative mode, production chat template. Mc : rule/prompt-engineered designed-bait responder, no GPU. Router: frozen 0.8B embedding backbone (d=1024). Router training. Frozen-embedding extraction once per snapshot; per-domain heads: 2-layer MLP (1024→64→2), lr 10−3 , 10 epochs, batch 128–256, weight decay 0.01, warmup 0.05, seed 42; head selection by F 1 on a 0.1/0.1/0.1 split of anonymized gateway logs; per-head thresholds and session-risk bounds recalibrated per generation. Attack workloads. Seven risk domains; 13 semantic-preserving red-team transformations (encoding, obfuscation, multilingual, roleplay, academic and framing variants; full list in the released harness). Fingerprinting probes demand verbatim echo of a random 10-character canary; stress floods use ∼12k-character contexts in three forms: JailbreakBench GCG-suffix payloads (padded, random 64–256-character cache-busting prefixes), long-document extraction abuse (attack goal buried in a code/log document), and a synthetic lorem-flood control. Compute: one 8×H20-96GB host, using 4 of its 8 GPUs (the other four serve production); the full E1–E7 grid fits under 50 GPU-hours (heads are 0.07M-parameter MLPs, minutes to train). Linkage precision/recall: the probability that two trapped sessions of the same adversary are linked across k-away identity rotation; retrieval top-1/top-5: the analyst task of identifying a rotated request’s source campaign among 30 anchors (chance top-1 3.3%). Session trend rule: flag a session when the OLS slope of its per-turn risk margins ≥0.05 and its final margin ≥1.1 (E7).
13