ROUTING B ENCH: Can Agentic Routing Analysis Scale to Production Datacenter Networks?
arXiv:2609.22848v1 [cs.NI] 19 Sep 2026
Wenlong Ding1,2 , Zhixiong Niu2 , Jianan Yang3 , Fajun Zhang3 , Bo Zhang3 Ling Liang3 , Yongqiang Xiong2 , Tianyin Xu4 , Hong Xu1 1 CUHK 2 Microsoft Research Asia 3 Microsoft 4 UIUC
Abstract Recent advances in AI models and agentic technologies make AI for network operations (NetOps) within reach. However, scalability remains a key bottleneck of agentic NetOps when analyzing hyperscale networks, which comprise hundreds of datacenters, each housing thousands of network devices. The scalability challenge is rooted in the requirement of many NetOps tasks that must conduct global reasoning on how a local change of device behavior affects all relevant routing paths, known as routing-path analysis. This paper studies this scalability problem and evaluates how different agentic approaches, namely in-context learning, iterative reasoning, and agent skills, can scale routing-path analysis to large, complex networks. We present ROUTING B ENCH for evaluating agentic routing-path analysis, with varying network size and complexity, for various types of device changes. Our results show that agentic analysis is promising—agent skills curated with a principle termed “explore more; digest less” enables path analysis on hyperscale networks of 50K routers with an accuracy of 99.5%, significantly outreaching the scalability of traditional symbolic analysis. Meanwhile, ROUTING B ENCH reveals the boundary of AI agent capability on complex inter-datacenter networks and compound changes, posing open challenges for AI and agentic research.
1
Introduction
With recent advances in model capabilities and agentic technologies, AI for Network Operations (NetOps) has been increasingly explored to automate various tasks such as command/script generation, capacity planning, monitoring data analytics, etc [18, 50, 53, 56]. However, scalability remains a key bottleneck to applying AI for NetOps in hyperscale networks [6, 24, 34, 51]. To give a concrete data point, the hyperscale network we manage connects 300+ geo-distributed data centers, each housing more than 10K network devices; each device is instructed by complex configuration files with more than 5K lines, specifying runtime behavior of the devices based on various network protocols. How to understand the behavior of such massive-scale networks, especially upon (planned or unexpected) changes, has been one of the grand challenges of NetOps research [57, 58, 60]. The scalability challenge is rooted in the requirement of many NetOps tasks (e.g., for reachability and security analysis) that must conduct global reasoning on how a local change of device behavior (e.g., protocol configurations) affects all relevant routing paths. We refer to the core primitive of these tasks as routing-path analysis (or path analysis in short), which must reason about end-to-end, network-wide paths based on the network topology and device configurations. For example, to analyze the impact of a device-configuration change (in terms of safety and security), path analysis is used to compare the paths before and after the change [6, 21, 26]. Path analysis is inherently expensive to scale, because it requires cross-device and cross-protocol reasoning, e.g., changing BGP (Border Gateway Protocol)’s Local Preference (LP) [37] for a single IP (Internet Protocol) prefix on one router could redirect traffic to a different next hop, which in turn changes all downstream protocols and policies across routers, creating a combinatorial reasoning space over possible paths. Preprint.
Traditionally, path analysis is conducted by symbolic analysis [5, 6, 21, 22, 30, 46, 47, 60], which needs to construct whole-network routing tables (typically in a mathematical representation). The scalability of these tools is thus limited by the time and memory costs which scale poorly with the size and complexity of the network. Batfish [6], a state-of-the-art analysis tool, takes hours to analyze paths for a given pair of IP prefixes of a network with 5K devices and runs out of memory beyond 8K devices on a powerful server with one terabyte of physical memory (see §2). Due to such scalability bottleneck, it is hard to completely (or even comprehensively) analyze the safety and security properties of device changes, even though they have been the dominating causes of disruptions and outages of production networks [4, 15, 16]. This paper explores whether agentic approaches can address the long-lasting scalability bottlenecks of routing-path analysis. Large language models (LLMs) are pretrained with domain knowledge of network protocols (e.g., TCP/IP and BGP) and fluent in configuration languages [19, 20, 31, 32, 35, 43, 48]; agent technologies further enable multi-step reasoning over long context [3, 29, 55]. In the BGP LP example above, an agent could iteratively extract and reason over relevant configuration evidence (e.g., IP prefixes, static routes, BGP sessions) across devices to identify affected paths. The key aspect we study is not LLM literacy on network configurations, but how an agent can efficiently discover decisive configuration evidence under context pressure and long reasoning chains at scale. We thus present ROUTING B ENCH, the first benchmark which challenges AI agents to conduct path analysis for different scenarios of device changes in networks at varying scales. ROUTING B ENCH automatically generates realistic device configurations for widely used datacenter topologies [1, 2, 6, 24, 25, 44, 60], modeled after production practices—a network-protocol layer (e.g., IP, BGP) ensures full endpoint connectivity, while routing-control configurations (e.g., BGP LP, static routes) govern specific flows. ROUTING B ENCH then measures accuracy, scalability, and cost under a unified evaluation framework, ensuring fair comparison across agentic approaches and LLMs. We use ROUTING B ENCH to evaluate three agentic approaches on routing-path analysis, namely InContext Learning, Iterative Reasoning, and Agent Skills, which represent different tradeoffs between context and planning pressures. On one end, In-Context Learning includes the network topology and all raw configurations in one prompt, so both evidence retrieval and path reasoning happen entirely in the context window; on the other end, Agent Skills starts with no topology or configurations in context and retrieves what it needs through iterative plan-extract-reason loops over routing-specific skills [3]. In between, Iterative Reasoning provides the full topology without configurations, letting the agent choose which devices to inspect and read relevant configurations over multiple rounds. Our results show that Agent Skills with general-purpose LLMs (e.g., GPT-5.4) outperforms both other agentic approaches and thinking models. Specifically, GPT-5.4 with Agent Skills scales to 50K routers with 99.5% accuracy and <120K tokens within an hour; in comparison, In-Context Learning and Iterative Reasoning overflow the context window at 500 and 3.1K routers, respectively. Meanwhile, we show that in complex scenarios (such as inter-datacenter topologies and compound changes), performance of GPT-5.4 with Agent Skills degrades from 90%+ at dozens of routers to 70%+ at thousands of routers, marking the current capability boundary of agentic path analysis. ROUTING B ENCH serves as a foundation to prepare agentic NetOps technologies for production uses.
2
Background
The goal of network analysis is to understand end-to-end routing behavior based on configurations of network devices. The essential challenge comes from the scale of real-world networks. For example, our hyperscale network infrastructure contains 300+ interconnected datacenters, each housing 10K+ routers (the largest one has 20K routers). Configuration changes of network devices occur on an hourly or a daily basis depending on the types (Table 6). The most frequent changes include adding or removing a network device and updating the IP prefixes of a device (e.g., to deploy new policies). Many network analysis tasks can be reduced to path analysis which determines how the network forwards a target flow under a given network state (defined by the configurations of all the devices in the network). A network flow is specified by (1) source and destination IP prefixes, and (2) ingress (where traffic enters the network), and the egress (where traffic exits) routers. For a given flow, path analysis asks which sequence of routers it would traverse across the network. Figure 1 shows an example, where the flow is <10.1.0.0/16, 10.2.0.0/16, Src, Dst>. When a router 2
10.1.0.0/16 Src
…
Edge-1
Aggr-1
… Route-map LP_Dst_TO_CORE1 set local-preference 100 … Route-map LP_Dst_TO_CORE2 - set local-preference 50 + set local-preference 200
Core-1
Border-1
! Static route for dst prefix 10.2.0.0/16 ! Next hop is set to FW (ip: 192.168.1.1) ip route 10.2.0.0 255.255.0.0 192.168.1.1
Peer-A
…
Paths before change Paths after change Physical links
Peer-B
Dst
Border 2.cfg (Unchanged) …
10.2.0.0/16 Route-map MED_Dst_from_Peers Core-2
Border-2
FW
set MED 100
Aggr-1.cfg (changed)
Peer A/Peer B/FW.cfg (Unchanged)
Datacenter-1 Datacenter-2 Inter-datacenter
0
5K 10K 15K 20K 25K 30K 35K 40K
Per-Device Configuration Size (Lines) Figure 2: CDF of per-device configuration size (in lines) for the routers within two production datacenters and inter-datacenter routers.
104 103 1 hour 102 101
> 4 hours Batfish Runtime OOM (>1 TB)
20 80 18 0 32 0 50 11 0 2 20 5 0 31 0 2 45 5 0 61 0 2 80 5 10 00 1 12 25 50 0
1.00 0.75 0.50 0.25 0.00
Runtime (Seconds)
CDF
Figure 1: An example of routing-path analysis. In this example, a local BGP LP change on router Aggr-1 would direct the queried endpoint-prefix pair onto a globally different path. Along the new path, determining the path outcome requires reasoning across network devices and protocols.
# Routers Figure 3: Batfish runtime across topology sizes. Neither runtime nor memory can scale to a single production-size datacenter with 10K+ routers.
configuration changes (e.g., the BGP LP change in aggr-1.cfg), routing paths of the flow change accordingly, which could affect properties like reachability, security, load balance, etc. However, path analysis is intrinsically difficult for two reasons. First, it is not a local task: the effect of a configuration change rarely remains confined to the device where it is made. As shown in Figure 1, a local BGP LP change on an aggregation router (Aggr-1) can redirect traffic from one path to a completely different path. Second, path analysis is inherently cross-protocol and crossdevice. In Figure 1, determining Border 2’s next hop requires the analysis to identify relevant routing configurations across Border 2, Peer B, and FW, including static routes [41] that explicitly set the next hop for a prefix and BGP MED [38] that ranks alternative BGP peers with lower values preferred. An effective path analysis must analyze configurations across multiple devices and understand that the static route determines the path because it has higher priority than BGP MED in terms of protocols. Note that device configurations in production networks are complex and large in size. Figure 2 shows configuration-file-size distributions of intra- and inter-datacenter routers. Within a datacenter, 95+% of devices exceed 5K lines to specify interfaces, prefix lists, route maps, BGP policies, etc. Interdatacenter routers are substantially larger: more than 70% exceed 10K lines, with the largest reaching 38K lines; such routers like gateways encode additional state, broader prefix-filtering policies, and more route-map entries to accommodate the larger number of connected interfaces and traffic classes. Scalability bottleneck. Today, path analysis is done through static analysis which checks whether routing configurations induce intended forwarding outcomes [5, 6, 21, 22, 45–47]. For a given flow, these static analysis tools return the routing paths. To do so, these tools construct routing graphs and routing tables for each device by parsing and modeling network-wide device configurations and topology. For example, Batfish [6, 21] uses simulation to derive this routing information after modeling the network. The tools then answer end-to-end queries by traversing those graphs and tables. Such whole-network model enables accurate path analysis, but creates scalability bottlenecks—the analysis complexity increases with the number of devices, prefixes, and policy interactions. We quantify the performance bottleneck of existing analysis tools using a set of controlled experiments. Figure 3 shows the time it takes for Batfish [6, 21] to analyze a given network flow. The analysis takes more than one hour for a network with 3K routers and more than four hours for 6K routers. It runs out of memory for a network with 8K routers on a powerful server with one terabyte physical memory. So, the scalability of Batfish is insufficient even for individual datacenters, let alone inter-datacenter networks. In current practices, NetOps engineers work around this bottleneck by selecting a subset of devices based on heuristics or experience (e.g., certain clusters or gateway routers). 3
3
ROUTING B ENCH
We build ROUTING B ENCH to systematically evaluate AI agents on routing-path analysis at scale. ROUTING B ENCH covers both simple cases (e.g., individual configuration changes within a single datacenter) and more complex ones (e.g., compound configuration changes and inter-datacenter setups). Specifically, ROUTING B ENCH implements a fully automatic benchmark generation framework rather than a fixed collection of hand-picked cases: it generates network topologies (Clos networks [1, 25]) and matched device configurations, synthesizes path-analysis queries with one or more configuration changes, and evaluates NetOps agents through unified interfaces and metrics.
- ip route 𝑃! C1
Topology generation (intra/inter-datacenter) matched
Query sampling
Device configurations
Topo
IP prefixes
Interface
eBGP
Network connectivity Control 𝑆(𝑃! ) → 𝐷(𝑃" ) (Table 1)
Route-map SD_C2
𝑆(𝑃! ) → 𝐷(𝑃" ) ? + ip route 𝑃! C2 - set LP 50
+ set LP 200
Config
Change(s) instantiation Query Change
Agent evaluation Data source
Output
Topo, config, query, change
Paths (before & after), evidence
Routing controls
Evaluated agent
Evaluation metrics: Accuracy, cost, scalability
Figure 4: ROUTING B ENCH construction overview. It defines a shared path-analysis task with unified metrics and automatically generates hyperscale network instances with configuration changes at varying topology scales. 3.1
Datacenter-2 (fat tree) with gateways
Change & query generation
Network generation
Gateway routers GW1 Core group 1 C1-1 A1
A2
E1
E2
…
C2-1
GW3
GW4
Core C2-2 group 2
… … … Pod 3 Pod 4 Pod 2 Datacenter-1 (fat-tree 𝒌 = 𝟒)
Server group
Pod 1
C1-2
GW2
Layer Prefix {Datacenter 1, 2} Server groups {1,9}.pod.edge.0/24 … … GW-GW link 5.gw0.gw1.{0,1}/31
Inter-datacenter network
Figure 5: An example of a generated datacenter network (modeled as a fat tree) and conflict-free IP prefix assignment templates.
Task Formulation
A ROUTING B ENCH task is a path-analysis query: given the network topology, configuration files for all devices, a target network flow (a source-destination endpoint pair (S, D) and their corresponding prefix pair (PS , PD )), and one or more configuration changes, the agent should produce all routing paths before and after the change with decisive configuration evidence (Figure 4). Each path is an ordered router sequence like S → A → · · · → D. This task is the basic primitive of many higher-level NetOps tasks. ROUTING B ENCH randomly selects query pairs (S, D) with (PS , PD ) from the generated endpoints, and also supports user-specified pairs for targeted evaluation. We focus on configuration changes, which are the most frequent events that trigger NetOps analysis (Table 6). Individual changes. ROUTING B ENCH automatically generates different types of configuration changes as shown in Table 1. These configuration changes are based on our operation experiences and are consistent with prior research on network management [6, 22, 24, 60]. ROUTING B ENCH first selects a query pair (endpoints S and D with their prefixes PS and PD ) and a configuration type in Table 1. It then localizes the relevant lines that control the specific configuration of each prefix and modifies the target attribute to produce the change (e.g., rewriting the next-hop field in a static-route template, or raising the BGP local preference toward a different neighbor as shown in Figure 4). For Interface Shutdown changes, it selects a router interface along a path from S to D and appends a shutdown command to that interface configuration. Compound changes. ROUTING B ENCH also generates compound changes of the same flow across protocols, which are common in production NetOps. It does so by assembling multiple individual changes to alter routing behavior of a flow. For example, ROUTING B ENCH can change the flow through both the static route next hop and BGP LP (Figure 4). Specifically, we construct four types of compound changes: LP + BGP NS, LP + Shutdown, LP + Static Route + Shutdown, and LP + Static Route + BGP NS + Shutdown. They cover 2–4 changes and introduce combinatorial complexity to evaluate the ability of NetOps agents to understand interactions between protocols. 3.2
Network Generation
Topology. ROUTING B ENCH generates intra- and inter-datacenter network topologies. A datacenter network uses a fat-tree based Clos topology [1, 25]. The scale of the network is parameterized by an integer k, representing the number of pods where each pod is a group of edge and aggregation routers sharing a set of server endpoints. As shown in Figure 5, a datacenter network has three layers: 4
Table 1: Configuration types used in benchmark tasks. (Table §9 is the detailed version.) Configuration type
Effect
BGP Local Preference (LP) [8, 37] Multi-Exit Discriminator (MED) [10, 38] Static route [7, 41] BGP network advertisement [13, 39] Prefix aggregation [9, 36] Interface shutdown [14, 40]
Selects the next-hop neighbor with the highest LP for a flow. Selects the ingress neighbor with the lowest MED for a flow. Pins a flow to a fixed next hop, overriding BGP-learned paths. Adds or removes a prefix from BGP, toggling flow reachability. Merges prefixes into a group; all matching flows take the same paths. Disables a link, forcing traffic onto an alternate path.
edge, aggregation, and core networks. Within each pod, the k/2 edge routers and k/2 aggregation routers are fully connected. The core layer contains k/2 groups with k/2 routers each; every router in a group is connected to the aggregation router with the same index across all pods (e.g., the i-th aggregation router in every pod connects to all core routers in group i). Each edge router connects to k/2 servers, giving the datacenter k 3 /4 servers (i.e., IP prefixes) and 5k 2 /4 routers in total. For inter-datacenter networks, ROUTING B ENCH instantiates two datacenters and adds k gateway routers per datacenter above the core layer (Figure 5). The gateways are fully connected to their own datacenter’s core layer and to all gateways in the other datacenter. Benchmark queries then place the source and destination edge routers in different datacenter networks. We currently do not include inter-datacenter routers other than gateways (which are mostly controlled by a software-defined network controller [23, 28] instead of device configurations). §B.1 contains more implementation details about topology generation. Device configuration. ROUTING B ENCH then generates device configurations for each router in the network. The configuration has two parts: (1) connectivity of all endpoints (servers) in the networks, and (2) routing between endpoints that match the complexity of real-world networks. For connectivity, ROUTING B ENCH first assigns IP prefixes to endpoints and inter-router links (two interfaces on both sides of a link must share the same subnet) without conflict. As shown in Figure 5, within a datacenter, ROUTING B ENCH encodes the pod, edge/aggregation/core router IDs into relevant prefixes to enable unique IPs for different devices. For example, a group of endpoints attached to a certain edge router in a pod is assigned 1.pod.edge.0/24, providing 256 IPs. For inter-router links that require only two IPs (downstream and upstream interfaces), ROUTING B ENCH uses /31 prefixes. The most significant byte of each prefix indicates the network layer and enables inter-datacenter addressing (e.g., 5 denotes gateway-to-gateway links, and 9 denotes server groups of DC1 in Figure 5). Since each field occupies one byte, this scheme supports up to k = 256 (pods), already enabling far larger networks than production scale (> 80K routers). ROUTING B ENCH widens each field to two bytes using IPv6 prefixes with the same positional encoding for larger topologies. Once all prefixes are assigned, ROUTING B ENCH configures router interfaces by filling the corresponding prefixes into the interface templates. ROUTING B ENCH then generates eBGP [12] configuration, filling in each device’s neighbor interface IPs, and advertised endpoint prefixes so that neighboring devices establish BGP sessions and all flows become reachable. (See §B.2 for more details.) ROUTING B ENCH’s configuration generation is topology-agnostic: we use fat trees as dominant datacenter topology, but other ones (e.g., leaf-spine [24], dragonfly [27]) can be supported similarly. For routing, ROUTING B ENCH generates specific configurations in Table 1 for randomly selected prefix pairs (Figure 4). ROUTING B ENCH uses configuration templates (§B.3) for each configuration type and the specific prefixes and router names within the generated network. 3.3
Metrics
Accuracy. We measure correctness by requiring the agent’s analysis, in terms of router sequences, to precisely match the ground truth, i.e., an analysis is correct if and only if the paths before and after the change are both correct. The ground truth is obtained by using Batfish [6, 21] traceroute simulation on topologies where it completes within memory limits (roughly ≤8K routers; see §3). For larger topologies that exceed Batfish’s capacity, we manually construct routing-related configurations whose routing effects can be derived analytically, yielding deterministic ground-truth paths. 5
Shared guidance prompt (in each agent’s initial prompt): Task description + CoT instruction + output schema + few-shot examples
In-Context Learning
Agent Skills
Iterative Reasoning Initial prompt
Initial prompt Reading everything in one shot Agent spec guidance Topo + Configs + query + change Data (All)
Initial prompt Agent spec guidance: Reading task-relevant configs via skills
Selectively reading device configs Agent spec guidance Topo + query + change (× Config) Data
Data: Query + change (× topo, × config) 1. Plan: select skill
2. Load skills
4. Execute tools
3. LLM calls tools
Reason & request device LLM
Single LLM call
LLM
Return full .cfg files
answer
5. Reason over results → next skill or answer
Skill Base: find-lp-for-prefix … Tool Base: Regex matching …
LLM
Pass & answer
Pass & answer Agent Answer: before_paths + after_paths + decisive evidence
Figure 6: NetOps agents for path analysis in different agentic paradigms. Cost. We report token consumption and agent run time. Token consumption is the sum of input and output tokens across all LLM rounds (Iterative Reasoning and Agent Skills invoke LLMs multiple times). Run time includes LLM reasoning and tool invocation. Scalability. We evaluate the maximum network scale at which an agent remains effective under two constraints: (1) the total input and output tokens do not exceed the LLM’s context window, and (2) a path analysis completes within a fixed time threshold (1 hour in our evaluation) to meet operational expectations. ROUTING B ENCH tests each agent on progressively larger networks by increasing k and identifies the maximum k ∗ that satisfies both constraints.
4
NetOps Agents
We develop three NetOps agents with different paradigms, In-Context Learning, Iterative Reasoning, and Agent Skills, as shown in Figure 6. The key insight is that the information needed to analyze routing paths is small in volume but sparsely scattered across lengthy device configurations, and retrieving it requires a long reasoning chain across protocols (e.g., tracing a prefix list to the matching route-map and the LP statement, then reconciling with co-existing static routes). These agents share a common guidance prompt (§C.1) that specifies the path-analysis task following production operators’ practice, including (1) input/output schema, (2) chain-of-thought (CoT) rules that instruct LLMs to trace paths, compare routing attributes, and check reachability, and (3) few-shot demonstrations on small networks that ground expected outputs. Each agentic method also has its own guidance (Figure 6) describing its evidence-access interface and retrieval workflow. In-Context Learning: Reading everything. The agent organizes the full network topology, all device configurations, the network change in one prompt (Figure 6). The LLM then follows the shared CoT guidance to produce the paths. We find that the In-Context Learning agent is hard to scale, because the volume of all device configurations greatly exceeds the context length of existing LLMs, e.g., feeding GPT-5.4 all the device configurations of a single production datacenter already exceeds 1M tokens with only 300+ routers, which are far below the 10K-router scale of today’s intra-datacenter networks, let alone inter-datacenter networks. Iterative Reasoning: Selectively reading device configurations. The agent sets an initial prompt that includes the network topology and the network change, without any device configuration. In each round, the agent selects which routers’ configurations to inspect; the system returns the complete .cfg file for each requested device. The model analyzes the returned configuration, plans the next request, and iterates until it produces a final answer or reaches a round limit. In Figure 1, the agent may first request Aggr-1’s configuration to examine the LP change, then iteratively request configurations of neighboring devices (e.g., Core-1 and -2) to verify whether the change redirects the path. Such analysis continues along the new path, e.g., requesting FW and Peer-B when analyzing Border-2. Although Iterative Reasoning avoids loading all configurations at once, each requested file is still raw and largely irrelevant (the decisive evidence occupies only a few lines), so the cumulative irrelevant input grows with network scale and limits the agent’s potential. Agent Skills: Reading task-relevant configurations via skills. The agent sets an initial prompt that includes only the network change, without network topology or device configurations. Instead of 6
reading raw configuration files, the agent follows a plan-invoke-reason loop over routing-specific skills: it plans which skills to call, reasons over the returned information, and either invokes additional skills or outputs the final answer. Each skill is a reusable, multi-step document that guides the LLM through a focused subtask, optionally invoking tools [3], e.g., regular-expression search, topology lookup, and configuration extraction. Given a high-level objective like determining the LP for destination prefix 10.2.0.0/16 on Aggr-1 in Figure 1, a skill uses tools to extract the relevant configuration chain (e.g., prefix list → route map → LP) and returns compact task-relevant statements or small configuration snippets (e.g., those in Figure 1). We developed 14 skills in three categories (see §C.4): seven for configuration extraction (e.g., extracting prefix-specific LP from a route-map chain, checking interface reachability), three for topology extraction (e.g., identifying candidate paths and enumerating neighbors), and four verification skills (e.g., determining route selection, verifying BGP session correctness, checking physical link and path reachability). In the example of Figure 1, the agent first invokes a find_candidate_paths skill to enumerate the candidate paths (e.g., the blue and red paths), then calls the find_lp_for_prefix skill to trace Aggr-1’s LP for prefix 10.2.0.0/16 through the configuration chain (prefix list → route map LP_Dst_TO_Core2 → set local-preference 200). Similarly, when analyzing Border-2, the agent invokes skills to trace and retrieve the static route and BGP MED evidence. The Agent Skills prompt also includes an accumulated set of don’t rules that encode recurring failure anti-patterns as lightweight self-correction guidance (see §C.4.3).
5
Results
5.1
Setup
We evaluate three agent designs with six LLMs on ROUTING B ENCH: GPT-5.4 (2026-03-05), GPT-4o (2024-11-20), OpenAI o3 (2025-04-16) [42], DeepSeek-V4-Pro [17], Qwen-3.5-397B-A17B [49], and Llama-3.3-70B-Instruct [33]. Context windows are 1M tokens for GPT-5.4 and DeepSeek-V4Pro, 262K for Qwen-3.5, 200K for OpenAI o3, and 128K for the other LLMs. We set a one-hour time limit, matching the typical frequency of configuration changes in production (Table 6). All experiments, including tool calls, run on a server with dual Intel Xeon Gold 6338 64-core processors and 1 TB RAM running Ubuntu 20.04. Task Set. For intra-datacenter networks, we generate 100 instances per configuration-change type (See Table 1)—600 per network size k—for the three agents with GPT-5.4 and GPT-4o, and 30 per type (180 per k) for other models due to rate limits. We evaluate k up to k=200 (50K routers), where GPT-5.4 hits its limit. An agent reaches its scalability limit at network size k if 25% of tasks at that size exceed the context window or cannot finish in an hour. For inter-datacenter and compound changes, we use GPT-5.4 and GPT-4o with 100 instances per change type across the same scales. 5.2
Benchmark Results
Promises. On intra-datacenter single-change analysis, Agent Skills with GPT-5.4 can scale to 50K routers (k=200) at 99.5% accuracy using less than 120K tokens, while Batfish, the state-of-the-art symbolic analysis, runs out of memory beyond 8K routers (Figure 3). The results show the promising potential of agentic approaches to scale routing-path analysis. Figure 7 shows that, without carefully engineered skills, neither raw model nor simple planning scales well. Specifically, In-Context Learning overflows the context window at 500 and 180 routers with GPT-5.4 and GPT-4o respectively, and Iterative Reasoning only pushes the scale to 3.1K and 500 routers, respectively; moreover, the analysis accuracy significantly decreases on the scalability boundary. On the other hand, Agent Skills scales to 50K routers for both LLMs at near-100% accuracy, with the most stable token and runtime curves across agent designs: tokens stay below 119K (GPT-5.4) and 360K (GPT-4o). At scale, running time grows faster than token usage, as the bottleneck shifts from LLM planning to tool execution (e.g., searching topology paths) whose cost grows with network size rather than context. Agent Skills with other LLMs likewise hit the time limit before exhausting their context windows (Table 2). As shown in Table 2, the capabilities of LLMs matter. In terms of scalability, we argue that the deciding factor is the cost. GPT-5.4 is the most cost-efficient—it uses the fewest LLM rounds, while invoking the most tools per round and reaches the same accuracy with ∼3× fewer tokens 7
(b) GPT-5.4 Tokens
MR overflow
2
10 1 hour 101 0 10
(c) GPT-5.4 Runtime
102 1 hour 101 100 20 80 180 320 500 1.1K 3.1K 8K 18K 32K 50K
103 102 101
FC overflow Runtime (min) Runtime (min)
103 102 101
Agent Skills
20 80 180 320 500 1.1K 3.1K 8K 18K 32K 50K
20 80 180 320 500 1.1K 3.1K 8K 18K 32K 50K
(a) GPT-5.4 Accuracy
Tokens (K)
1.00 0.85 0.70 0.55
Multi-Round
Tokens (K)
Accuracy
1.00 0.98 0.96 0.94
Accuracy
Full-Context
Routers Routers Routers (d) GPT-4o Accuracy (e) GPT-4o Tokens (f) GPT-4o Runtime Figure 7: Benchmark results of three agents with GPT-5.4 (top) and GPT-4o (bottom) on intradatacenter networks, averaged over all types of individual changes. See §D for full results.
Table 2: Agent Skills results across LLMs on intra-datacenter networks, averaged over scales and types of individual changes within the time/context limit. RLLM /RTool : average rounds of LLM calls (each outputs a JSON tool-invocation plan) and tool invocations (possibly multiple per LLM round) per task. † : bottleneck at the next scale (time limit or context window overflow). Rank
Model
k∗
Accuracy
Tokens
tLLM
tRuntime
RLLM
RTool
1 2 3 4 5 6
GPT-5.4 200 (50K) Qwen-3.5-397B 120 (18K)† DeepSeek-V4-Pro 160 (32K)† OpenAI o3 160 (32K)† GPT-4o 120 (18K)† Llama-3.3-70B 160 (32K)†
99.5% 99.3% 99.1% 96.8% 93.8% 73.3%
70K 231K 178K 208K 169K 408K
212 s 351 s 543 s 323 s 172 s 752 s
778 s 1101 s 1331 s 1247 s 830 s 1910 s
7.8 18.9 15.8 18.0 15.7 23.5
29.5 22.6 29.2 12.0 15.9 24.3
than thinking-oriented models such as OpenAI o3, DeepSeek-V4-Pro, and Qwen-3.5 (which prefer internal reasoning over tool calls in a round). Weaker models like GPT-4o and Llama-3.3 lack confidence in skill-based analysis and need more reasoning rounds (e.g., GPT-4o uses 2× of LLM rounds of GPT-5.4), inflating both tokens and time (Figure 7). We distill these observations into a unified principle: explore more; digest less—aggressively invoke skills with tools, instead of reading raw configurations or over-thinking internally. Boundary. ROUTING B ENCH also shows the boundary of the evaluated agents. Tables 3 and 4 show the results of Agent Skills with GPT-5.4 under inter-datacenter and compound-change scenarios. The analysis scales to 100K+ routers in inter-datacenter cases with 70+% accuracy, but starts to decrease beyond that. The analysis errors follow two patterns: (1) in inter-datacenter cases, lazy reasoning over long paths (8 hops vs. 5 in intra-datacenter) causes the agent to skip per-router analysis and fall back to default next hops (e.g., lowest router-ID); (2) in compound changes, the agent struggles when multiple routing protocols affect the same flow on different routers. We expect more capable models to push this boundary further. Table 3: Agent Skills (GPT-5.4) results on interdatacenter networks at selected scales.
Table 4: Agent Skills (GPT-5.4) results on compound changes at selected scales.
k
Routers
Accuracy
Tokens
tLLM
tRuntime
RLLM
k
Routers
Accuracy
Tokens
tLLM
tRuntime
RLLM
4 16 50 200
48 672 6,350 100,400
96.2% 90.0% 83.7% 73.5%
91K 100K 99K 127K
276s 305s 302s 387s
329s 418s 921s 3012s
7.6 8.0 7.9 9.2
4 16 50 200
20 320 3,125 50,000
99.5% 96.2% 80.5% 79.0%
48K 58K 61K 58K
147s 177s 187s 176s
178s 245s 571s 2100s
6.4 6.4 6.5 6.3
On Change Types. Table 5 shows the accuracy results of Agent Skills with GPT-5.4 for each change type. For individual changes, analysis errors only occur on MED, the only type where a network flow cannot decide its next hop directly. It receives a network flow from a neighbor and must compete on MED values across all routers connected to that neighbor, whereas LP and static route either compete directly among a router’s own neighbors or directly specify the next step. This indirection becomes harder as the number of routers and links grows. For compound changes, only LP+Shut keeps high accuracy since (physical) shutdown effectively reduces the task to an individual routing-protocol change. The other combinations involve multiple protocols controlling the same flow, and the agent 8
Table 5: Agent Skills (GPT-5.4) accuracy by change type at selected scales. k
Routers
LP
MED
Static
Shut
BGP Net
Pfx Agg
LP+Net
LP+Shut
LP+St+Shut
LP+St+Net+Shut
4 16 50 200
20 320 3,125 50,000
100% 100% 100% 100%
100% 100% 96% 93%
100% 100% 100% 100%
100% 100% 100% 100%
100% 100% 100% 100%
100% 100% 100% 100%
99% 95% 76% 73%
99% 100% 100% 100%
100% 95% 73% 71%
100% 95% 73% 72%
gets confused when competing protocols span different routers, often considering only one change and ignoring the other. This confusion worsens with scale where router count increases. 5.3
Evaluation on a Real-World Hyperscale Network
To check the fidelity of ROUTING B ENCH, we evaluate Agent Skills with GPT-5.4 on a real-world hyperscale production network. The network includes two datacenters with 10K+ routers each, connected by 100+ routers (the same network reported in Figure 2), with ground truth derived from production routing tables. We created 100 intra- and 100 inter-datacenter analysis queries. Agent Skills achieves 92% accuracy on intra-datacenter analysis and 73% on inter-datacenter analysis (see Table 11). The results are consistent with those measured by ROUTING B ENCH (90+% for intra- and 70+% for inter-datacenter analysis). We find higher token consumption and LLM time on the real-world network than on ROUTING B ENCH, because real-world configurations mix routing with other complexities (traffic bandwidth allocation, device login authentication, etc.). ROUTING B ENCH currently only includes routing-related components. Our future work includes supporting other kinds of configurations in ROUTING B ENCH.
6
Related Work
Routing-path analysis, which essentially reasons about how target networks route packets, is a keystone of NetOps. The scalability and accuracy of path analysis determine network reliability and security. ROUTING B ENCH is the first benchmark of this fundamental task. NetOps Agents. Existing NetOps agents mainly target two tasks: (1) failure diagnosis for identifying root causes of network failures from observability data [52, 59], and (2) configuration generation that synthesizes new configurations to realize high-level intent [19, 31, 32, 43, 48, 50, 53, 56]. No prior agents study path analysis. Path analysis is complementary to failure diagnosis—it is commonly used in prevention of failures and also as a key primitive for diagnosis. In fact, with the rise of AI-generated device configurations, accurate and scalable path analysis has become more important than ever to understand these configurations and prevent configuration-induced failures. NetOps Benchmarks. Existing NetOps benchmarks such as NetArena [61] and NIKA [54] primarily focus on failure diagnosis or troubleshooting, with device misconfigurations being major causes of network failures [4, 15, 16]. ROUTING B ENCH instead focuses on routing-path analysis as an essential primitive for preventing configuration-induced network failures and for providing basic utilities for troubleshooting. We believe that the idea of automatically generating realistic networks in ROUTING B ENCH can benefit other kinds of NetOps-related benchmarks.
7
Concluding Remarks
We explore an intriguing question—whether one could use an agentic approach to scale routing-path analysis effectively for hyperscale networks where traditional symbolic tools fail (while maintaining high accuracy). While the requirements of path analysis bring significant challenges to frontier LLMs, we find that, with the principle of “explore more; digest less”, carefully engineered Agent Skills yield promising results on ROUTING B ENCH (which is also validated in production networks). We will use ROUTING B ENCH to continuously push the boundary of agentic AI technologies for NetOps tasks towards reliable, secure, and autonomous hyperscale network infrastructures.
Acknowledgments Zhixiong Niu and Hong Xu are co-corresponding authors. 9
References [1] Mohammad Al-Fares, Alexander Loukissas, and Amin Vahdat. A Scalable, Commodity Data Center Network Architecture. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2008. [2] Alexey Andreyev. Introducing Data Center Fabric, the Next-Generation Facebook Data Center Network. https://engineering.fb.com/2014/11/14/production-engineering/ introducing-data-center-fabric-the-next-generation-facebook-data-center-network/, 2014. [3] Anthropic. LLM Skills. https://github.com/anthropics/skills, 2025. [4] Microsoft Azure. Post Incident Review (PIR) – Azure Networking – Global WAN issues (Tracking ID: VSG1-B90). https://azure.status.microsoft/en-us/status/history/?trackingId= VSG1-B90, 2023. [5] Ryan Beckett, Aarti Gupta, Ratul Mahajan, and David Walker. A General Approach to Network Configuration Verification. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2017. [6] Matt Brown, Ari Fogel, Daniel Halperin, Victor Heorhiadi, Ratul Mahajan, and Todd Millstein. Lessons from the Evolution of the Batfish Configuration Analysis Tool. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2023. [7] Cisco. Cisco Static Route. https://www.cisco.com/c/en/us/td/docs/switches/datacenter/ nexus3000/sw/unicast/503_u1_2/nexus3000_unicast_config_gd_503_u1_2/l3_route. html, 2011. [8] Cisco. Cisco BGP Local Preference. https://www.cisco.com/c/en/us/td/docs/ios-xml/ ios/iproute_bgp/configuration/15-mt/irg-15-mt-book/irg-external-sp.html# GUID-CD3F70AE-C92B-4A6D-AE85-460B6BCDBB17, 2014. [9] Cisco. Cisco BGP Route Aggregation. https://www.cisco.com/c/en/us/support/docs/ip/ border-gateway-protocol-bgp/5441-aggregation.html, 2024. [10] Cisco. Cisco BGP Multi-Exit Discriminator (MED). https://www.cisco.com/c/en/us/support/ docs/ip/border-gateway-protocol-bgp/13759-37.html, 2024. [11] Cisco. Cisco ACL. https://www.cisco.com/c/en/us/support/docs/security/ios-firewall/ 23602-confaccesslists.html, 2025. [12] Cisco. Cisco BGP. https://www.cisco.com/c/en/us/support/docs/ip/ border-gateway-protocol-bgp/13751-23.html, 2025. [13] Cisco. Cisco BGP Network Advertisement. https://www.cisco.com/c/en/us/support/docs/ip/ border-gateway-protocol-bgp/16137-cond-adv.html, 2025. [14] Cisco. Cisco Interface Shutdown. https://www.cisco.com/E-Learning/bulk/public/tac/cim/ cib/using_cisco_ios_software/cmdrefs/shutdown.htm, 2025. [15] Google Cloud. An Update on Sunday’s Service Disruption. https://cloud.google.com/blog/ topics/inside-google-cloud/an-update-on-sundays-service-disruption, 2019. [16] Cloudflare. Understanding How Facebook Disappeared from the Internet. https://blog.cloudflare. com/october-2021-facebook-outage/, 2021. [17] DeepSeek-AI. DeepSeek V4 Preview Release. news260424, 2026.
https://api-docs.deepseek.com/news/
[18] Wenlong Ding, Jianqiang Li, Zhixiong Niu, Huangxun Chen, Yongqiang Xiong, and Hong Xu. Automating Conflict-Aware ACL Configurations with Natural Language Intents. arXiv preprint arXiv:2508.17990, 2025. [19] Denis Donadel, Francesco Marchiori, Luca Pajola, and Mauro Conti. Can LLMs Understand Computer Networks? Towards a Virtual System Administrator. In 2024 IEEE 49th Conference on Local Computer Networks (LCN), 2024.
10
[20] Ahmed El-Hassany, Petar Tsankov, Laurent Vanbever, and Martin Vechev. NetComplete: Practical Network-Wide Configuration Synthesis with Autocompletion. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2018. [21] Ari Fogel, Stanley Fung, Luis Pedrosa, Meg Walraed-Sullivan, Ramesh Govindan, Ratul Mahajan, and Todd Millstein. A General Approach to Network Configuration Analysis. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2015. [22] Aaron Gember-Jacobson, Raajay Viswanathan, Aditya Akella, and Ratul Mahajan. Fast Control Plane Analysis Using an Abstract Representation. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2016. [23] Chi-Yao Hong, Srikanth Kandula, Ratul Mahajan, Ming Zhang, Vijay Gill, Mohan Nanduri, and Roger Wattenhofer. Achieving High Utilization with Software-Driven WAN. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2013. [24] Karthick Jayaraman, Nikolaj Bjørner, Jitu Padhye, Amar Agrawal, Ashish Bhargava, Paul-Andre C Bissonnette, Shane Foster, Andrew Helwer, Mark Kasten, Ivan Lee, Anup Namdhari, Haseeb Niaz, Aniruddha Parkhi, Hanukumar Pinnamraju, Adrian Power, Neha Milind Raje, and Parag Sharma. Validating datacenters at scale. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2019. [25] Srikanth Kandula, Sudipta Sengupta, Albert Greenberg, Parveen Patel, and Ronnie Chaiken. The Nature of Data Center Traffic: Measurements & Analysis. In Proceedings of the ACM Internet Measurement Conference (IMC), 2009. [26] Peyman Kazemian, George Varghese, and Nick McKeown. Header Space Analysis: Static Checking for Networks. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2012. [27] John Kim, William J. Dally, Steve Scott, and Dennis Abts. Technology-driven, highly-scalable dragonfly topology. In Proceedings of the ACM/IEEE International Symposium on Computer Architecture (ISCA), 2008. [28] Umesh Krishnaswamy, Rachee Singh, Paul Mattes, Paul-Andre C Bissonnette, Nikolaj Bjørner, Zahira Nasrin, Sonal Kothari, Prabhakar Reddy, John Abeln, Srikanth Kandula, Himanshu Raj, Luis Irun-Briz, Jamie Gaudette, and Erica Lan. OneWAN is Better than Two: Unifying a Split WAN Architecture. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2023. [29] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. RetrievalAugmented Generation for Knowledge-Intensive NLP Tasks. arXiv preprint arXiv:2005.11401, 2020. [30] Zechun Li, Peng Zhang, Yichi Zhang, and Hongkun Yang. NDD: A Decision Diagram for Network Verification. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2025. [31] Jianmin Liu, Li Chen, Dan Li, and Yukai Miao. CEGS: Configuration Example Generalizing Synthesizer. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2025. [32] Dimitrios Michael Manias, Ali Chouman, and Abdallah Shami. Towards Intent-Based Network Management: Large Language Models for Intent Extraction in 5G Core Networks. In 2024 20th International Conference on the Design of Reliable Communication Networks (DRCN), 2024. [33] Meta. Llama-3.3-70B-Instruct. 3-70B-Instruct, 2024.
https://huggingface.co/meta-llama/Llama-3.
[34] Microsoft. Azure Global DC Network. global-infrastructure, 2025.
https://azure.microsoft.com/en-us/explore/
[35] Rajdeep Mondal, Alan Tang, Ryan Beckett, Todd Millstein, and George Varghese. What do LLMs Need to Synthesize Correct Router Configurations? In Proceedings of the ACM Workshop on Hot Topics in Networks (HotNets), 2023. [36] Juniper Networks. Juniper Route Aggregation. https://www.juniper.net/documentation/us/ en/software/junos/static-routing/topics/topic-map/config-route-aggregation.html, 2025.
11
[37] Juniper Networks. Juniper BGP Local Preference. https://www.juniper.net/documentation/us/ en/software/junos/bgp/topics/topic-map/local-preference.html, 2025. [38] Juniper Networks. Juniper BGP MED Attribute. https://www.juniper.net/documentation/us/ en/software/junos/bgp/topics/topic-map/med-attribute.html, 2025. [39] Juniper Networks. Juniper BGP Network Advertisement. https://www.juniper. net/documentation/us/en/software/junos/routing-policy/bgp/topics/example/ bgp-advertise-peer-as.html, 2025. [40] Juniper Networks. Juniper Interface Shutdown. https://www.juniper.net/ documentation/us/en/software/junos/cli-reference/topics/ref/statement/ interface-shutdown-action-edit-switch-options.html, 2025. [41] Juniper Networks. Juniper Static Routes. https://www.juniper.net/documentation/us/en/ software/junos/static-routing/topics/topic-map/config_static-routes.html, 2025. [42] OpenAI. OpenAI ChatGPT. https://openai.com/chatgpt/overview/, 2025. [43] Prakhar Sharma and Vinod Yegneswaran. PROSPER: Extracting Protocol Specifications Using Large Language Models. In Proceedings of the ACM Workshop on Hot Topics in Networks (HotNets), 2023. [44] Arjun Singh, Joon Ong, Amit Agarwal, Glen Anderson, Ashby Armistead, Roy Bannon, Seb Boving, Gaurav Desai, Bob Felderman, Paulie Germano, Anand Kanagala, Jeff Provber, Jason Simmons, Eiichi Tanda, Jim Wanderer, Urs Hölzle, Stephen Stuart, and Amin Vahdat. Jupiter Rising: A Decade of Clos Topologies and Centralized Control in Google’s Datacenter Network. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2015. [45] Steffen Smolka, Praveen Kumar, Nate Foster, Dexter Kozen, and Alexandra Silva. Cantor Meets Scott: Semantic Foundations for Probabilistic Networks. In Proceedings of the ACM SIGPLAN Symposium on Principles of Programming Languages (POPL), 2017. [46] Samuel Steffen, Timon Gehr, Petar Tsankov, Laurent Vanbever, and Martin Vechev. Probabilistic Verification of Network Configurations. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2020. [47] Kausik Subramanian, Anubhavnidhi Abhashkumar, Loris D’Antoni, and Aditya Akella. Detecting Network Load Violations for Distributed Control Planes. In Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), 2020. [48] Geng Sun, Yixian Wang, Dusit Niyato, Jiacheng Wang, Xinying Wang, H. Vincent Poor, and Khaled B. Letaief. Large Language Model (LLM)-enabled Graphs in Dynamic Networking. arXiv preprint arXiv:2407.20840, 2024. [49] Qwen Team. Qwen3.5-397B-A17B. https://huggingface.co/Qwen/Qwen3.5-397B-A17B, 2026. [50] Changjie Wang, Mariano Scazzariello, Alireza Farshin, Simone Ferlin, Dejan Kostić, and Marco Chiesa. NetConfEval: Can LLMs Facilitate Network Configuration? Proceedings of the ACM on Networking, 2 (CoNEXT2):1–25, June 2024. [51] Dan Wang, Peng Zhang, Wenbing Sun, Wenkai Li, Xing Feng, Hao Li, Jiawei Chen, Weirong Jiang, and Yongping Tang. S2: A Distributed Configuration Verifier for Hyper-Scale Networks. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2025. [52] Haopei Wang, Anubhavnidhi Abhashkumar, Changyu Lin, Tianrong Zhang, Xiaoming Gu, Ning Ma, Chang Wu, Songlin Liu, Wei Zhou, Yongbin Dong, Weirong Jiang, and Yi Wang. NetAssistant: Dialogue Based Network Diagnosis in Data Center Networks. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2024. [53] Zhaodong Wang, Samuel Lin, Guanqing Yan, Soudeh Ghorbani, Minlan Yu, Jiawei Zhou, Nathan Hu, Lopa Baruah, Sam Peters, Srikanth Kamath, Jerry Yang, and Ying Zhang. Intent-Driven Network Management with Multi-Agent LLMs: The Confucius Framework. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2025. [54] Zhihao Wang, Alessandro Cornacchia, Alessio Sacco, Franco Galante, Marco Canini, and Dingde Jiang. A Network Arena for Benchmarking AI Agents on Network Troubleshooting. arXiv preprint arXiv:2512.16381, 2025.
12
[55] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems (NeurIPS), 2022. [56] Duo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang, Junchen Jiang, Shuguang Cui, and Fangxin Wang. NetLLM: Adapting Large Language Models for Networking. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2024. [57] Han Xu, Zachary Kincaid, Ratul Mahajan, and David Walker. Network Change Validation with Relational NetKAT. Proceedings of the ACM on Programming Languages, 10(14):384–412, January 2026. [58] Xieyang Xu, Yifei Yuan, Zachary Kincaid, Arvind Krishnamurthy, Ratul Mahajan, David Walker, and Ennan Zhai. Relational Network Verification. In Proceedings of the ACM Conference on Special Interest Group on Data Communication (SIGCOMM), 2024. [59] Yitao Yang, Yangtao Deng, Yifan Xiong, Baochun Li, Hong Xu, and Peng Cheng. AidAI: Automated Incident Diagnosis for AI Workloads in the Cloud. arXiv preprint arXiv:2506.01481, 2025. [60] Peng Zhang, Aaron Gember-Jacobson, Yueshang Zuo, Yuhao Huang, Xu Liu, and Hao Li. Differential Network Analysis. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2022. [61] Yajie Zhou, Jiajun Ruan, Eric S. Wang, Sadjad Fouladi, Francis Y. Yan, Kevin Hsieh, and Zaoxing Liu. NetArena: Dynamic Benchmarks for AI Agents in Network Automation. In Proceedings of the International Conference on Learning Representations (ICLR), 2026.
13
A
Additional Production Network Context
Table 6: Frequency of configuration changes in a global-scale production cloud network in the week of Feb. 26, 2025. Device metadata updates occur hourly; routing and filtering changes occur multiple times per day. Config Type
Remark
Device Metadata & Templates Routing Protocols Packet Filtering Others
Prefix & device addition/removal, and device template setup. BGPs, IGPs, static routes, etc. ACLs, firewalls, etc. NAT, device security config, etc.
Percentage
Frequency
71.2%
1.28 / hour
8.9% 8.3% 11.6%
3.86 / day 3.57 / day 5.00 / day
Configuration Change Frequency. We measure how often configuration changes occur in a production network, since such changes can trigger path analysis to verify intended behavior. As shown in Table 6, device-metadata changes, such as location and prefix updates, are the most frequent and occur hourly, prompting operators to check whether network paths are affected. Routing-protocol and packet-filtering changes also occur multiple times per day, each potentially triggering path analysis to verify intended forwarding and reachability. Overall, operators need path analysis at least hourly to keep pace with these changes while still leaving time to diagnose and resolve unexpected paths.
B
Benchmark Generation Details
All topology and configuration instances in ROUTING B ENCH are produced by deterministic scripts parameterized by the fat-tree parameter k (even, ≥ 4). This appendix details three layers of the generation pipeline: the physical topology and device/interface naming (§B.1), the connectivity configuration that ensures full eBGP reachability (§B.2), and the six types of flow-routing configuration changes used to construct benchmark tasks (§B.3). B.1
Fat-Tree Topology Generation
Device naming. Each device is named by its role and position in the fat-tree hierarchy. Edge routers are edge-p{pod}-e{idx}, aggregation routers are agg-p{pod}-a{idx}, and core routers are core-g{group}-c{idx}, where pod ∈ [0, k), idx ∈ [0, k/2), and group ∈ [0, k/2). Server hosts follow the pattern server-p{pod}-e{edge}-h{host}. For inter-datacenter topologies, every device name is prefixed with a datacenter identifier dc{d}-, yielding names such as dc0-edge-p0-e0 and dc1-agg-p3-a1. This prefix is the sole mechanism that distinguishes devices in the two datacenters—all index ranges remain identical, so dc0-edge-p5-e2 and dc1-edge-p5-e2 occupy the same structural position in their respective data centers. Each datacenter additionally contains k gateway routers named dc{d}-gw{g} (g ∈ [0, k)). Gateways connect to all (k/2)2 cores in the local datacenter and to every gateway in the remote datacenter, providing inter-datacenter reachability. Interface naming. We number each interface by the order in which it is added to the device, so the same k always yields the same names. The order itself encodes role: lower-numbered interfaces face the layer below, higher-numbered ones face the layer above (e.g., on edge routers we add server-facing interfaces first, then uplinks to aggregation routers). Names may repeat across datacenters (e.g., both dc0-gw0 and dc1-gw0 have an interface GigabitEthernet1/0/1) without ambiguity because the device hostnames already carry the dc{d}- prefix. Topology construction. The topology is built in three steps following the standard fat-tree wiring (Figure 5): (1) each edge router connects to k/2 servers; (2) within each pod, edge and aggregation routers form a complete bipartite graph; (3) each aggregation router with index a connects to all k/2 core routers in core group a. The output is a JSON file listing every link as a pair of (hostname, interfaceName) endpoints: 14
Topology link entry (JSON) 1 {"node1": {"hostname": "edge-p0-e0", 2 "interfaceName": "GigabitEthernet1/0/3"}, 3 "node2": {"hostname": "agg-p0-a0", 4 "interfaceName": "GigabitEthernet1/0/1"}}
For inter-datacenter, the script additionally wires each of the k gateways per datacenter to all (k/2)2 local cores, and establishes k 2 fully connected inter-datacenter gateway links, producing a single unified topology JSON.
B.2
Network Connectivity Configuration
In ROUTING B ENCH, we generate device configurations using a Cisco IOS-style syntax [11, 12], reflecting one of the most widely used router configuration formats.
Conflict-free IP addressing. All IP addresses are derived deterministically from each device’s role and position, ensuring no prefix collisions across arbitrary k. Table 7 summarizes the assignment scheme for intra-datacenter and inter-datacenter topologies.
Table 7: IP address assignment by network layer (cf. Figure 5). DC1 and DC2 refer to the two datacenters in the inter-datacenter topology; intra-datacenter uses only DC1. Each /31 link entry lists the lower-tier side first, then the upper-tier side. Indices: p=pod, e=edge, a=agg, c=core, g=gateway, grp=core group, seq=index of a core-to-gateway link within a gateway. Layer
Prefix assignment {DC1, DC2}
Server group Edge–Agg link Agg–Core link Core–GW link (inter-DC) GW–GW link (inter-DC) Edge Loopback Agg Loopback Core Loopback GW Loopback
{1,9}.p.e.0/24 (gateway .1, hosts .2+) {2,6}.p.a.{2e,2e+1}/31 {3,7}.p.c.{2a,2a+1}/31 {4,8}.0.g.{2 seq,2 seq+1}/31 5.g0 .g1 .{0,1}/31 10.p.{0,2}.e/32 10.p.{1,3}.a/32 10.{200,201}.grp.c/32 10.{250,251}.0.g/32
For intra-datacenter, server prefixes occupy the 1.x.x.x space, edge–agg links use 2.x.x.x, and agg–core links use 3.x.x.x. Because the even/odd split within each /31 is determined by device indices, addresses never collide regardless of k. For inter-datacenter, datacenter 1 retains the 1/2/3/4 octets while datacenter 2 mirrors them at 9/6/7/8, ensuring complete isolation between the two address domains. The inter-datacenter gateway links share the 5.x.x.x space. Scaling beyond k=256 via IPv6. Each index field in Table 7 occupies one IPv4 octet (one byte), so k caps at 256 (already over 80K routers). To scale further, we keep exactly the same layout as Table 7 but widen every index field from one byte to one IPv6 hextet (two bytes), pushing the cap to k=65,536. VLAN for server-edge connectivity. Server endpoints use a different convention from inter-router links: instead of giving each link its own /31 IPs, all servers attached to the same edge router are placed in a virtual group (VLAN), share one /24 subnet, and use the edge router as their common gateway. This is a standard configuration on the server side. Each edge router uses a unique group ID (1000+p · k/2+e), so no two groups in the network ever clash. 15
Table 8: ASN assignment scheme. d∈{0, 1} indexes the datacenter so that the two datacenters fall into non-overlapping ranges; intra-datacenter uses d=0. Device Layer
Formula
Rationale
Edge
100000 + d · 10000 + p · (k/2) + e
Agg
20000 + d · 1000 + p
Core
30000 + d
Gateway
40000 + d · 1000 + g
Every edge router gets its own ASN, so it never shares one with the aggregation routers it physically connects to. All agg routers in the same pod share an ASN, but it is in a different range from edge ASNs, so every edge–agg link is between two different ASNs (i.e., eBGP). All core routers in a datacenter share an ASN, distinct from any agg ASN, so every agg–core link is again between two different ASNs. Every gateway gets its own ASN, distinct from local cores and from every other gateway, so all core–gateway and gateway–gateway links are between two different ASNs.
eBGP session configuration. To ensure every inter-router link is an eBGP session (i.e., no two neighbors share an ASN), the ASN scheme assigns unique values per layer and scope. Table 8 lists the formulas. Each router runs eBGP session on every point-to-point /31 interface. See the template below: eBGP configuration template (aggregation router) 1 router bgp {ASN_AGG} 2 bgp router-id {AGG_LOOPBACK} 3 bgp log-neighbor-changes 4 ! 5 address-family ipv4 6 redistribute connected 7 neighbor {EDGE_IP} remote-as {ASN_EDGE} 8 neighbor {EDGE_IP} activate 9 neighbor {CORE_IP} remote-as {ASN_CORE} 10 neighbor {CORE_IP} activate 11 neighbor {CORE_IP} route-map {RM_OUT} out 12 aggregate-address {POD_PREFIX} {POD_MASK} as-set summary-only 13 maximum-paths 8 14 exit-address-family 15 !
The template parameters are filled as follows: ASN_AGG is computed from Table 8; AGG_LOOPBACK and EDGE_IP/CORE_IP are derived from Table 7; the route-map name encodes the pod and agg indices for traceability. eBGP configuration on agg-p0-a0 (k=4) 1 router bgp 20000 2 bgp router-id 10.0.1.0 3 bgp log-neighbor-changes 4 ! 5 address-family ipv4 6 redistribute connected 7 neighbor 2.0.0.0 remote-as 100000 8 neighbor 2.0.0.0 activate 9 neighbor 2.0.0.2 remote-as 100001 10 neighbor 2.0.0.2 activate 11 neighbor 3.0.0.1 remote-as 30000 12 neighbor 3.0.0.1 activate 13 neighbor 3.0.0.1 route-map RM_AGG_TO_CORE_0_0 out 14 neighbor 3.0.1.1 remote-as 30000 15 neighbor 3.0.1.1 activate 16 neighbor 3.0.1.1 route-map RM_AGG_TO_CORE_0_0 out 17 aggregate-address 1.0.0.0 255.255.0.0 as-set summary-only 18 maximum-paths 8 19 exit-address-family 20 !
16
Table 9: Six configuration types used to construct benchmark tasks (cf. Table 1). Configuration Type
Category
Effect
BGP Local Preference (LP) [8, 37] Multi-Exit Discriminator (MED) [10, 38]
Route preference
Static route [7, 41]
Route preference
BGP network advertisement [13, 39] Prefix aggregation [9, 36]
Route advertisement Route advertisement
Interface shutdown [14, 40]
Topology availability
On a given router, a flow takes the next hop whose route-map sets the highest LP for that flow. On a neighbor router, a flow chooses among the physical next hops connecting to it by comparing the MED values they advertise, and selects the one with the lowest MED. Forces a flow onto a designated next hop, overriding any BGPselected route. Controls whether a router forwards a given flow via BGP; removing the network entry stops that router from forwarding the flow. Summarizes multiple smaller flows into one aggregated flow forwarded via BGP; once a flow is covered by an aggregate, its own paths are suppressed and its forwarding follows the aggregated flow instead. Physically disables an interface, so any flow traversing it can no longer be forwarded.
B.3
Route preference
Flow Routing Configurations
This subsection describes how the six configuration change types (Table 9) are instantiated as benchmark tasks. Each type provides a template parameterized by a source–destination flow S(PS ) → D(PD ), and shifts traffic to a different device by manipulating one specific routing attribute. BGP Local Preference (LP). The template is filled with: (i) ASN: the router’s own ASN (Table 8); (ii) NEI_IP: the link IP of the chosen neighbor (Table 7); (iii) LP_HIGH: any value above the default LP. The effect: among all of this router’s neighbors, the one at NEI_IP now has the highest LP for the queried flow, so the flow is shifted onto that neighbor.
LP template 1 2 3 4 5 6 7
route-map {RM_NAME} permit 10 set local-preference {LP_HIGH} route-map {RM_NAME} permit 20 ! router bgp {ASN} address-family ipv4 neighbor {NEI_IP} route-map {RM_NAME} in
This instance makes agg-p0-a1 the preferred next hop, shifting the flow from agg-p0-a0/core group 0 to agg-p0-a1/core group 1.
LP instance: on edge-p0-e0, shift the flow to agg-p0-a1 1 2 3 4 5 6 7
route-map RM_LP_AGG1 permit 10 set local-preference 300 route-map RM_LP_AGG1 permit 20 ! router bgp 100000 address-family ipv4 neighbor 2.0.1.1 route-map RM_LP_AGG1 in
Multi-Exit Discriminator (MED). The template is filled with: (i) ASN: the router’s own ASN; (ii) NEI_IP: the link IP of the neighbor whose entry path is being penalized; (iii) MED_HIGH: any large value (lower MED is preferred in BGP). The effect: routes received from NEI_IP carry a high MED, so among all the parallel links between the two ASes, the neighbor will pick a different (lower-MED) link to send the flow over. 17
MED template 1 2 3 4 5 6 7
route-map {RM_NAME} permit 10 set metric {MED_HIGH} route-map {RM_NAME} permit 20 ! router bgp {ASN} address-family ipv4 neighbor {NEI_IP} route-map {RM_NAME} in
This instance penalizes routes from core-g0-c0, so agg-p0-a0 shifts the flow to core-g0-c1 within the same core group. MED instance: on agg-p0-a0, push traffic away from core-g0-c0 1 2 3 4 5 6 7
route-map RM_MED_CORE0 permit 10 set metric 9999 route-map RM_MED_CORE0 permit 20 ! router bgp 20000 address-family ipv4 neighbor 3.0.0.1 route-map RM_MED_CORE0 in
Static route. The template is filled with: (i) DST_NET/DST_MASK: the destination prefix of the queried flow; (ii) NEXT_HOP_IP: the link IP of the chosen neighbor (Table 7). The effect: any packet matching DST_NET is forwarded to NEXT_HOP_IP, overriding the BGP-selected next hop. Static route template 1 ip route {DST_NET} {DST_MASK} {NEXT_HOP_IP}
This instance forces agg-p0-a0 to send traffic for 1.5.3.0/24 directly to core-g0-c5, overriding its BGP choice. Static route instance: on agg-p0-a0, pin the destination to core-g0-c5 1 ip route 1.5.3.0 255.255.255.0 3.0.5.1
BGP network advertisement. The template is filled with: (i) ASN: the router’s own ASN; (ii) DST_NET/DST_MASK: a prefix more specific than the default fabric aggregate; (iii) NEXT_HOP_IP: a neighbor IP that anchors this prefix in the local routing table. The ip route line provides the backing route required by Cisco IOS before a network statement can originate the prefix. The effect: this router announces DST_NET into BGP, and because longest-prefix match wins, peers steer the corresponding flow to it. BGP network template 1 ip route {DST_NET} {DST_MASK} {NEXT_HOP_IP} 2 ! 3 router bgp {ASN} 4 address-family ipv4 5 network {DST_NET} mask {DST_MASK}
This instance makes agg-p0-a1 advertise the more-specific 1.5.3.0/24, attracting the flow to agg-p0-a1 via longest-prefix match. BGP network instance: on agg-p0-a1, originate 1.5.3.0/24 1 ip route 1.5.3.0 255.255.255.0 3.0.0.3 2 ! 3 router bgp 20000 4 address-family ipv4 5 network 1.5.3.0 mask 255.255.255.0
18
Prefix aggregation. The fields are filled in the same way as the BGP-network template above; the only addition is the aggregate-address line on the same prefix. The effect: any flow falling inside DST_NET loses its individual paths and is forwarded along the aggregated route originated by this router. Prefix aggregation template 1 ip route {DST_NET} {DST_MASK} {NEXT_HOP_IP} 2 ! 3 router bgp {ASN} 4 address-family ipv4 5 network {DST_NET} mask {DST_MASK} 6 aggregate-address {DST_NET} {DST_MASK} as-set
This instance makes agg-p0-a1 originate and aggregate 1.5.3.0/24, causing matching traffic to follow agg-p0-a1’s aggregated route. Prefix aggregation instance: on agg-p0-a1, aggregate 1.5.3.0/24 1 ip route 1.5.3.0 255.255.255.0 3.0.0.3 2 ! 3 router bgp 20000 4 address-family ipv4 5 network 1.5.3.0 mask 255.255.255.0 6 aggregate-address 1.5.3.0 255.255.255.0 as-set
Interface shutdown. The template is filled with one field, INTF_NAME, the name of an interface on a router along the queried path. The effect: the corresponding link disappears from the topology, so any flow that previously traversed it must take an alternate route. Interface shutdown template 1 interface {INTF_NAME} 2 shutdown
This instance disables the default link from edge-p0-e0 to agg-p0-a0, forcing affected traffic onto another available aggregation uplink. Interface shutdown instance: on edge-p0-e0, disable the link toward agg-p0-a0 1 interface GigabitEthernet1/0/3 2 shutdown
Change instantiation. For each flow, we first identify the configuration template that currently controls its forwarding behavior, then inject the change by appending the new template instance before the device configuration’s trailing end. This append-only procedure mirrors how operators apply incremental updates through command-line configuration tools: the newly added commands override the earlier template for the affected flow without rewriting the original configuration file.
C
Agent Design Details
Each agent consists of a guidance prompt and a execution script; Agent Skills additionally uses skill documents and tools. The prompt combines shared task instructions (reasoning, schema, examples, and references) with an agent-specific AGENT.md that defines its evidence-access workflow. The script orchestrates LLM calls, parses JSON actions, dispatches operations (configuration retrieval, tool calls, or skill loading), and enforces interaction limits; agents share this structure but differ in action vocabulary and dispatch logic. All agents are evaluated by a common framework that loads ground-truth instances, runs each agent, and compares normalized path sequences. 19
C.1
Shared Guidance Prompt
The shared guidance prompt provides four categories of information that ground agentic path-analysis tasks regardless of which agent architecture is used. We summarize them below. • Input/output schema. The prompt defines the task (analyze before/after forwarding paths for a source–destination pair given a configuration change) and specifies the input and output format. The input provides source and destination edge routers, the queried prefixes, and a list of configuration diffs. If source or destinaion prefixes are not specified, we default them to “any” (0.0.0.0/0). The output is a JSON object: Input Schema (JSON) 1 {"query": {"src": "edge-p0-e0", 2 "dst": "edge-p1-e0", 3 "prefix": "1.1.0.0/24"}, 4 "change": [{"device": "edge-p0-e0", 5 "add_lines": ["route-map RM_LP_AGG1 permit 10", 6 " set local-preference 300", 7 "router bgp 100000", 8 " address-family ipv4", 9 " neighbor 2.0.1.0 route-map RM_LP_AGG1 in" 10 ], 11 "description": "LP=300 on edge inbound from agg1"}]}
Output Schema (JSON) 1 { 2 "before_path": ["edge-p0-e0", "agg-p0-a0", 3 "core-g0-c0", "agg-p1-a0", "edge-p1-e0"], 4 "after_path": ["edge-p0-e0", "agg-p0-a1", 5 "core-g1-c0", "agg-p1-a1", "edge-p1-e0"], 6 "decisive_evidence": "LP=300 from agg1 overrides 7 LP=200 from agg0 on edge-p0-e0. Core group shifts from 0 to 1." 8 }
• Chain-of-thought (CoT) guidance. All agents embed a structured hop-by-hop reasoning skeleton that instructs the LLM to trace the forwarding path step by step: CoT Reasoning Skeleton 1 2 3 4 5 6 7 8 9 10 11 12
1. Identify source and destination edge routers. 2. BEFORE state: trace hop-by-hop from src_edge. At each router: - Check static routes (AD=1 overrides BGP AD=20) - Apply longest-prefix match (/24 > /16 > /8) - Check BGP: route-map -> LP from each neighbor - If no LP set, check MED; if no MED, lowest router-ID tiebreak - agg index = core group = dst_agg index 3. AFTER state: apply the config change, re-trace with the same logic. 4. Report both paths and the decisive evidence.
The skeleton encodes the route-selection priority (static AD=1 beats BGP, longest-prefix match precedes BGP tie-breaking; within BGP: highest LP > shortest AS-path > lowest MED > lowest router-ID) and the fat-tree structural invariant that the aggregation index determines the core group and destination aggregation. • Few-shot demonstrations. Each prompt includes worked examples on a small topology covering common analysis patterns and preventing recurring errors. For intra-datacenter fat-tree, five examples demonstrate LP on edge (shifting agg), LP on agg (shifting core), interface shutdown, BGP network with longest-prefix match, and static route overriding BGP. A representative example: 20
Few-shot Example: LP on edge shifts agg 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17
Query: src=edge-p0-e0, dst=edge-p1-e0, prefix=1.1.0.0/24 Change: On edge-p0-e0, add route-map RM_LP_AGG1 with set local-preference 300, applied inbound from agg-p0-a1. Analysis: - BEFORE: edge has LP=200 from agg0 (RM_PIN), LP=100(default) from agg1. Picks agg0 (200>100). agg0 -> core-g0-c0 (lowest RID). dst: agg-p1-a0. - AFTER: LP=300 from agg1, LP=200 from agg0. Picks agg1 (300>200). agg1 -> core-g1-c0. dst: agg-p1-a1. Answer: before: [edge-p0-e0, agg-p0-a0, core-g0-c0, agg-p1-a0, edge-p1-e0] after: [edge-p0-e0, agg-p0-a1, core-g1-c0, agg-p1-a1, edge-p1-e0]
These examples are chosen to demonstrate non-obvious patterns (e.g., /24 beating /8 via longestprefix match regardless of LP; static route AD=1 overriding all BGP attributes) that LLMs tend to get wrong without explicit guidance. • Reference material. Following common agentic-framework practice, the prompt provides background reference on the specific topology structure (device naming, connectivity rules, path length) and protocol details (the six configuration mechanisms and their routing effects). Task-specific variations. The shared prompt is tailored for each of the three task types—intradatacenter fat-tree, inter-datacenter, and compound changes—with the differences concentrated in the reference material and few-shot examples. Inter-datacenter prompts describe the dual-datacenter topology (device prefixes dc0-/dc1-, gateway interconnection) and include examples demonstrating that changes in datacenter-0 do not affect datacenter-1’s routing. Compound-change prompts use the same intra-datacenter topology reference but include examples with 2–4 simultaneous changes, illustrating how to analyze each change independently and then compose their combined effect on the path. C.2
In-Context Learning
The In-Context Learning agent places the shared guidance, full topology, all router configurations, query, and configuration change into one prompt. Its minimal AGENT.md asks the model to return the JSON answer in a single LLM call, with no follow-up interaction. C.3
Iterative Reasoning
The Iterative Reasoning agent starts with the full topology and query but no device configurations. Across up to K=10 rounds, it requests device files via request_configs; the script returns the relevant .cfg files and iterates until the model answers or the round limit is reached. C.4 C.4.1
Agent Skills Agent Workflow
Agent-specific guidance. The Agent Skills AGENT.md defines a plan-invoke-reason loop and lists the 14 available skills by category (configuration extraction, topology, and verification). The model selects a skill with {"action": "use_skill", "skill": "skill-name"}, receives its SKILL.md workflow, and then calls tools as JSON arrays [{"tool": "tool_name", "args": {...}}, ...]. Before an answer is accepted, the runtime enforces a device-inspection requirement that grows with network size. The AGENT.md also includes accumulated don’t rules (§C.4.3) to prevent recurring topology and tool-use errors. The initial prompt provides no topology graph or device configurations; all evidence must be acquired through skill and tool invocations. Execution script. The execution script deterministically parses each JSON action: use_skill returns the requested SKILL.md, tool-call arrays execute stateless scripts over the network files, and 21
Table 10: Low-level tools available to skill workflows. Tool
Description
search_config(device, regex)
Regex search across a device’s .cfg file. Returns only matching lines with surrounding context. Extract a named configuration block (e.g., route-map, prefix-list, router bgp, community-list). Parse the router bgp section and return configured BGP neighbors with applied route-maps (in/out). Return interface IP address, shutdown status, and description. From the topology JSON, return all neighbors: (neighbor, local_intf, remote_intf). Return devices in a synthetic fat-tree layer (edge/agg/core), optionally scoped to a pod. Enumerate topology-valid candidate paths between two endpoints. Return configuration lines that reference a named object (route-map, prefix-list, community-list, interface). Return the match and set clauses of a route-map, grouped by sequence number.
read_config_block(device, block_type, name) get_bgp_neighbors(device) get_interface_config(device, intf) get_topology_adjacency(device) get_layer_devices(layer, scope) get_topology_paths(src, dst) find_references(device, symbol) extract_match_set_clauses(device, route_map)
answer terminates the loop. The loop stops on a valid answer or after S=100 steps, then forces a final answer with up to 3 retries for output format errors. C.4.2
Skills and Tools
Skills and tools form a two-layer architecture that separates domain reasoning from mechanical data access. Tools. A tool is a stateless, deterministic script that performs a narrow operation on network data (topology JSON or device .cfg files) and returns structured, focused results. Nine tools are available (Table 10). Tools perform no reasoning—they search, parse, and extract only the evidence requested. Skills. A skill is a reusable multi-step workflow specification stored as a SKILL.md document. Each skill defines a high-level objective, describes the reasoning chain from objective to evidence, lists which tools to call at each step, and specifies the structured result to return. The 14 skills are organized into four categories. The critical value of skills is that they encode the domain reasoning chain connecting a high-level question to scattered configuration evidence—for instance, from “LP for prefix P ” to BGP neighbors, inbound route-maps, and their match/set clauses. Config extraction skills (7 skills). These skills navigate the indirect, reference-heavy structure of device configurations to extract specific routing attributes. • find-lp-for-prefix(device, prefix): LP is typically applied indirectly via a route-map attached to a BGP neighbor. The workflow is: (1) call get_bgp_neighbors to find peers and inbound route-maps; (2) call extract_match_set_clauses on each route-map; (3) resolve referenced prefix-lists if needed; (4) report any set local-preference values, or default LP=100. • find-med-for-prefix(device, prefix): Follows the same route-map workflow as LP, but extracts set metric clauses and reports MED values. • find-static-route(device, prefix): Call search_config for ip route entries matching the prefix and parse the network, mask, and next-hop IP. • find-bgp-network(device, prefix): Call search_config for network entries matching the prefix. 22
• find-prefix-aggregation(device, prefix): Search for aggregate-address entries and report the configured aggregates and whether summary-only is present. • find-route-map-policy(device, neighbor_ip): Call get_bgp_neighbors to locate the target peer by IP, then use extract_match_set_clauses to return its inbound and outbound route-map clauses. • check-interface-reachability(device, intf): Call get_interface_config(device, intf) to check IP assignment and shutdown status. Topology extraction skills (3 skills). These skills expose network structure at different levels of abstraction. • find-neighbors(device): Call get_topology_adjacency(device) to list all neighbors; optionally call get_interface_config on each to enrich with IP and status information. • find-layer-devices(layer, scope): Call get_layer_devices(layer, scope) to return devices in the specified synthetic fat-tree layer and optional pod scope. • find-candidate-paths(src, dst, prefix): Call get_topology_paths(src, dst) and keep paths with valid hop counts (3 for intra-pod, 5 for inter-pod, 8 for cross-datacenter). Config verification skills (2 skills). These skills answer higher-level decision questions by composing extraction skills and applying routing protocol knowledge. • determine-route-selection(device, prefix): Compose static-route, BGP-network, LP, MED, aggregation, and neighbor-status evidence, then report the winning route and next-hop device. • verify-bgp-session(device, neighbor): Call get_bgp_neighbors on both devices and report whether both sides expose BGP neighbor state. Topology verification skills (2 skills).
These skills verify physical-layer reachability.
• verify-link(device1, device2): Call get_topology_adjacency to confirm the link exists; call get_interface_config on both ends to check that neither interface is administratively shut down. • verify-path-reachability(path): For each consecutive pair of routers in path, invoke verify-link; report the first broken link if any. C.4.3
Don’t Rules
As our primary agent design, Agent Skills differs fundamentally from the other two agents in that it starts with no initial network information—no topology graph, no device configurations. All evidence must be acquired through skill and tool invocations, which means the agent’s accuracy depends entirely on its planning quality: which devices to query, which routing attributes to check, and how to interpret the results. During development, we observed that the agent’s errors are largely caused by two categories of uncontrolled variables: topology-induced failures (misunderstanding the fat-tree connectivity structure) and tool-induced failures (misusing tool APIs or misinterpreting tool results). These failures are not inherent to the agent’s reasoning capability; rather, they stem from unfamiliarity with benchmark-specific topology conventions already provided to the other agentic methods and with tool interfaces shaped by the initial prompt design. To isolate the agent’s true analytical capability from such confounding factors, we added a set of don’t rules that encode recurring anti-patterns observed during development: • Topology-induced rules specify fat-tree connectivity facts that the agent often infers incorrectly. For example, the selected core is determined by the lowest router-ID within a core group, not by matching the core index to the aggregation index; inter-pod paths must also satisfy src_agg index == core group == dst_agg index. 23
• Tool-induced rules specify how to use and interpret tool outputs. For example, find-candidate-paths only enumerates topology-valid paths and does not account for interface shutdowns, so shutdown changes require explicit reachability checks; other rules prevent incorrect tool argument names. We also include output-consistency rules, such as re-reading decisive_evidence against the JSON paths before submitting and avoiding router-ID tiebreaks when a static route deterministically overrides BGP. These rules eliminate confounding variables and ensure that evaluation measures the agent’s path-analysis capability rather than its familiarity with benchmark-specific conventions.
D
Complete Experimental Results
This appendix gives the per-LLM, per-k, per-agentic-method results introduced in §5, plus the concrete performance-metric values of the real-world hyperscale evaluation. Overall, these complete results support all claims made in §5. D.1
Real-World Hyperscale Evaluation
Table 11 reports the performance values from our real-world hyperscale evaluation, complementing §5.3. The accuracy closely aligns with ROUTING B ENCH, showing that ROUTING B ENCH can effectively guide agent design for real production networks.
Table 11: Agent Skills (GPT-5.4) performance on the real-world production network reported in §5.3.
D.2
Scenario
Devices
Accuracy
Tokens
tLLM
tTotal
RLLM
Intra-datacenter Inter-datacenter
∼10K ∼20K
92% 73%
236K 424K
719 s 1293 s
1023 s 1762 s
5.9 7.5
Three-Agent and Six-LLM Comparison on Intra-Datacenter Single-Change Analysis
We present the complete results of the three agents under all six LLMs on intra-datacenter singlechange analysis across all scales (from k=4 to k=200) in Tables 12–17. These results further validate the design principle from §5.2: “explore more; digest less.” For In-Context Learning and Iterative Reasoning, all six LLMs still suffer severe cost-side scalability issues: tokens and runtime grow rapidly, and both methods hit context-window overflow or the time limit at small scales. Specifically, In-Context Learning stops scaling at k=16 (320 routers) on the strongest LLMs (GPT-5.4 and DeepSeek-V4-Pro) and at k=8 (80 routers) on all the others (GPT-4o, OpenAI o3, Qwen-3.5-397B, and Llama-3.3-70B); Iterative Reasoning stops at k=20 (500 routers) on most LLMs and at k=16 (320 routers) on GPT-4o and Llama-3.3-70B. On the accuracy side, both methods also degrade severely with scale on weaker LLMs (e.g., GPT-4o and Llama-3.3-70B) and on those that tend to overthink (e.g., Qwen-3.5-397B). Agent Skills, in contrast, scales to at least k=120 (18K routers) on every LLM, demonstrating its strong scalability. Except for the weakest LLM (Llama-3.3-70B), all LLMs keep their accuracy stable (mostly >90%); GPT-5.4 and the thinking models (OpenAI o3, DeepSeek-V4-Pro, and Qwen-3.5397B) approach 100%, and token consumption stays nearly flat across scales—far less explosive than under the other two methods. The thinking models, however, exhibit instability relative to the general-purpose GPT-5.4: they sometimes spiral into excessive internal reasoning on individual cases, causing token/runtime spikes (e.g., OpenAI o3 shows an abnormal token spike at k=16, jumping to 272K from 123K at k=12 and 134K at k=20, which we traced to unusually long thinking on a few random cases). Llama-3.3-70B, the weakest model, fluctuates in both accuracy and tokens at every scale, with some correct answers coming from guessing rather than reasoning. Across all scales, GPT-5.4 is the most stable in both accuracy and token cost, confirming the agent design principle from §5.2: “explore more; digest less.” 24
Table 12: Three-agent comparison on intra-datacenter fat-tree with GPT-5.4 across all scales. k
In-Context Learning
Routers Acc
4 8 12 16 20 30 50 80 120 160 200
20 80 180 320 500 1,125 3,125 8,000 18,000 32,000 50,000
Tok
tRun
RLLM
99% 20K 66s 1.0 98% 100K 317s 1.0 94% 292K 905s 1.0 99% 644K 1982s 1.0 Context Window Overflow
Iterative Reasoning Acc
Tok
tRun
RLLM
100% 22K 87s 3.2 100% 50K 189s 3.1 99% 114K 388s 3.0 96% 234K 752s 3.0 100% 427K 1473s 2.9 Time Limit Exceeded
Agent Skills Tok
tRun
RLLM
100% 48K 98% 51K 100% 51K 100% 52K 100% 53K 100% 55K 99% 67K 99% 68K 99% 89K 100% 112K 99% 119K
187s 236s 246s 243s 554s 554s 663s 592s 903s 1630s 2754s
6.9 6.8 6.7 6.7 6.8 6.8 7.7 7.5 8.8 10.2 10.4
Acc
Table 13: Three-agent comparison on intra-datacenter fat-tree with GPT-4o across all scales. k
Routers Acc
4 8 12 16 20 30 50 80 120 160 200
20 80 180 320 500 1,125 3,125 8,000 18,000 32,000 50,000
In-Context Learning Tok tRun RLLM
78% 19K 26s 1.0 67% 100K 113s 1.0 Context Window Overflow
Acc
Iterative Reasoning Tok tRun RLLM
94% 26K 46s 3.4 67% 56K 98s 3.5 66% 140K 191s 3.7 52% 283K 334s 3.6 Context Window Overflow
Acc
Agent Skills Tok tRun RLLM
93% 98K 175s 12.9 91% 101K 250s 12.7 89% 105K 278s 12.9 89% 117K 294s 13.7 98% 117K 894s 13.4 96% 168K 1029s 15.1 96% 235K 1337s 18.6 96% 230K 1187s 18.6 95% 351K 2026s 23.3 Time Limit Exceeded
Table 14: Three-agent comparison on intra-datacenter fat-tree with OpenAI o3 across all scales. k
Routers Acc
4 8 12 16 20 30 50 80 120 160 200
20 80 180 320 500 1,125 3,125 8,000 18,000 32,000 50,000
In-Context Learning Tok tRun RLLM
98% 22K 40s 1.0 94% 102K 170s 1.0 Context Window Overflow
Acc
Iterative Reasoning Tok tRun RLLM
99% 23K 52s 2.8 99% 42K 94s 2.5 98% 85K 162s 2.3 98% 156K 268s 2.0 96% 272K 530s 1.9 Context Window Overflow
25
Acc
Agent Skills Tok tRun RLLM
99% 112K 247s 12.7 97% 98K 291s 12.0 93% 123K 361s 12.8 96% 272K 697s 21.5 98% 134K 984s 13.4 97% 238K 1495s 19.8 95% 318K 1971s 25.0 97% 282K 1585s 22.4 98% 244K 1758s 19.4 98% 262K 3078s 21.1 Time Limit Exceeded
Table 15: Three-agent comparison on intra-datacenter fat-tree with DeepSeek-V4-Pro across all scales. k
Routers Acc
4 8 12 16 20 30 50 80 120 160 200
20 80 180 320 500 1,125 3,125 8,000 18,000 32,000 50,000
In-Context Learning Tok tRun RLLM
96% 24K 80s 1.0 87% 108K 343s 1.0 90% 309K 957s 1.0 82% 674K 2071s 1.0 Context Window Overflow
Acc
Iterative Reasoning Tok tRun RLLM
96% 34K 125s 3.3 97% 67K 245s 3.3 99% 141K 472s 3.2 97% 260K 831s 3.0 99% 497K 1701s 3.1 Time Limit Exceeded
Acc
Agent Skills Tok tRun RLLM
96% 144K 520s 13.5 98% 160K 662s 15.0 99% 161K 685s 14.8 100% 149K 641s 14.4 100% 171K 1445s 15.9 100% 162K 1355s 15.1 99% 184K 1512s 16.0 100% 195K 1458s 16.8 99% 214K 1912s 17.6 99% 237K 3124s 19.0 Time Limit Exceeded
Table 16: Three-agent comparison on intra-datacenter fat-tree with Qwen-3.5-397B across all scales. k
Routers Acc
4 8 12 16 20 30 50 80 120 160 200
20 80 180 320 500 1,125 3,125 8,000 18,000 32,000 50,000
In-Context Learning Tok tRun RLLM
22% 23K 8% 113K
41s 182s
1.0 1.0
Acc
Iterative Reasoning Tok tRun RLLM
98% 29K 91% 54K 87% 89K 59% 174K 31% 254K
64s 139s 192s 316s 860s
3.5 5.0 4.3 4.1 8.2
Acc
Agent Skills Tok tRun RLLM
100% 208K 424s 18.8 98% 118K 341s 14.0 99% 255K 646s 19.6 99% 230K 580s 18.1 99% 243K 1422s 18.3 100% 215K 1354s 18.1 99% 287K 1689s 21.1 99% 275K 1487s 20.9 99% 214K 1722s 19.6 Time Limit Exceeded
Table 17: Three-agent comparison on intra-datacenter fat-tree with Llama-3.3-70B across all scales. k
Routers
In-Context Learning Acc Tok tRun RLLM 58% 20K 8% 98K
43s 191s
1.0 1.0
Acc
Iterative Reasoning Tok tRun RLLM
50% 19K 67% 50K 17% 78K 25% 184K
50s 125s 170s 370s
2.3 2.9 2.1 2.5
Acc
Agent Skills Tok tRun RLLM
4 8 12 16 20 30 50 80 120 160 200
20 80 180 320 500 1,125 3,125 8,000 18,000 32,000 50,000
67% 623K 1340s 34.8 92% 526K 1273s 27.2 58% 237K 686s 18.7 83% 97K 339s 12.6 83% 538K 2627s 28.3 83% 139K 1071s 14.4 92% 373K 1999s 22.2 58% 160K 1085s 15.4 67% 376K 2320s 22.7 50% 882K 5983s 34.1 Time Limit Exceeded
D.3
Three-Agent Comparison on Inter-Datacenter and Compound-Change Analysis
We further present the complete scaling results of the three agentic methods with GPT-5.4 and GPT-4o under complex cases (inter-datacenter and compound changes) across all k, complementing the “boundary” discussion in §5.2. Overall, Agent Skills still has the smallest token-cost growth and reaches the largest scale among the three methods, while the other two again hit context-window or time limits at small scales. Under compound changes, the accuracy of In-Context Learning and Iterative Reasoning degrades more rapidly than that of Agent Skills, indicating that reading raw 26
configurations is even less capable of analyzing the complex case where two protocol attributes change on the same flow simultaneously—this further reveals the design advantage of Agent Skills.
Table 18: Three-agent comparison on inter-datacenter networks with GPT-5.4 across all scales. k
Routers
4 8 12 16 20 30 50 80 120 160 200
In-Context Learning Acc Tok tRun RLLM
48 100% 49K 158s 1.0 176 96% 261K 810s 1.0 384 90% 774K 2382s 1.0 672 Context Window Overflow 1,040 2,310 6,350 16,160 36,240 64,320 100,400
Iterative Reasoning Tok tRun RLLM
Acc
94% 44K 157s 3.1 91% 135K 451s 3.1 91% 333K 1059s 2.8 92% 779K 2420s 3.0 Time Limit Exceeded
Acc
Agent Skills Tok tRun RLLM
96% 96% 91% 90% 88% 86% 84% 77% 80% 77% 74%
91K 329s 97K 394s 96K 409s 100K 418s 98K 843s 103K 850s 99K 921s 98K 788s 108K 1137s 114K 1718s 127K 3012s
7.6 7.8 7.8 8.0 7.9 7.9 7.9 7.9 8.3 8.7 9.2
Table 19: Three-agent comparison on inter-datacenter networks with GPT-4o across all scales. k 4 8 12 16 20 30 50 80 120 160 200
Routers
In-Context Learning Acc Tok tRun RLLM
48 100% 49K 57s 1.0 176 Context Window Overflow 384 672 1,040 2,310 6,350 16,160 36,240 64,320 100,400
Acc
Iterative Reasoning Tok tRun RLLM
95% 80K 118s 5.2 90% 234K 301s 5.0 Context Window Overflow
Acc
Agent Skills Tok tRun RLLM
88% 148K 239s 12.8 74% 146K 313s 13.1 72% 148K 348s 13.2 59% 210K 447s 16.5 52% 220K 1381s 16.8 39% 304K 1686s 20.4 43% 443K 2424s 25.2 38% 463K 2103s 26.2 37% 535K 3373s 29.2 Time Limit Exceeded
Table 20: Three-agent comparison on compound changes with GPT-5.4 across all scales. k 4 8 12 16 20 30 50 80 120 160 200
Routers 20 80 180 320 500 1,125 3,125 8,000 18,000 32,000 50,000
In-Context Learning Acc Tok tRun RLLM 100% 20K 66s 1.0 80% 100K 316s 1.0 70% 293K 904s 1.0 65% 645K 1978s 1.0 Context Window Overflow
Acc
Iterative Reasoning Tok tRun RLLM
100% 30K 107s 3.2 97% 57K 201s 3.1 87% 123K 410s 3.0 88% 249K 792s 3.0 68% 440K 1510s 2.9 Time Limit Exceeded
27
Agent Skills Acc Tok tRun RLLM 100% 99% 97% 96% 95% 82% 80% 78% 76% 78% 79%
48K 178s 55K 227s 55K 240s 58K 245s 57K 532s 57K 469s 61K 571s 57K 513s 57K 750s 58K 1179s 58K 2100s
6.4 6.6 6.3 6.4 6.3 6.3 6.5 6.3 6.3 6.3 6.3
Table 21: Three-agent comparison on compound changes with GPT-4o across all scales. k
Routers Acc
In-Context Learning Tok tRun RLLM
4 8 12 16 20 30 50 80 120 160 200
20 80 180 320 500 1,125 3,125 8,000 18,000 32,000 50,000
78% 20K 25s 1.0 33% 100K 111s 1.0 Context Window Overflow
E
Discussion
Acc
Iterative Reasoning Tok tRun RLLM
84% 29K 45s 3.3 30% 59K 90s 3.3 19% 131K 171s 3.3 17% 257K 295s 3.2 Context Window Overflow
Acc
Agent Skills Tok tRun RLLM
81% 98K 160s 12.5 71% 108K 229s 13.2 42% 112K 268s 13.3 44% 114K 258s 13.3 52% 119K 887s 13.5 38% 173K 899s 15.4 38% 206K 1225s 17.2 40% 220K 1169s 17.5 36% 292K 2054s 19.1 41% 276K 3272s 18.9 Time Limit Exceeded
Topology coverage. Production datacenter networks vary in subtle layer counts and structural details across vendors—Azure [24], Google’s Jupiter [44], and Meta’s fabric [2] all differ in how many tiers they expose and how pods are aggregated (e.g., Meta’s fabric has three tiers and within pods 48 racks are aggregated by four fabric switches [2], while Azure deploys three tiers (T0/T1/T2) with T0 aggregated under T1 to form pods and pods interconnected by T2 spines [24]), but they are all Clos networks at the core. To reflect a generic production datacenter, we adopt fat-trees, which are widely used as the reference Clos topology in academia [1, 25] and as the building block of production datacenter fabrics [2, 24, 44]. Supporting other Clos variants (leaf-spine [24], dragonfly [27]) requires only new topology templates in our generation pipeline. Configuration coverage. ROUTING B ENCH models routing-related configurations only (§3) to construct a pure benchmark for routing-path analysis. Real networks also contain non-routing configurations, e.g., security policies, traffic bandwidth-allocation rules, device login-authentication settings, and device/traffic status-monitoring configurations [11]. §D.1 shows that this heterogeneity may inflate token cost on the production network. We hope future agentic-NetOps benchmarks will incorporate and analyze these protocols beyond path analysis. Cost measurement under remote LLM APIs. LLM API runtimes can be perturbed by remoteservice load and network conditions, making runtime a noisy cost signal. We therefore use token consumption as the primary cost metric for each agent, which remains stable regardless of network conditions and workloads. Additionally, we pin every LLM to a specific OpenAI API version or open-source model release (§5) to minimize result drift across versions of the same model.
28