Conceptio › Archive › arXiv CS
arXiv CSopen access

Benchmarking LLM-Driven Network Configuration Repair

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

Benchmarking LLM-Driven Network Configuration Repair Ioannis Protogeros, Rufat Asadli, Benjamin Hoffman, Laurent Vanbever ETH Zürich Zürich, Switzerland {iprotogeros,rasadli,bhoffman,lvanbever}@ethz.ch with their impressive capabilities across domains [11, 12], LLMs appear promising for facilitating complex workflows that are bottlenecked by human reasoning. On the other hand, as probabilistic models, they remain prone to errors and hallucinations [13, 14], precluding their adoption for managing critical infrastructure. The tension between the tremendous potential for automating network operations and their associated risks makes principled evaluation indispensable. Doing so requires designing a benchmark that challenges LLM reasoning with complex and diverse configuration tasks, alongside with a proper methodology to assess their capabilities on said tasks. Constructing such a diverse benchmark with plausible, well-posed problems is nontrivial. Unlike domains such as mathematics or software engineering, where benchmarks can leverage vast publicly available datasets, network configurations are proprietary and sensitive. Consequently, we must synthesize the dataset. Crucially, this synthesis cannot be random; to ensure relevance, test cases must reflect the complexities that operators face in production networks, including large-scale configurations with feature and protocol dependencies. Lastly, the evaluation must verify the functional correctness of solutions, scalably and automatically.

arXiv:2604.22513v1 [cs.NI] 24 Apr 2026

ABSTRACT There is a rapidly growing interest in using Large Language Models (LLMs) to automate complex network operations, but their reliable adoption requires rigorous assessment of their effectiveness and safety. Existing benchmarks do not address whether LLMs can successfully resolve errors in large-scale, interdependent network configurations without introducing new disruptions. Developing such a benchmark is challenging: scenarios must be diverse and increasingly complex, yet their evaluation must be straightforward and meaningful. In this paper, we present Cornetto, the first benchmark to evaluate LLM-driven network configuration repair functionally and at scale. Cornetto features a generation pipeline that synthesizes representative and plausible misconfiguration scenarios, coupled with an evaluation framework that uses formal verification to assess functional correctness of proposed fixes against ground-truth specifications. Using this pipeline, we synthesize a dataset of 231 problems for fixing configurations across varying network topologies (20–754 nodes) and diverse protocols. We evaluate 9 state-of-the-art LLMs and find that while they show promise, they often introduce regressions and their performance degrades at scale. Our results indicate that reliable LLM-powered network automation requires integrating LLMs into iterative workflows guided by formal verification.

1

The gap in LLM network configuration benchmarking. LLMs’ progress outpaces our abilities to evaluate them effectively in fields that require deep domain expertise and complex reasoning. Despite growing interest in using LLMs for network operations, we still lack benchmarks to evaluate their performance and explore strategies to improve them. While previous works such as NetConfEval [15] and NetLLMBench [16] have established baselines for evaluating LLMs on network configuration tasks, they are severely constrained in scale and complexity relative to the capabilities of current models. Additionally, their proposed evaluation methods rely on proxy metrics (e.g., textual similarity or pingtest validation) that do not guarantee the functional correctness of a configuration. A more recent benchmark, NIKA [17], evaluates the diagnostic capabilities of LLM Agents in dynamic, emulated network environments. Yet, it does not support the evaluation of proposed fixes for network faults.

INTRODUCTION

Network correctness, while paramount, remains extremely difficult to achieve and maintain. While network verification and synthesis [1–5] have made significant strides towards eliminating human-induced misconfigurations, they are not a silver bullet. In particular, their adoption is hindered by limited protocol coverage and inaccuracies in modelling complex network behaviour [6, 7]. More recently, there has been a surge of interest in leveraging Large Language Models (LLMs) as a more flexible approach to automating network operations. Hyperscalers have already begun deploying LLM-based frameworks that assist with such tasks, including ByteDance’s NetAssistant [8], Alibaba’s BiAn [9], and Meta’s Confucius [10]. On the one hand, 1

Preprint, ,

Protogeros et al.

Consequently, evaluating the performance of LLMs in repairing realistic, large-scale network configurations correctly and safely remains underexplored.

by optimizing fault diversity within a minimal number of scenarios. • We design a verification-based evaluation pipeline that utilizes data-plane analysis to automatically infer ground-truth specifications and assess the functional correctness and safety of fixes. • We evaluate 9 state-of-the-art LLMs on Cornetto and comprehensively analyze their performance, examining the effects of scenario complexity on reconfiguration correctness and safety.

Cornetto: Correct & safe configuration repair. To address this gap, we introduce Cornetto, a benchmark that evaluates end-to-end configuration repair in representative, large-scale networks.1 The task of resolving misconfigurations encapsulates critical challenges in network operations: understanding the interplay among interdependent features and protocols, and bridging the semantic gap between lowlevel configurations and high-level intent [5]. To capture this complexity, Cornetto employs a generation pipeline that produces syntactically valid configurations subject to logical and semantic constraints, thereby ensuring structural coherence and consistency. We use this pipeline to synthesize 231 challenging misconfiguration scenarios spanning diverse protocols and scales, where the intended network state is unambiguously defined. To evaluate correctness and safety, Cornetto formally verifies the data plane of the reconfigured network against ground-truth specifications. This ensures that success depends on functional correctness, requiring restoration of intended behaviour. The introduction of new bugs is penalised based on the extent of disruption to previously functional behaviour. Crucially, we also evaluate the diagnostic reasoning (localization and root-cause analysis) that led to the fix.

Outlook. We will open source Cornetto to the community as an extensible and modular framework. 2 Cornetto enables the generation of more challenging tasks, and its architecture supports the evaluation of any LLM-based system, including advanced scaffolds such as Retrieval-Augmented Generation (RAG) systems and agentic setups. By providing a standardized testbed, Cornetto contributes to the continual evaluation of LLMs on configuration repair.

2

OVERVIEW

Cornetto is a benchmark for assessing LLMs’ capabilities in automated network configuration repair. A Cornetto scenario simulates an end-to-end troubleshooting task: Given a misconfigured network state (including topology and configuration files) and a set of high-level intents (e.g., Reachability between A and B), the model must localize and diagnose the misconfiguration, and propose a correct reconfiguration that restores the network’s intended function. Cornetto comprises two pipelines. First, the Dataset Generation focuses on creating realistic, diverse, and complex misconfiguration scenarios. Second, the Evaluation Framework provides an automated system that enables meaningful evaluation of proposed solutions against intended network behaviour. This section outlines Cornetto design goals, the key insights required to address emerging challenges, and the system components that implement these solutions.

Key findings. Our evaluation of 9 LLMs reveals that while current models can diagnose and fix faults (restoring up to 60% of network state on average), they cannot reliably act as monolithic solvers—the best-performing model successfully resolved only 25% of scenarios. We find that performance degrades in large-scale settings with noisy data, and models often propose partial or unsafe solutions. These insights indicate that reliable automation requires integrating LLMs into systems that filter noisy data, preserve necessary context, and iteratively verify the safety of repairs before application. Key contributions. Our summarized contributions are: • We formulate the problem of automated configuration repair, enabling the quantification of functional disruption caused by misconfigurations and the assessment of fix correctness and safety. • We develop a scenario generation pipeline that synthesizes logically valid network configurations for any topology and systematically injects faults across diverse protocols. • Using this pipeline, we curate a dataset of 231 misconfiguration scenarios of varying scale and complexity

2.1

Design Goals

Diversity and Complexity. To gain meaningful insights into the performance of LLMs, the benchmark must cover a wide variety of scenarios. This includes growing topology scales, features across diverse protocols, and complex fault scenarios that reflect real-world issues. At the same time, the benchmark should remain small enough to keep its execution feasible without incurring high costs. Additionally, the complexity of these scenarios should be sufficient to challenge models’ reasoning capabilities and ensure resilience against benchmark saturation [19]. A key

1 The idea behind the system was presented in a previously-accepted 2 available at https://github.com/nsg-ethz/cornetto

poster [18]. 2

Benchmarking LLM-driven configuration repair

Preprint, ,

I. Dataset Generation & Problem Formulation (§3) Topology collection

231 Scenarios

Scenario Coordination (§3.2)

Network specification and violated specifications

Configuration Generation + Fault Injection (§3.3)

Feature dependency modeling

Fault Library (§3.2)

eBGP

class DuplicateLoopbacks(Fault): def applicable(): class DuplicateLoopbacks(Fault): def applicable(): ... class DuplicateLoopbacks(Fault): ... def targets(): def applicable(): ... def targets(): ... ... def apply(): def targets(): def apply(): ... ... ... def apply():

Route Refl.

Reachability(n1,p2) Waypoint(n3,p1,n1)

-RemoveRouteReflector -RemoveRouteReflector

iBGP

RemoveRouteReflector

IGP

-DisableOSPF2BGPRedist -DisableOSPF2BGPRedist

DisableBGP2OSPFRedist

Redistr.

...

Isolation(n2,p2)

Scenario #1 Scenario Scenario#1 #1

OSPF

System Under Test

Reachability(n1,p2) Isolation(n2,p2) Waypoint(n3,p1,n1)

✘ ✔ ✘

Differential Data Plane Analysis (§4.2) p1 p1 p1

Problem description

✘ ✔ ✘

p1 p1 p1

Proposed Reconfiguration

Metrics & Report ✘ ✘

✔ ✔ ✔✔

✘ ✔ Waypoint(n3,p1,n1)✘ Reachability(n1,p2)

Diagnosis: 0.72 Fix score: 0.43

✘ ✘

Isolation(n2,p2)

Regression: 0.12

II. Evaluation Framework (§4)

Figure 1: Cornetto architecture. (I) The Dataset Generation pipeline coordinates the scenarios to ensure a diverse and complex test suite, generates sensible configurations and misconfigurations, and provides a standardized problem definition. (II) The Evaluation Framework enables automated and meaningful evaluation of the created scenarios by validating the reconfigured network’s behaviour against ground-truth specifications. challenge lies in identifying the dimensions that affect a scenario’s complexity.

and misconfigurations is prohibitively large. To create a representative benchmark that can be run on a minimal compute budget, we identify the key dimensions that affect scenario difficulty (e.g., topology size, fault types) and employ sampling strategies and combinatorial testing techniques [20] to efficiently cover them. This enables Cornetto to test LLMs against a diverse suite of scenarios without requiring thousands of redundant test cases.

Sensible problems. Scenarios must be sensible so that a model’s performance reliably proxies its real-world applicability. However, actual network-wide configurations and historical bug data are proprietary and hence unavailable. Therefore, the system must generate scenarios synthetically. This includes base configurations that are syntactically and semantically valid, maintain internal coherence, and respect dependencies between protocols.

Producing sensible configurations with grammar-based generation and semantic constraints (Providing sensible problems). For a benchmark to effectively evaluate LLM performance on resolving misconfigurations, it must contain sensible configurations that can exist in reality. This means that configurations must be syntactically and semantically valid and realize a plausible intent. To achieve that, we first define a high-level logical network plan that specifies its functionality. Then, we synthesize configs using a grammarbased approach and enable features iteratively and contextually. This approach ensures that semantic constraints (e.g., that a prefix list is defined before it is referenced) and feature dependencies are respected.

Well-posed problems. Despite the scenarios’ complexity, the evaluation of their proposed solutions must be straightforward and provide concrete, interpretable metrics for success, efficiency and safety. This requires the concrete formulation of a misconfiguration scenario, including the definition and quantification of success metrics. Meaningful evaluation. Evaluation should take into account the emergent network behaviour, rather than relying solely on the comparison of the configuration files. This requires analyzing the network’s data plane behaviour post-fix, and comparing it against a specification that captures the intended behaviour. Furthermore, all the different facets of troubleshooting should be assessed, including localization and diagnosis abilities.

2.2

Formulating concrete problems through differential data plane analysis (Providing well-posed problems). We cannot proxy the effect of a misconfiguration by a textual diff; a perturbation of just a few Lines of Code (LoC) may cause significant disruptions in the network, while a larger change can be functionally benign. To rigorously define the problem, Cornetto employs differential data-plane analysis. We compare the forwarding behaviour of the faulty network

Key Insights

Efficiently representing a massive task space (Enabling diversity). The space of all possible network configurations 3

Preprint, ,

Protogeros et al.

against the “golden" reference state using Batfish [1]. The resulting set of behavioural differences (e.g., “A cannot reach B") yields a concrete, symptom-based problem description.

tasks the model under test with a standard troubleshooting workflow: localizing the fault, diagnosing the root cause, and proposing a reconfiguration to restore intended behaviour. The evaluation pipeline parses the proposed solution, simulates the new data plane, and compares the network state against the desired specification set. This process yields metrics for reconfiguration Efficacy (did the proposed reconfiguration fix the violations?) and Safety (did it introduce any new violations?), along with performance metrics for diagnosis and localization.

Evaluating functional correctness and reasoning (Enabling meaningful evaluation). To automate evaluation, we treat the network’s functional requirements as a suite of “unit tests". We mine the specific data-plane properties (Reachability, Isolation, Waypointing) satisfied by the golden network [21] to create a ground-truth specification. A correct solution will produce a reconfigured network that restores the previously violated predicates (i.e., it is efficacious), without violating any previously satisfied predicates (i.e., it is safe). Crucially, we also evaluate the intermediate steps of a model’s reasoning (localization, root-cause diagnosis) to gain a holistic view into its troubleshooting effectiveness.

2.3

Output and results. For each test case, Cornetto generates a structured report primarily containing: • Fix Rate (Efficacy): The proportion of initially violated specifications that are successfully restored by the configuration • Regression Rate (Safety): Quantifies unintended side effects by measuring new violations introduced by the reconfiguration. Additionally, the testbed evaluates diagnostic quality using both objective metrics (precision/recall on faulty devices for localization) and an LLM-as-a-Judge [22] approach to assess provided textual diagnoses against ground-truth misconfigurations.

Cornetto

Cornetto comprises two primary pipelines, as shown in Fig. 1. The Dataset Generation (§3) pipeline accepts a topology collection and a fault library to generate a minimal yet diverse suite of misconfiguration scenarios, along with their problem formulations. With the generated scenarios, the Evaluation Framework (§4) assesses repair capabilities by verifying configurations against a ground-truth specification.

3

Scenario coordination and creation (§3.2). To drive fault generation, we implement a diverse fault library, where each fault is a function that slightly perturbs a configuration to break or alter the functionality of a protocol. The scenario coordinator utilizes this library and the topology collection to orchestrate the creation of diverse scenarios, ensuring representation across all topology scales and fault combinations. For each scenario, the system generates a valid configuration that enables the specific protocols and features targeted by the fault. This process yields two network states per scenario: the Golden (healthy) state and the Broken (faulty) states.

DATASET GENERATION PIPELINE

The test suite of our benchmark must cover a wide variety of high-quality scenarios of varying complexity. Yet, it should include a minimal number of test cases so the research community can test their methods without incurring prohibitive LLM inference costs. In this section, we define the network configuration troubleshooting task space and show how to represent it effectively, constructing a benchmark that meets our design goals.

3.1

Task Definition

Formally, we define a Cornetto benchmark scenario as a tuple (𝑇 , 𝐶 gold, 𝐶 broken, Φ), where: • 𝑇 represents the network topology (undirected graph of devices and links between interfaces). • 𝐶 gold is the configuration at its “Golden" state. • 𝐶 broken is the faulty configuration that is derived from 𝐶 gold by applying a fault function 𝑓 , so 𝐶 broken = 𝑓 (𝐶 gold ). • Φ is the set of data plane specifications (or intent) that the Golden Configuration satisfies. We denote satisfaction as 𝐶 gold |= Φ. Since the forwarding plane of 𝐶 broken deviates from intended behaviour, it holds that 𝐶 broken ̸ |= Φ.

Data plane analysis and problem formulation (§3.3). The system invokes Batfish to simulate the forwarding behaviour for both the Golden and Broken states. It then distills the behaviour of the Golden network into a set of invariants, or predicates, which constitute the ground-truth specification. By verifying the Broken network state against this specification, the system identifies which specifications are violated. These violations are the “symptoms" that indicate the problem in network behaviour. Benchmark testbed (§4). The testbed manages the interaction with the LLM-based system under test. The benchmarked system receives a description of the network problem created by the Dataset Generation pipeline. This description contains the network topology, the faulty configurations, and the violated specifications (symptoms). Cornetto

Specifications. We define a specification 𝜙 ∈ Φ as a boolean predicate that describes a property in the network’s forwarding behaviour. We consider four types of predicates that 4

Benchmarking LLM-driven configuration repair

Preprint, ,

encompass the most common requirements about the network’s function [21]:

multiple independent root causes, potentially causing masking or compounding symptoms. • Fault types: Different faults from the library F will cause different types of symptoms that will be more or less difficult to link to the root cause.

• Reachability(r,p): Traffic from router r can reach prefix p. • Isolation(r,p): Traffic from r cannot reach p. • Waypointing(r,p,w): Traffic from r destined to p always passes through router w. • LoadBalancing(r,p,n): Traffic from r destined to p is load-balanced across n paths.

These three dimensions will be used to select the scenarios that will comprise the benchmark. The fault library. To ensure the benchmark contains troubleshooting tasks across a wide array of misconfigurations, we curated a collection F of 27 fault functions that target specific protocol functionalities. This fault library spans across the following dimensions that we expect to affect scenario complexity:

Specification violations. The fault 𝑓 introduces a disruption in the data plane. We define the set V of violated specifications as the subset of specifications that are satisfied by 𝐶 gold but violated by 𝐶 broken :

• Protocols and features affected: The library includes faults that target features of eBGP, iBGP (incl. route reflection), OSPF and IS-IS (single and multi area), redistribution, ACLs, route-maps, and static routes • Configuration impact: From perturbing a single parameter (e.g., the subnet mask of an interface) to performing “organized" alterations like removing a route reflector functionality from a router. • Operational impact: Faults that disrupt a protocol’s functionality (e.g., mismatched remote-as numbers preventing BGP session establishment), and faults that only change the intent of the used feature (e.g., stripping an export policy)

V = {𝜙 ∈ Φ | 𝐶 broken ̸ |= 𝜙 } The objective. The system under test’s task is twofold. Given the input tuple I = (𝑇 , 𝐶 broken, V), it must: • Localize and Diagnose: Provide (i) a list of the misconfigured routers and (ii) a textual description of the faults in the network that correspond to the misconfiguration Δ(𝐶 gold, 𝐶 broken ). • Repair: Act as a repair function R that produces a reconfiguration 𝐶 fix = R (I) such that 𝐶 fix |= Φ.

3.2

Effective Task Space Representation

For a given topology set T , the theoretical space of benchmark scenarios is defined by all the configurations that could take the place of 𝐶 gold and 𝐶 broken . If we consider both configs to be a part of a configuration space C that fits each topology, then the scenario space would be contained in T × C × C. This space is prohibitively large and dominated by unreasonable elements, i.e., random configuration pairs that cannot represent either operational networks or realistic misconfiguration scenarios. To construct a useful benchmark, we must restrict the space to a meaningful subset of scenarios and strategically represent it with minimal samples.

Appendix A contains the comprehensive list of all faults. Representative Sampling Strategy. To balance scenario diversity with a manageable test set size, we employ a sampling strategy centred on pairwise coverage [20]. Our goal is to generate complex scenarios with up to 𝑁 = 8 simultaneous faults, where every possible pair of fault types appears together at least once. This ensures we test model performance on scenarios with varying disruptions and multi-root-cause failures. 1. Fault selection procedure. We construct a compact collection of fault sets S by iteratively sampling from the fault library F until 100% pairwise coverage is achieved. For each generation step:

Benchmark diversity. Covering the entire task space is neither possible nor relevant to our goals. However, to ensure robust evaluation, the benchmark should challenge the tested systems across different complexity dimensions. We identify the following controllable characteristics that are expected to critically affect the difficulty and nature of a misconfiguration scenario:

(1) We randomly select a number 𝑘 ∈ {2, . . . , 8} of simultaneous faults to apply. This variation ensures the benchmark includes scenarios from few to many root causes. (2) Select a subset 𝐹𝑠 ⊂ F of 𝑘 faults |𝐹𝑠 | = 𝑘 that greedily maximizes the number of newly covered fault pairs (3) Repeat this process until the set of applied scenarios covers 100% of feasible fault pairs.

• Topology scale: Varying the size of the network up to hundreds of nodes will stress-test the models’ ability to handle large inputs and identify the information that points to the issue. • Number of applied faults: Applying multiple simultaneous faults will challenge models with having to detect

This optimization yields a compact collection of 50 distinct fault sets S = {𝐹 1, 𝐹 2, . . . 𝐹 50 } that describe which faults are to be applied in each scenario. To this collection, we add all 5

Preprint, ,

Protogeros et al.

monosets of single faults (𝑘 = 1) to include each fault type individually.

To ensure that generated configurations are operationally viable, we model the dependencies between network protocols. Network features are rarely independent; for example, testing a route-reflection fault requires the AS to support iBGP, which, in turn, relies on an underlying IGP (e.g., OSPF) for loopback reachability. We enforce these constraints during the generation phase: for each selected fault, the generator enables the target features and recursively satisfies all its prerequisite protocol dependencies.

2. Topology selection. For the topology collection T , we use real-world topologies from the Topology Zoo [23]. To ensure these fault patterns are tested across varying scales, we stratify our topology dataset T into three tiers: Small (<50 nodes), Medium (50–100 nodes), and Large (>100 nodes). For each generated fault set 𝐹𝑖 ∈ S, we instantiate the benchmark scenario by applying the faults to three distinct topologies, one randomly sampled from each tier. This yields a final dataset of 231 scenarios that cover diverse combinations of fault types across all scale tiers.

3.3

Configuration generation. The syntactic and semantic validity of the “Golden" state configurations is critical for the realism and utility of the benchmark. While syntactic validity ensures the configuration can be parsed, semantic validity ensures the configuration is logically consistent. We enforce two types of semantic constraints, as stipulated in Metha:

Scenario Generation

The scenario selection procedure described before yields for each scenario a descriptor (𝐹,𝑇 ) with (i) the set of faults 𝐹 ⊂ F and (ii) the topology of the network 𝑇 ∈ T . We now describe the pipeline that synthesizes the configurations themselves and applies the fault functions to obtain the configurations 𝐶 gold and 𝐶 broken . To ensure the benchmark scenarios are plausible and useful, the generated configurations must satisfy two constraints:

• Intra-device constraints: Dependencies within a single configuration file. For example, a BGP neighbour statement cannot apply a route-map that has not been defined, and an interface cannot be assigned to an OSPF area if the OSPF process is not active. • Inter-device constraints: Dependencies across the network. For example, two routers connected via a link must have IP addresses in the same subnet, and eBGP peers must have matching remote-as declarations.

• The base configuration 𝐶 gold must be syntactically and semantically valid. Crucially, the features that are enabled in the network must not exist vacuously (e.g. BGP processes without peers or unreferenced routemaps), but they must be structured to realize a functional intent within the network context. • The “broken" config 𝐶 broken must be derived from 𝐶 gold after a perturbation that is minimal, so it can represent a plausible misconfiguration.

To satisfy these, we construct a high-level logical plan of the network. This process is iterative and context-aware: First, we extend the physical topology with logical groupings to define the control plane hierarchy. We split the topology into ASes and assign OSPF/IS-IS areas to router interfaces according to the chosen IGP in each domain. We also define peering relationships, including iBGP full-meshes and route-reflection clusters. Then, we assign subnets to links and IP addresses to the interfaces on those links. Finally, additional resources are generated based on the selected features and their dependencies, as defined by the logical topology. For instance, BGP advertisements are generated strictly for subnets assigned to the router’s local interfaces. Once the logical plan is defined, we render it using a template that follows a Context-Free Grammar for vendorspecific configurations (e.g., Cisco IOS). This part ensures the syntactical correctness of the produced configurations, and is decoupled from the process of enforcing semantic constraints (which are not context-free).

Existing synthesizers like NetComplete [4] or Propane [5, 24] are ill-suited for this task because they optimize for a fundamentally different objective: finding any valid configuration that satisfies a specific high-level intent, often resulting in simple, uniform implementations. Our benchmark requires configurations that vary in protocols and specific, low-level configuration features. High-level intent is insufficient for this purpose, as a single intent (e.g., reachability) can be satisfied by many combinations of features. Therefore, we adopt a grammar-based generation approach building on Metha [6], allowing for the selection of features that the fault functions in F can affect.

Fault injection. We apply the fault functions to the produced logical plan, so when it passes through the renderer, it results in the broken configuration 𝐶 broken . After these faults, the configurations may either violate semantic constraints (e.g., mismatched remote-as parameters) or remain semantically valid, only deviating from intended behaviour (e.g., changing a route-map action from permit to deny). In

Configuration feature selection. Each fault function acts upon specific configuration features. But for the fault to be applicable, the appropriate "attack surface" must exist in the first place. For instance, removing a route reflector requires that the AS is configured with route reflection clusters. 6

Benchmarking LLM-driven configuration repair

Preprint, ,

3.4

Violated specifications (% of total)

both cases, the specifications violated after fault injection are verified through data-plane analysis, ensuring that all faults induce a tangible change in forwarding behaviour.

Dataset Statistics

The resulting Cornetto dataset comprises 231 network misconfiguration scenarios across topologies with 20 to 754 nodes. Table 1 reports the key features of the scenarios. Notably, the network-wide configurations are large (>16K lines of code on average), yet contain only a few lines actually affected by the fault. Consequently, solving Cornetto requires navigating massive, distributed configuration files to locate the few relevant lines that point to an issue. At the same time, the actual change in the configuration is small and buried under the volume of code. The effect of a misconfiguration is not related to its textual difference either; most of the faults affect <1% of the total configuration lines, but may greatly affect the functionality of the network as shown in Fig. 2. We expect the varying levels of perturbations and symptoms to affect scenario difficulty.

Topology

Fault impact

Impact (%)

4

Mean

Nodes (#) Configuration lines (LoC) Routes Data plane predicates

86.4 754 16.1K 200.0K 4.6K 130.4K 12.7K 598.0K

Lines edited Routers affected Routes changed Predicates changed

50.5 5.5 568.4 1.5K

345 20 9.2K 38.4K

LoC% changed Routes% changed Predicates% changed

0.44 6.34 6.82

5.93 57.3 49.8

Large (>100 nodes)

50% 10% 1% 0.1%

0.01%

0.1%

1%

Lines of code changed (% of total)

decouples evaluation logic from the specific solver. Hence, Cornetto can evaluate any proposed system that can generate a configuration 𝐶 fix . To carry out experiments across different models, we implement an evaluation framework that (i) builds a structured prompt containing the problem description to elicit a solution, and (ii) parses a model-generated patch to build the configuration 𝐶 fix . Context Construction. To provide the necessary info to diagnose and fix the problem in the configuration, the system constructs a prompt containing three main information sources:

Max

• The physical topology 𝑇 • The list of violated specifications V • The faulty configuration files 𝐶 broken A key challenge for the benchmarked system is handling the volume and sparsity of the available raw data; among thousands of configuration lines, only a very small subset points to the root cause in the network. At the same time, it is possible that network-wide configurations, along with the specifications and topology data, cannot fit within the LLMs’ context windows, and their performance is known to degrade well before that limit [25, 26]. Consequently, handling the context constitutes a core experimental dimension (§5). We define the system under test to include not only the generation model but also the context strategy, that is used to derive, filter, or retrieve relevant information from the available raw data (𝑇 , 𝐶 broken , Φ, V). For obtaining the solution, we prompt the models to output the following:

EVALUATION FRAMEWORK

To derive meaningful insights into the diagnostic capabilities of LLMs, the evaluation system must distil concrete, interpretable metrics that quantify the functional success of a solution. In this section, we delineate the Cornetto evaluation pipeline that enables automatic evaluation of proposed reconfigurations against the ground truth of the network’s data-plane behaviour.

4.1

Medium (51-100 nodes)

Figure 2: While configuration perturbations are minimal, disruption in network behaviour varies greatly across scenarios.

Table 1: Cornetto spans diverse scales and misconfiguration severities. Metric

Small (<50 nodes)

(1) A textual diagnosis, containing all the detected faults in the network configuration 𝐶 broken (2) A list of all the routers that need to be reconfigured, and the needed reconfigurations to resolve the specification violations

LLM-Benchmark Interface

To enable testing and comparing different models and systems around them, we build a standardized interface that 7

Preprint, ,

Protogeros et al.

Reconfiguration Parser. The raw text output of the model needs to be processed in order to apply the fixes and obtain the proposed reconfiguration 𝐶 fix . We decided against requiring Unix diff patches [27], since they require precise line arithmetic, which language models famously struggle with [28, 29]. Instead, following common practice in coding agents [30] that perform edit operations, we expect the answer to include, for each reconfigured file:

Node R1 R1 R2 R3 R4

Action Fwd R2 Fwd R3 Fwd R4 Drop Accept

R2 R1

R4 R3

Reachability(R1,10.1/24) Reachability(R2,10.1/24) Isolation(R3,10.1/24) Reachability(R4,10.1/24) Waypoint(R1,10.1/24,R2)

Figure 3: For each unique destination prefix, the pipeline uses the forwarding behaviour table (calculated by Batfish) to construct a forwarding graph, from which it derives the specifications of the network.

• A search block, containing a snippet that is uniquely present in the configuration file • A replace block, containing the snippet that will replace the search block

• The set of regressions as originally healthy specifications that are violated by the fix Φregressed = {𝜙 ∈ (Φ \ V) | 𝐶 fix ̸ |= 𝜙 }

Since unparseable output is still possible (wrong format, non-existent search block), we draw inspiration from the same practices and robustify the pipeline by (i) allowing fuzzy matching of search blocks, if there is a block that differs only in whitespaces or has a small enough Levenshtein distance [31], and (ii) by providing feedback from the parser in case of invalid outputs.

4.2

Dest. 10.1/24 10.1/24 10.1/24 10.1/24 10.1/24

• The set of violations that remained unresolved: Φunfixed = V \ Φfixed And we calculate the following scores that describe different performance aspects of the solutions: • Safety (Regression Rate): The proportion of specifications that were violated because of the proposed misconfiguration: |Φregressed | Regression = |Φfixed | + |Φunfixed | + |Φregressed |

Differential Data Plane Analysis

Evaluating the functional correctness of the proposed reconfiguration 𝐶 fix requires extracting the network’s emergent high-level behaviour from its low-level configuration files and comparing it against a “ground truth" behaviour. We standardize a network’s high-level function with the following procedure: First, we use Batfish to simulate the data plane of each network, including forwarding decisions for each (src, dst prefix) pair. Then, we construct a forwarding graph for each prefix, and use the graph algorithms proposed in Config2Spec [21] to extract the set of predicates that describe the specifications of the network. Fig. 3 shows an example of this flow.

• Efficacy (Fix Score): The proportion of resolved specifications relative to the total violations (initial and regressions): |Φfixed | Fix Score = |Φfixed | + |Φunfixed | + |Φregressed |

4.3

Diagnosis Evaluator

While the previous procedure quantifies the restoration of the intended network behaviour, we are also interested in evaluating LLMs’ ability to localize (at the router-level) and diagnose root causes. The following parts of a proposed solution are related to this: • The list of routers that are detected • The textual description of the detected faults We use these artifacts to gain insights into how sound (are only real faults detected?) and how complete (are all the faults detected) the proposed diagnoses are. For the task of localizing faulty routers, computing precision and recall is straightforward because they can be compared against the ground truth. To evaluate diagnostic accuracy, a naïve approach would be to have the models classify each fault into a class from the fault library F . However, this would require exposing the model to the list of potential fault types, thereby contaminating the reasoning process and compromising the generalizability of the open-ended diagnosis problem.

Specification-based reconfiguration evaluation. With the previous process we obtain from 𝐶 gold the set of specifications Φ that the reconfigured network must satisfy. To evaluate a proposed reconfiguration, we also need to compare the behaviour between each network state. We do this by extracting and comparing the high-level specifications across networks. To improve efficiency, we construct the forwarding graphs only for prefixes whose entries in the forwarding behaviour table differ from the golden state. We calculate the set of violated specifications V, and after performing the predicate extraction pipeline for 𝐶 fix , we calculate the following sets on which we will base our scoring: • The set of successfully resolved violations Φfixed = {𝜙 ∈ V | 𝐶 fix |= 𝜙 } 8

Benchmarking LLM-driven configuration repair

Preprint, ,

Instead, we use the LLM-as-a-Judge method [22], as a viable and scalable alternative to expert human annotation, where we provide 3 different high-capability LLMs (GPT-5.1, Claude 4.5 Opus, Gemini 2.5 Pro) with (i) the proposed textual diagnosis of the benchmarked LLM, and (ii) the ground truth list of misconfigurations in the network; both the types of faults and the textual differences. We request two separate scores that quantify the completeness (i.e., the percentage of faults correctly identified) and the soundness (i.e., the percentage of faults hallucinated) of the diagnosis. To obtain the final scores, we aggregate scores across the judge models to further ensure robustness and mitigate potential biases.

5

problem into a structured workflow. We instruct the model to use a Chain-of-Thought (CoT) process that mirrors the formal fault management lifecycle standard [37]. We implement a parser that expects three distinct solution parts that map to each goal of the process: • Localization: The list of detected faulty routers. • Diagnosis: A textual diagnosis of the root causes in the configuration. • Reconfiguration: A list of configuration changes for the detected faulty routers, using the mandated searchand-replace format (as described in §4.1). This structured elicitation allows Cornetto to evaluate the accuracy of the intermediate reasoning steps (localization and diagnosis) independently of the final repair quality. Finally, to ensure that format following is not a confounding factor in our evaluation, we configure the parser to allow one retry attempt per scenario if the model produces a solution that cannot be parsed.

EXPERIMENTAL SETUP

In this section, we review the models examined and describe how inputs are constructed to evaluate LLMs on Cornetto. Model selection. We evaluate 9 LLMs of varying sizes on Cornetto: 8 state-of-the-art proprietary LLMs (including GPT-5.2, Gemini 3.0, and Claude 4.5 Opus) that consistently dominate benchmark leaderboards [29, 32] and a single opensource model: GPT-OSS-20B. We include the latter to test the viability of smaller, lower-cost models, though we expect it to be outperformed by the larger proprietary ones.

Metrics and reporting. We evaluate performance for each stage leading up to the configuration repair. 1. Diagnostic Reasoning. To assess the models’ ability to isolate faults before fixing them, we report two key metrics:

Input and context construction. A model receives as input a topology 𝑇 , a set of violated specifications V, and a network-wide configuration 𝐶 broken , which is potentially very large. This data is noisy and voluminous, posing a distinct challenge for the LLM-based troubleshooting workflow [33, 34]. For evaluating LLMs’ robustness in handling this volume and noise, we test the following strategies for including configurations in the context:

• Localization F1: The harmonic mean of precision and recall for the set of faulty routers identified by the model against the ground truth. • Diagnosis Quality: Using the LLM-as-a-Judge method (§4.3), we quantify the soundness and completeness of the model’s natural language explanation against the ground truth misconfigurations provided to the LLMJudges.

• Full context: The model receives the entire networkwide configuration 𝐶 broken . Because of the high context window limits of current models, the vast majority (98%) of cases can fit entirely within the prompt. If a model cannot handle the entire input, configuration files are truncated. • Oracle context: The model receives only the files affected by the misconfiguration. This is an idealized scenario for analysis purposes, since realistically, this information is not known a priori. • Retrieval mode: For a subset of models, we evaluate a two-stage workflow where the LLM is first prompted to retrieve the necessary configuration files for diagnosis. This should yield a superset of the faulty configurations; therefore, we evaluate the success of this step using the recall metric.

2. Functional Correctness. We use differential data plane analysis and report the fix rate and regression rate metrics for each scenario, as defined in §4.2. We also report the percentage of cases correctly resolved, i.e., fixes that satisfy the intended specification without introducing new regressions.

6

EVALUATION

We use Cornetto to evaluate 9 state-of-the-art LLMs across 231 diverse troubleshooting scenarios. Through our analysis, we aim to answer two primary research questions: • RQ1: To what extent can current LLMs autonomously localize, diagnose, and repair network misconfigurations without introducing regressions? We present the performance of all LLMs across our core metrics for diagnostic accuracy and repair functional correctness, summarizing key insights into their capabilities. • RQ2: How do task factors such as topological and configuration scale, fault multiplicity, and extent of disruption impact

Prompting and parsing. Following standard practices in tackling holistic and multi-stage reasoning tasks [35, 36], we design a prompt that decomposes the reconfiguration repair 9

Preprint, ,

Protogeros et al. Fix Score ↑ Localization ↑ Diagnosis ↑ Regression ↓ Success Rate ↑ Cost ($/task) ↓

Model GPT-5.2 (High) Gemini 3 Flash Gemini 3 Pro GPT-5.1 (High) Claude 4.5 Opus Claude 4.5 Sonnet GPT-5 mini (High) Grok 4.1 Fast (R.) GPT-OSS-20B

57.8 55.4 47.2 45.4 44.2 37.2 33.9 4.5 1.7

76.5 73.7 70.9 60.7 59.8 58.1 57.5 17.1 15.2

76.8 70.2 66.7 65.6 64.1 60.6 58.3 45.2 12.5

8.6 11.3 13.9 5.3 3.5 5.4 12.8 4.9 2.0

25.5 24.2 18.6 22.9 24.3 17.3 16.9 0.03 0.01

0.16 0.04 0.18 0.11 0.42 0.28 0.02 0.01 -

Table 2: While models show promise in restoring network state, they rarely achieve complete resolution and frequently introduce regressions.

model performance? We quantify the degradation of model reliability across increasing difficulty gradients.

Fix Score

1.0

Table 2 and Figure 4 present the comprehensive evaluation of all 9 models on Cornetto. The histogram of Fig. 5 illustrates the distribution of fix scores — calculated as the pertask average of the top five models. This quasi-normal distribution confirms a desirable property of the benchmark [19]: it captures a spectrum of complexity rather than being saturated with impossible (score 0.0) or trivial (score 1.0) cases. While LLMs demonstrate potential for diagnostic and repair tasks, our analysis shows that they rarely produce fully correct fixes, with strictly correct resolutions (100% fix rate with zero regressions) occurring in at most 25.5% of cases. This ceiling in performance suggests that current models are best deployed as “Human-in-the-Loop" assistants, consistent with recent research [34] and industry practices [8–10]. We detail the specific factors affecting LLM performance through the following insights:

Full Oracle

0.8 0.6 0.4 0.2

5

B

OS

S-

(R GP

T-

st Fa

4.

1

mi

20

.)

h)

5 (H ni

nn So

ok

T-

Models perform better with access to global context. As shown in Fig. 4, providing the full network-wide configuration almost always outperforms the idealized oracle setting, which only contains the faulty configuration files. This indicates that for top-performing models, the ability to examine configurations contextually outweighs the noise introduced by all the irrelevant configuration data. We quantify this trade-off in the Retrieval mode results (Table 3). When tasked with autonomous context selection, GPT-5 mini achieves high recall (82.6%) of faulty routers, thereby effectively filtering noise and improving performance across all metrics. Gemini 3.0 Flash, however, misses critical configuration files (68.1% Recall), and its performance degrades. This suggests that smaller models can benefit from careful context selection. This finding is consistent with work on multi-agent systems powered by smaller language models [38, 39].

Gr

de au Cl

GP

de au

Cl

ig

4.

4.

et

us Op

1 5. T-

GP

5

h)

o 3

(H

ig

Pr

h as ni mi Ge

mi Ge

GP

T-

ni

5.

2

3

(H

Fl

ig

h)

0.0

Figure 4: Frontier LLMs benefit from access to global context.

25

Count

20 15 10 5 0 0.0

0.2

0.4

0.6

0.8

1.0

Fix Score

Most efficacious LLMs are not always the safest. A high fix rate does not guarantee preservation of previously satisfied specifications. While GPT-5.2 leads in fix rate (57.8%),

Figure 5: Cornetto is not saturated with either impossible or trivial tasks. 10

Benchmarking LLM-driven configuration repair Model 3 Flash 5 Mini

Preprint, ,

Recall

ΔFix

ΔDiag.

ΔRegr.

Δ$

68.1% 82.6%

-4.7% +5.7%

-2.0% +4.0%

+0.5% -2.0%

+0.03 +0.01

sparse, which often causes models to miss the issue entirely or even hallucinate faults — hurting both soundness and completeness. Counterintuitively, as the number of faults increases, the abundance of such signals makes it easier for the model to identify any of the real faults. Yet, models often stop at a partial diagnosis and reconfiguration, ignoring other disrupting misconfigurations.

Table 3: Accurate retrieval of critical configurations can filter out noisy data and improve performance

Major network disruptions impact model performance. As shown in Fig. 9, we observe a general negative correlation between network disruption (percentage of broken predicates) and fix score performance. Notably, GPT-5.1 exhibits the sharpest degradation, whereas more capable models such as GPT-5.2 show resilience in repairing networks under more severe disruptions, indicating the ability to link broader specification violations to their root causes.

several other models achieve lower regression rates, including its predecessor GPT-5.1. In contrast, Claude 4.5 Opus is more conservative in its repairs, achieving a lower score of 44.2% but the lowest regression rate at 3.6%. Accurate diagnoses lead to (but do not guarantee) effective fixes. We analyse the correlation between the diagnosis performance and final repair quality in Fig. 6. We observe a moderate positive correlation between diagnosis accuracy/localization and fix score. Crucially, the cluster of high diagnostic scores that yield poor repair metrics (upperleft quadrant in the scatter plot) represents cases in which models identified issues but failed to resolve them. Thus, while correct diagnoses often lead to correct fixes, they do not guarantee them. GPT-5.2 (High)

Gemini 3 Flash

Diagnosis

Takeaways. The results of our analysis show that: • Global context is critical but noisy: While excessive data volume degraded performance, we found that models benefited from access to global configuration context. A system that effectively handles configuration repair should include a stage in which critical context is filtered in a dependency-aware manner, akin to recent work in [33]. • Verification is a prerequisite for safety: Regressions frequently accompany fixes, which is prohibitive while configuring networks. We posit that a system that automates configuration requires a closed loop with a verifier that proves the safety of solutions before deployment. • Iterative repair is needed for completeness: Models struggle to resolve concurrent faults in a single pass, indicating that monolithic prompting fails at scale despite expansive context windows. Troubleshooting must be decomposed into an iterative agentic workflow that diagnoses, proposes, and verifies solutions until the desired result is achieved.

Gemini 3 Pro Localization

1.0

Score

0.8 0.6 0.4 0.2 r = 0.388 p < 0.001***

0.0 0.0

0.2

0.4

0.6

0.8

1.0

r = 0.358 p < 0.001***

0.0

0.2

0.4

0.6

0.8

1.0

Fix Score

Figure 6: Diagnostic accuracy is necessary but not sufficient for repair.

7 LLM performance degrades at scale. Topological scale, configuration length, and predicate set size are all factors that directly affect the amount of information included in the prompt. Hence, it is unsurprising that model performance consistently degrades with increasing input size, as shown in Fig. 7. This amplifies the need for selecting relevant information for a model’s limited context window.

RELATED WORK

Network verification. Two decades of research have established formal methods to mathematically prove network correctness. Data plane verification [3, 40] checks that the network’s forwarding behaviour satisfies some desired property, and control plane verification [1, 2], verifies that a network configuration will produce a data plane that satisfies some intent [7]. Specification mining [21] builds on these approaches to derive the set of forwarding specifications satisfied by a configuration. Cornetto uses Batfish [1] in conjunction with the specification mining algorithms proposed in Config2Spec [21] as an integral component of its problem formulation. Specifically, we mine the ground-truth specifications from the correct

LLMs struggle more to detect all faults. We observe a revealing divergence between diagnostic soundness (precision) and completeness (recall) as fault multiplicity increases. As seen in Fig. 8, completeness degrades sharply, while soundness exhibits a slight upward trend. We hypothesize the following: In single-fault scenarios, the misconfiguration “signal" is 11

Preprint, ,

Protogeros et al.

1.0

<50k 50k-100k 100k-150k >150k

0.8

Fix Score

in software engineering faces fundamental challenges similar to those in network configuration repair. First, ensuring usability requires automatic, functional verification of solutions. Second, realistic tasks require navigating large volumes of noisy data (whether from entire code repositories or network-wide configurations) to identify useful information and arrive at a solution. Cornetto contextualizes these challenges within the realm of network configuration repair, ensuring that it reflects the complexities of real-world network operations.

Input Tokens

0.6 0.4 0.2

B OS

S-

(R

20

.)

h)

GP

T-

st Fa

4.

1

mi

LLMs for network operations. Interest in LLMs to address the limitations of formal verification is growing, evidenced by recent industry assistants [8–10] and benchmarks [15– 17, 42]. Cornetto complements those efforts by rigorously evaluating end-to-end configuration repair, thereby advancing understanding of AI’s potential and applicability to automated network operations.

Gr

ok

T-

5

de

GP

au Cl

ig

ni

nn So

de au

(H

et

us

4.

4.

ig Op

1 5. T-

GP

5

5

h)

o (H

3 ni

Ge

mi

3 ni mi Ge

Cl

as Fl

ig (H 2 5. TGP

Pr

h

h)

0.0

Figure 7: Repair performance consistently degrades with increasing context length. GPT-5.2 (High) Gemini 3 Flash

Gemini 3 Pro GPT-5.1 (High)

GPT-5 mini (High) Grok 4.1 Fast (R.)

Claude Opus 4.5 Claude Sonnet 4.5

Diagnosis Soundness

Fix Score

GPT-OSS-20B

Diagnosis Completeness

8

1.0

DISCUSSION AND LIMITATIONS

Score

0.8

What about evaluating fancier LLM-based systems? The strategy space for LLM-based configuration repair is immense. Among others, it encompasses different prompting strategies that affect a model’s reasoning process [35, 43], retrieval methods that determine which relevant information is included in the context from large volumes of data [44], and agentic systems, which constitute a distinct design space of their own [39, 45, 46]. In this work, we assess the intrinsic capabilities of LLMs to reason about network state and resolve misconfigurations. We argue that establishing this baseline is critical, as it decouples performance attributable to the model’s reasoning power from gains attributed to complex-system scaffolding. Still, we design Cornetto as a modular platform that supports the evaluation of such advanced setups (including RAG and agentic systems).

0.6 0.4 0.2 0.0

1

2-4

5-8

1

2-4

5-8

1

2-4

5-8

Injected Faults

Figure 8: Models struggle to handle concurrent failures; as the number of root causes increases, diagnosis becomes partial and fix rate degrades. GPT-5.2 (High)

Gemini 3 Flash

Gemini 3 Pro

GPT-5.1 (High)

Diagnosis

Fix Score

Claude Opus 4.5

Localization

1.0

Score

0.8

0.6

What about specifications under failures? Network operators often care about invariants holding across multiple environments (e.g., maintaining reachability under any 2 link failures). While control-plane verification [2] can prove properties across all possible environments, we focus on specifications for a single environment; we posit that performance in this setting serves as a necessary upper bound on model capability. Since our results demonstrate that LLMs already struggle to reason about network state in a single environment, introducing the complexity of failure models is currently premature. Evaluating reasoning across all possible environments efficiently enough to support many benchmark cases will be considered in the future.

0.4

0.2

0.0 0-1%

1-5%

5-10%

>10%

0-1%

1-5%

5-10%

>10%

0-1%

1-5%

5-10%

>10%

Network Disruption (Violated specifications - % of total)

Figure 9: While some models remain robust, many perform poorly on more disruptive network faults.

reference network and use them to rigorously evaluate the functional correctness of the LLM-generated fixes. LLM Benchmarks. Benchmarks such as SWE-Bench [29] and BaxBench [41] indicate that rigorous evaluation of LLMs 12

Benchmarking LLM-driven configuration repair

9

Preprint, , page 496–511, New York, NY, USA, 2025. Association for Computing Machinery. [10] Zhaodong Wang, Samuel Lin, Guanqing Yan, Soudeh Ghorbani, Minlan Yu, Jiawei Zhou, Nathan Hu, Lopa Baruah, Sam Peters, Srikanth Kamath, Jerry Yang, and Ying Zhang. Intent-driven network management with multi-agent llms: The confucius framework. In Proceedings of the ACM SIGCOMM 2025 Conference, SIGCOMM ’25, page 347–362, New York, NY, USA, 2025. Association for Computing Machinery. [11] Google DeepMind. Gemini 3 Pro. https://deepmind.google/models/ gemini/pro/, 2025. Accessed: 2026-02-01. [12] OpenAI. GPT-5 System Card. Technical report, OpenAI, 2025. Accessed: 2026-02-01. [13] Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, NY, USA, 2021. Association for Computing Machinery. [14] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, March 2023. [15] Changjie Wang, Mariano Scazzariello, Alireza Farshin, Simone Ferlin, Dejan Kostić, and Marco Chiesa. Netconfeval: Can llms facilitate network configuration? Proc. ACM Netw., 2(CoNEXT2), June 2024. [16] Kaan Aykurt, Andreas Blenk, and Wolfgang Kellerer. Netllmbench: A benchmark framework for large language models in network configuration tasks. In 2024 IEEE Conference on Network Function Virtualization and Software Defined Networks (NFV-SDN), pages 1–6, 2024. [17] Zhihao Wang, Alessandro Cornacchia, Alessio Sacco, Franco Galante, Marco Canini, and Dingde Jiang. A network arena for benchmarking ai agents on network troubleshooting, 2025. [18] Ioannis Protogeros and Laurent Vanbever. Continual benchmarking of llm-based systems on networking operations. In Proceedings of the ACM SIGCOMM 2025 Posters and Demos, ACM SIGCOMM Posters and Demos ’25, page 70–72, New York, NY, USA, 2025. Association for Computing Machinery. [19] Moritz Hardt. The emerging science of machine learning benchmarks. Online at https://mlbenchmarks.org, 2025. Manuscript. [20] D.R. Kuhn, D.R. Wallace, and A.M. Gallo. Software fault interactions and implications for software testing. IEEE Transactions on Software Engineering, 30(6):418–421, 2004. [21] Rüdiger Birkner, Dana Drachsler-Cohen, Laurent Vanbever, and Martin Vechev. Config2spec: Mining network specifications from network configurations. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20). USENIX Association, 2020. [22] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. [23] Simon Knight, Hung X. Nguyen, Nickolas Falkner, Rhys Bowden, and Matthew Roughan. The internet topology zoo. IEEE Journal on Selected Areas in Communications, 29(9):1765–1775, 2011. [24] Ryan Beckett, Ratul Mahajan, Todd Millstein, Jitendra Padhye, and David Walker. Network configuration synthesis with abstract topologies. In Proceedings of the 38th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI 2017, page 437–451, New York, NY, USA, 2017. Association for Computing Machinery. [25] Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. Large language models can be easily distracted by irrelevant context, 2023. [26] Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How

CONCLUSION

We presented Cornetto, a comprehensive framework for evaluating LLM-driven configuration repair. Cornetto generates diverse scenarios and rigorously evaluates the end-toend troubleshooting process by assessing diagnostic accuracy and formally verifying the correctness of repairs. Our evaluation of 9 state-of-the-art LLMs on 231 generated scenarios reveals their potential to diagnose misconfigurations and their struggle to reliably synthesize correct and safe reconfigurations. By providing a platform for thorough assessment of these repair capabilities, Cornetto contributes towards the advancement of reliable, automated network operations. This work does not raise ethical issues.

REFERENCES [1] Ari Fogel, Stanley Fung, Luis Pedrosa, Meg Walraed-Sullivan, Ramesh Govindan, Ratul Mahajan, and Todd Millstein. A general approach to network configuration analysis. In 12th USENIX Symposium on Networked Systems Design and Implementation (NSDI 15), pages 469– 483, Oakland, CA, May 2015. USENIX Association. [2] Ryan Beckett, Aarti Gupta, Ratul Mahajan, and David Walker. A general approach to network configuration verification. In Proceedings of the Conference of the ACM Special Interest Group on Data Communication, SIGCOMM ’17, page 155–168, New York, NY, USA, 2017. Association for Computing Machinery. [3] Ahmed Khurshid, Xuan Zou, Wenxuan Zhou, Matthew Caesar, and P. Brighten Godfrey. VeriFlow: Verifying Network-Wide invariants in real time. In 10th USENIX Symposium on Networked Systems Design and Implementation (NSDI 13), pages 15–27, Lombard, IL, April 2013. USENIX Association. [4] Ahmed El-Hassany, Petar Tsankov, Laurent Vanbever, and Martin Vechev. NetComplete: Practical Network-Wide Configuration Synthesis with Autocompletion. In USENIX NSDI’18, Renton, WA, USA, 2018. [5] Ryan Beckett, Ratul Mahajan, Todd Millstein, Jitu Padhye, and David Walker. Don’t mind the gap: Bridging network-wide objectives and device-level configurations. In SIGCOMM 2016, August 2016. [6] Rudiger Birkner, Tobias Brodmann, Petar Tsankov, Laurent Vanbever, and Martin Vechev. Metha: Network verifiers need to be correct too! In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21), pages 99–113. USENIX Association, April 2021. [7] Alexander Krentsel, Oliver Ye, Anthony Tafoya, Xuqian Ma, Sylvia Ratnasamy, and Anees Shaikh. Towards accessible model-free verification. HotNets ’25, page 210–217, New York, NY, USA, 2025. Association for Computing Machinery. [8] Haopei Wang, Anubhavnidhi Abhashkumar, Changyu Lin, Tianrong Zhang, Xiaoming Gu, Ning Ma, Chang Wu, Songlin Liu, Wei Zhou, Yongbin Dong, Weirong Jiang, and Yi Wang. NetAssistant: Dialogue based network diagnosis in data center networks. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), pages 2011–2024, Santa Clara, CA, April 2024. USENIX Association. [9] Chenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin, Xuan Zeng, Xinlei Zhang, Zhe An, Gongwei Wu, Jiaqi Gao, Chen Tian, Guihai Chen, Guyue Liu, Yuhong Liao, Tao Lin, Dennis Cai, and Ennan Zhai. Towards llm-based failure localization in production-scale networks. In Proceedings of the ACM SIGCOMM 2025 Conference, SIGCOMM ’25, 13

Preprint, ,

Protogeros et al.

language models use long contexts, 2023. [27] David Mackenzie, Paul Eggert, and Jim Meyering. Comparing and Merging Files with GNU Diff and Patch. Free Software Foundation, 2002. [28] Evgeniy Glukhov, Michele Conti, Egor Bogomolov, Yaroslav Golubev, and Alexander Bezzubov. Diff-xyz: A benchmark for evaluating diff understanding, 2025. [29] Carlos E Jimenez, John Yang, et al. SWE-bench: Can language models resolve real-world github issues? In ICLR, 2024. [30] Paul Gauthier. Aider: Ai pair programming in your terminal, 2023. [31] Vladimir I Levenshtein. Binary codes capable of correcting deletions, insertions and reversals. Soviet Physics Doklady, 10:707, February 1966. [32] Mislav Balunović, Jasper Dekoninck, Ivo Petrov, Nikola Jovanović, and Martin Vechev. Matharena: Evaluating llms on uncontaminated math competitions, 2026. [33] Xi Jiang, Aaron Gember-Jacobson, and Nick Feamster. Caip: Detecting router misconfigurations with context-aware iterative prompting of llms, 2024. [34] Pouya Hamadanian, Behnaz Arzani, Sadjad Fouladi, Siva Kesava Reddy Kakarla, Rodrigo Fonseca, Denizcan Billor, Ahmad Cheema, Edet Nkposong, and Ranveer Chandra. A holistic view of ai-driven network incident management. In Proceedings of the 22nd ACM Workshop on Hot Topics in Networks, HotNets ’23, page 180–188, New York, NY, USA, 2023. Association for Computing Machinery. [35] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023. [36] Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi. Least-to-most prompting enables complex reasoning in large language models, 2023. [37] Information processing systems – Open Systems Interconnection – Basic Reference Model – Part 4: Management Framework, 1989. [38] Peter Belcak, Greg Heinrich, Shizhe Diao, Yonggan Fu, Xin Dong, Saurav Muralidharan, Yingyan Celine Lin, and Pavlo Molchanov. Small language models are the future of agentic ai, 2025. [39] Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. Autogen: Enabling next-gen llm applications via multi-agent conversation, 2023. [40] Haohui Mai, Ahmed Khurshid, Rachit Agarwal, Matthew Caesar, P. Brighten Godfrey, and Samuel Talmadge King. Debugging the data plane with anteater. SIGCOMM Comput. Commun. Rev., 41(4):290–301, August 2011. [41] Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev, Maximilian Baader, Nikola Jovanović, Jingxuan He, and Martin Vechev. Baxbench: Can llms generate correct and secure backends?, 2025. [42] Yajie Zhou, Jiajun Ruan, Eric S. Wang, Sadjad Fouladi, Francis Y. Yan, Kevin Hsieh, and Zaoxing Liu. Netarena: Dynamic benchmarks for ai agents in network automation, 2026. [43] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners, 2020. [44] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau

Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrievalaugmented generation for knowledge-intensive nlp tasks, 2021. [45] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models, 2023. [46] Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. Metagpt: Meta programming for a multi-agent collaborative framework, 2024.

14

Benchmarking LLM-driven configuration repair

A

Preprint, ,

FAULT LIBRARY

Table 4: Comprehensive Fault Catalog listing the protocols affected, the nature of the misconfiguration (Summary), and the resulting impact on the network (Expected Effect). Protocol / Type

Summary

Expected Effect

BGP

eBGP neighbor configured with incorrect remote AS Administratively shut down a BGP neighbor Node configured with incorrect local ASN Force invalid next-hop on eBGP advertisements

eBGP session reset due to ASN mismatch, cutting off inter-AS route exchange BGP Peering is administratively disabled, withdrawing all prefixes learnt via the neighbor Misaligned local ASN breaks iBGP/eBGP sessions and splits the AS control plane Outbound policy rewrites next-hop to an unreachable address, causing downstream traffic blackholes iBGP routes advertised to clients retain original eBGP next-hop, which may be unreachable from clients causing traffic blackholes Prefix is no longer originated, withdrawing reachability from downstream peers Export policy no longer enforced, allowing infrastructure routes (loopbacks, P2P) and unintended prefixes to leak to external peers Inbound filters begin applying outbound and vice versa, breaking intended import/export policy ASBR originates its loopback /32 into eBGP and the peer accepts it because inbound filtering was removed iBGP sessions removed between RR and up to 5 (exclusive) clients, orphaning them from iBGP reachability Conflicting cluster-ids cause route reflectors to drop one another’s updates, stranding clients that now depend on the misconfigured RR (cf. RFC 4456, Sec. 8)

Remove next-hop-self from RR → client iBGP session Withdraw a BGP network statement from the process Remove outbound route-map from eBGP neighbor Swap inbound and outbound route-maps on a neighbor Leak router loopback by stripping export/import policies Break RR sessions to orphan clients Duplicate cluster-id across route reflectors and isolate clients on one RR OSPF

OSPF interface cost set to extreme value Disable OSPF adjacency on a link Node missing OSPF area membership Assign duplicate OSPF router-ID to multiple routers

IS-IS

Disable IS-IS on an intra-AS link Demote a Level-1-2 IS-IS router to Level-1 Assign router to wrong IS-IS area

Addressing

Duplicate loopback IPv4 addresses Link interfaces disagree on prefix length Link interfaces reside in different subnets

Artificially high OSPF cost diverts traffic away from the link based on alternate SPF paths Removing the link from OSPF prevents adjacency formation and withdraws LSAs learned across it Router withdraws from all OSPF areas, tearing down adjacencies and LSAs OSPF adjacencies fail or LSAs rejected due to router-ID collision, fragmenting OSPF domain and blackholing traffic Removing the link from IS-IS prevents adjacency formation and withdraws LSPs learned across it Reduces inter-area reachability by removing a backbone-capable router, risking L2 partitioning Router in wrong area cannot form L1 adjacencies with its physical neighbors; causes partition of L1 domain and reachability loss Two routers share the same loopback, risking routing loops and control-plane instability One side of a point-to-point link uses a mismatched subnet mask, preventing adjacency formation Interfaces on a point-to-point link move to disjoint IPv4 subnets, breaking adjacency formation

Device

Remove supporting static route for advertised prefix

Advertised network disappears once the backing static route is withdrawn, causing a control-plane withdraw

Policy

Remove permit entry from prefix-list Convert BGP route-map permit clause into deny Lower BGP local-preference on inbound policy

Prefix-list no longer matches intended prefixes, causing route filtering to block previously allowed routes Previously exported prefixes are now filtered, withdrawing routes from neighbors Reduced local-preference makes an alternate egress the best path for affected prefixes

Redistribution

Drop BGP → OSPF redistribution on an ASBR

Internal OSPF loses external reachability because Type-5 LSAs are never originated

Security

Insert implicit deny at top of interface ACL

Ingress traffic on the protected interface is dropped before policy permits, breaking connectivity Egress traffic on the protected interface is dropped before policy permits, breaking connectivity

Insert implicit deny at top of outbound interface ACL

15

Preprint, ,

ADDITIONAL RESULTS

0.4 0.2

(a) Diagnosis

B

5

20

GP

T-

OS

S-

4.

.) (R

us

st Fa

ok

Cl

au

1

de

5.

Op

5

h)

4.

ig (H

et

1

nn So

TGP

de au

Gr

Cl

GP

(b) Localization

4.

h

h)

as

ig

Fl 3

5.

ni

T-

mi

GP

mi Ge

T-

2

Pr ni

ig (H

3

SOS

TGP

ok

(H

o

h)

B 20

.)

4.

(R

Fa 1

4.

de au Cl

st

et

(H

nn So

ni

1 5.

mi 5

T-

5

h) ig

ig (H

us Op GP

TGP

Gr

5

h)

o

4.

Pr 3

ni

de

mi

ok

Cl

au

mi

Ge

ni

5. TGP

h

h)

Fl

ig 2

3

(H

SOS

TGP

as

B 20

4.

(R st

Fa 1

4.

de au

0.0

Gr

Cl

.)

5

h) ig

ni

So

5.

mi 5

TGP

et

(H

(H 1

Op

TGP

Cl

nn

5

ig

4. us

3 ni

mi

de au

Ge

ni mi

h)

o

h

Pr

as

ig

Fl 3

(H 2 5. T-

Ge

GP

0.2

0.6

0.0 h)

0.0

0.4

ni

0.2

0.6

mi

0.4

Full Oracle

0.8

Ge

0.6

1.0

Full Oracle

0.8

Regression Rate

Localization

Diagnosis

1.0

Full Oracle

0.8

5

1.0

Ge

B

Protogeros et al.

(c) Regression Rate

Figure 10: Overview of the model leaderboard using other core performance metrics: diagnosis (left) and localization (center) scores, followed by regression rate (right).

1.0

Regression Rate

<50k 50k-100k 100k-150k >150k

0.8 0.6 0.4 0.2

0.6 0.4 0.2

B 20

5 GP

T-

OS

S-

4.

.)

us

Cl

au

de

Fa 1 4.

ok

Op

st

(R

4. et

nn So Gr

au

de

TCl

(a) Diagnosis

5

h) ig

h) (H 5.

1

(H 2 5. GP

T-

GP

GP

ni mi Ge

T-

ig

as 3

(H

Fl

ig

Pr 3 Ge

mi

mi

ni

ni

OS TGP

h

h)

o

B S-

(R

20

.)

h) Fa 4.

1

mi ok Gr

GP

T-

5

de

st

(H ni

nn So

de au

au Cl

ig

4. et

us Op

1 5. Cl

5

5 4.

ig (H

3 ni

TGP

3

mi Ge

ni mi Ge

h)

o Pr

h as Fl

ig (H 2 5. TGP

<50k 50k-100k 100k-150k >150k

0.0 h)

0.0

Input Tokens

0.8

5

Diagnosis

1.0

Input Tokens

(b) Regression Rate

1.0

1.0

0.8

0.8

0.6

Gemini 3 Flash

GPT-5.2 (High)

Gemini 3 Pro GPT-5.1 (High)

0.4

GPT-5 mini (High)

Regression Rate

Fix Score

Figure 11: Diagnosis performance (left) consistently degrades with increasing input prompt tokens. The same trend is noticeable for regression rates (right) too; this potentially stems from the fact that with smaller context models might hallucinate more and break correct predicates.

Claude Opus 4.5 Claude Sonnet 4.5

0.2

0.6 0.4 0.2

Gemini 3 Pro GPT-5.2 (High) Claude Sonnet 4.5 GPT-5.1 (High)

GPT-5Gemini mini (High) 3 Flash

Grok 4.1 Fast (R.)

0.0

Claude Opus 4.5

0.0 0.0

0.1

0.2 0.3 Task Cost ($)

0.4

0.5

0.0

(a) Fix Score

0.1

0.2 0.3 Task Cost ($)

0.4

(b) Regression Rate

Figure 12: Cost-Pareto frontier with respect to average fix score (left) and regression rate (right). 16

0.5

Record · ID 134517 · SHA-256 c844fc20b886a5b5
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.