Conceptio › Archive › arXiv CS
arXiv CSopen access

RuleAutoPilot: Synthesizing Deployable Suricata Rules from Network Traffic

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

RULE AUTO P ILOT: Synthesizing Deployable Suricata Rules from Network Traffic

arXiv:2609.16231v1 [cs.CR] 14 Sep 2026

Mughees Ur Rehman Purdue University [email protected]

Aritran Piplai University of Texas at El Paso [email protected]

Abstract

network traffic against databases of rules (structured signatures that encode protocol-level characteristics of known threats) and raise alerts or block connections when matches are found. Suricata is well suited for high-throughput deployments, partly due to its multi-threaded design [46, 51]. Community-maintained rulesets such as Emerging Threats Open (ET Open) [41] provide curated signatures for known malicious behavior. Suricata documentation describes ET Open as a free ruleset and useful signature reference [3], while prior work shows that NIDS effectiveness depends heavily on rule quality [45]. Despite their effectiveness, signature-based IDS are constrained by a persistent rule creation bottleneck. Writing a high-quality Suricata rule requires substantial analyst effort: analysts must inspect traffic, infer the underlying behavior, and design specific rules that provide useful coverage without creating noisy alerts [17, 47]. The resulting rule must then be tested to ensure that it loads correctly, detects the intended malware traffic, and avoids excessive false positives [32]. This process is difficult to scale: AV-TEST reports that more than 450,000 new malware and potentially unwanted application samples are registered each day [9]. Recent work has explored Large Language Models (LLMs) for automating IDS rule generation. For instance, RuleMaster+ [27] fine-tunes an LLM on instruction datasets derived from proof-of-concept exploit descriptions. FALCON [31] generates Snort and YARA rules from Cyber Threat Intelligence (CTI) reports using an agentic framework. Moreno et al. [32] study zero-shot and few-shot prompting for Suricata rule generation from pre-segmented malicious flows in ICS/SCADA settings, with syntax-only validation. While promising, these approaches still assume that rule generation starts from an already processed security artifact, such as a PoC write-up, CTI report, or analyst-labeled malicious network segment. This assumption moves the hardest part of the workflow outside the system: before a rule can be generated, someone must first analyze the raw evidence and identify the security-relevant behavior. In early-stage incidents, however, security analysts may only have a suspicious

Rule-based Intrusion Detection Systems (IDS) such as Suricata are central to network security, yet crafting effective detection rules demands deep expert knowledge and cannot keep pace with emerging threats. Existing LLM-based approaches can reduce analyst effort, but they either rely on curated threat intelligence that is produced only after the underlying traffic artifacts already exist, or they require costly LLM use without sufficient quality control. We present RULE AUTO P I LOT, an end-to-end agentic framework that generates deployable Suricata rules directly from malware network traffic, with no prior threat intelligence required. A key challenge is noise: network traffic captures often contain a small amount of security-relevant traffic mixed with large volumes of background traffic, which reduces LLM reasoning quality and increases cost. RULE AUTO P ILOT addresses this challenge with a Benign Traffic Fingerprinting stage that removes known benign background flows before LLM processing. Rules that fail syntax checks, do not trigger on the source traffic, or generate false positives on a benign corpus are automatically repaired using structured feedback. Across 1,296 malware PCAPs, execution-grounded verification raises rule quality (F1) from 0.443 to 0.539. On a stratified 200-PCAP subset, RULE AU TO P ILOT on the open-weight gpt-oss-120b reaches nearfrontier quality, 0.524 F1 against Claude Opus 5 under Claude Code’s 0.623, at 52× lower billed-token cost. Swapping only the backbone to Claude Opus 5, RULE AUTO P ILOT surpasses Claude Code outright, 0.656 F1 against 0.623, at 40× fewer tokens. A stronger backbone raises RULE AUTO P ILOT’s own ceiling, but at the same backbone, our scaffold still outperforms Claude Code’s, showing the scaffold contributes independently of the backbone.

1

Murat Kantarcioglu Virginia Tech [email protected]

Introduction

Signature-based Intrusion Detection Systems (IDS) such as Suricata [38] and Snort [12] remain widely deployed components of operational network defense [6, 45]. They inspect 1

binary or a malware execution trace. The malicious behavior has not yet been summarized into CTI, distilled into curated indicators, or separated from benign background traffic. This gap is not a historical artifact. On 3 December 2025, researchers publicly disclosed CVE-2025-55182 (“React2Shell”), a CVSS 10.0 pre-authentication remote code execution flaw in the Flight protocol used by React Server Components and Next.js [15, 48]. No active exploitation had been reported at disclosure; within five days, post-exploitation activity was observed on vulnerable hosts [48]. Among the concrete detection artifacts circulated by defenders during that window were a packet capture of exploit traffic together with Suricata and Snort signatures derived from it [50], which is precisely the traffic-to-rule workflow this paper automates. More broadly, CTI gathering, analysis, and reporting can take days or weeks [44]. Another challenge is TLS encryption, which leaves a rule only the packet headers and the handshake to work with. The problem is new and consequential, since traditional Suricata rules match on payload as well. Yet established rulesets such as ET Open continue to derive substantial value from plaintext inspection: of the 13,645 rules ET Open published over the past two years, 36% target encrypted traffic while 64% read plaintext. A rule-generation system must therefore produce high-quality rules for both encrypted and unencrypted traffic. This gap motivates a central research question: How can we automatically synthesize deployable Suricata rules directly from network traffic, without relying on curated threat intelligence, manually labeled malicious flows, or a security analyst in the loop? To address this problem we propose RULE AUTO P ILOT, an agentic framework for end-to-end Suricata rule synthesis from network traces, designed around three operational requirements: generated rules must be syntactically valid, must trigger on the traffic they were derived from, and must not fire on benign traffic, because false alarms make a rule undeployable regardless of what it detects. The first challenge is noise. Malware traces contain a few security-relevant flows among many background ones, and passing all of them to an LLM raises cost, can exceed the context window, and distracts the model from the flows that matter. RULE AUTO P ILOT addresses this with Benign Traffic Fingerprinting, a protocolaware filtering stage that removes recurring background flows before synthesis. The second challenge is deployability, and RULE AUTO P I LOT answers it with a verification-first design. A synthesis agent identifies security-relevant flows and writes candidate rules; a verification agent then loads each rule in Suricata, replays it against the source malware PCAP, and replays it against a benign corpus. Rules failing any check go to a repair agent that revises them from structured verifier feedback before re-verification. The pipeline therefore does not produce plausible rule text but executable rules tested against real malware and benign traffic.

Building the system makes a second question answerable, and it is the one the paper is ultimately about. Rule generation needs a model, a prompt, and a scaffold: the code around the model that decides what it sees and what happens to its output. It is not obvious which of the three determines quality. If capability dominates, the answer is to wait for larger models. If the scaffold does, an open model that can be run in-house on sensitive traffic already suffices. We answer this two ways. First, we hold RULE AUTO P ILOT’s scaffold fixed and vary the backbone, running it on gpt-oss-120b and on Claude Opus 5, to measure how much the backbone alone contributes. Second, we hold the backbone fixed and vary the scaffold: at gpt-oss-120b, we compare RULE AUTO P ILOT against two general-purpose coding agents, OpenClaw and Hermes Agent; at Claude Opus 5, we compare RULE AUTO P ILOT against Claude Code. Together these tell us whether the backbone or the scaffold drives rule quality. The prompt is measured separately, by sweeping six variants inside one scaffold. These questions structure the paper: RQ1: How well can an agentic LLM pipeline generate deployable Suricata rules from malware traffic, and how closely do they align with ground-truth rules? RQ2: What determines rule quality, the backbone model or the scaffold around it, and does a general-purpose scaffold suffice? RQ3: When traffic is encrypted, what evidence remains visible, and how well can Suricata rules still be generated from it? We evaluate RULE AUTO P ILOT on 1,296 malware PCAPs across 192 malware families, spanning samples from 2014 to 2026, along with 1,172 benign PCAPs across the benign fingerprinting and benign evaluation corpora. The cross-system comparison uses a 200-capture stratified subset drawn from the same corpus. We make the following core contributions: • End-to-end rule generation from network traffic. RULE AUTO P ILOT generates Suricata rules directly from malware PCAPs, with no CTI report, CVE description, or analyst-labeled flow as input. • Benign traffic fingerprinting. We introduce a protocolaware filtering stage built from 981 benign PCAPs. It removes 88.8% of background flows from malware PCAPs while retaining 98.0% of security-relevant flows, reducing average LLM token usage by 1.99×. • Execution-grounded validation and repair. Every candidate rule is loaded in Suricata, replayed on the source malware PCAP, and replayed on benign traffic, with failures sent to a repair agent. Together these raise flowalignment F1 (how closely a generated rule’s alerts match the flows an expert-authored rule would flag) from 2

2.2 Deep Packet Inspection and Stateful Matching

alert dns $HOME_NET any -> any any ( msg:"ET MALWARE Diezen/Sakabota CnC Domain Observed in DNS Query"; dns.query; content:"antivirus-update.top"; nocase; endswith; classtype:domain-c2; sid:2029326; rev:2; metadata:created_at 2020_01_29, deployment Perimeter, confidence High, signature_severity Major, updated_at 2020_01_29; )

Suricata rules match in two distinct ways, and this paper’s results turn on the difference. Content matching. Most rules inspect bytes, either raw payload or one of the many parsed protocol buffers Suricata exposes, including http.uri, dns.query, and tls.sni. This is deep packet inspection, and it requires those bytes to be readable. Transport encryption hides everything above the TLS record layer, including HTTP headers, paths, and bodies, while the TLS handshake itself, including the server name, stays visible before encryption begins. What a capture exposes therefore bounds what a content rule can key on (§7). Stateful matching. A minority of rules express conditions over accumulated network state rather than a single packet. threshold and detection_filter keywords count events within a time window: ET Open rule sid:2001581, for instance, alerts only when one internal host opens at least 70 TCP connections to port 135 within 60 seconds. flowbits sets and tests named bits so that one rule can condition on another having matched earlier in the same flow.

Figure 1: Rule matching a malware C2 domain in DNS traffic. 0.443 to 0.539 and cut the benign false-positive rate from 0.128% to 0.006%. • At a fixed backbone, the scaffold determines rule quality. At a fixed gpt-oss-120b backbone, RULE AU TO P ILOT reaches 0.524 flow-alignment F1 where the same model driving the OpenClaw [39] and Hermes [34] coding agents, on an identical prompt, reaches 0.416 and 0.400. Every system receives the same benign falsepositive control, so the comparison does not reward RULE AUTO P ILOT for a stage the baselines lack. • Frontier-model rule quality at open-model cost. Claude Opus 5 under Claude Code [7] reaches 0.623 F1 against RULE AUTO P ILOT’s 0.524 on gpt-oss-120b, at 52× more billed tokens per capture. A scaffold built for rule generation therefore comes close to frontier-model quality at a fraction of the cost.

2.3 Rule Maintenance and Community Rulesets Rule authors must balance specificity against generality: broad rules produce false positives, narrow ones miss variants of the same behavior. Community-maintained feeds such as URLhaus and Feodo Tracker publish Suricata-compatible rules for known malicious infrastructure [2,4]. These feeds operationalize known indicators quickly, but they are indicatordriven: they match specific URLs, domains, IP addresses, or IP:port pairs, so their coverage degrades when operators rotate infrastructure. Emerging Threats Open (ET Open) [41] instead provides rules that inspect protocol fields, payload content, and malware-specific network behavior. The Suricata documentation describes it as a free ruleset with a wide range of signature examples, and the ET community publishes frequent updates [1, 3]. ET Open has been adopted as a baseline or comparison ruleset in prior IDS evaluations [13, 26, 52], which makes it a reliable reference for studying how highquality signatures capture malicious network behavior. In this paper we use it as our ground truth.

• At equal backbone, our scaffold outperforms Claude Code’s. Pairing RULE AUTO P ILOT’s own scaffold with the same Claude Opus 5 backbone that powers Claude Code surpasses Claude Code’s F1 outright, 0.656 against 0.623, while using 40× fewer tokens. Holding the backbone fixed isolates the scaffold as the source of the gap.

2

Background

We provide background on Suricata rule structure and the community-maintained rulesets used in operational deployments.

2.1

Rule Structure

Suricata analyzes traffic against structured rules to raise alerts [38]. Each rule has an action and a header defining protocol, address, port scope, and direction. It also has a set of options that define matching conditions over flows, payloads, and protocol fields, plus metadata for identification and maintenance. Figure 1 shows these parts on a C2 domain rule. alert is the action, and the header restricts matching to DNS traffic leaving the monitored network. Among the options, dns.query selects the query field and content looks for the domain antivirus-update.top within it. The classtype, sid, rev, and metadata fields identify and classify the rule.

3

Problem Statement and Threat Model

Problem Statement. Rule authoring requires expertise in traffic analysis, malware behavior, and rule construction, so manual generation struggles to keep pace. Deployment needs more than plausible rule text: candidates must be executable by Suricata and validated against both the source malware traffic and benign traffic. Given a trace, RULE AUTO P ILOT pro3

duces rules that are syntactically valid, trigger on the source traffic, and stay quiet on benign traffic. Assumptions. We assume that RULE AUTO P ILOT is given either an existing network traffic capture from malware execution or a malware binary that can be executed in a controlled environment to obtain such traffic. RULE AUTO P ILOT does not require curated CTI reports, CVE write-ups, attacker-IP lists, proof-of-concept descriptions, or malware-family labels as input. In addition, RULE AUTO P ILOT requires access to a Suricata environment for rule validation and a benign traffic corpus for false-positive testing. We assume that rule generation is performed in a trusted environment. Adversary Model. We consider malware that exhibits network-visible behavior, such as command-and-control communication, staged payload delivery, credential theft, or exfiltration. We assume that this behavior is at least partially expressible in Suricata’s rule language. We do not consider an adaptive adversary that observes RULE AUTO P ILOT and then crafts traffic specifically to evade the generated rules. Analyzing such adversarial adaptation is beyond the scope of this work.

4

and execute them in Cuckoo Sandbox [11] to collect PCAPs. From these we extract recurring protocol-level patterns that characterize normal execution: DNS query names, HTTP hosts, URIs and User-Agents, TLS SNI values and stable handshake identifiers, and recurring artifacts such as NTP destinations, OCSP behavior, and broadcast traffic. Table 12 lists the categories. When processing a PCAP, RULE AUTO P I LOT extracts the same features per flow and removes library matches, retaining the rest as candidates. The filter is deliberately conservative because a discarded security-relevant flow cannot contribute to any rule. The filter removes 88.8% of background flows while discarding 2.0% of security-relevant ones; §6.4.1 and Table 13 quantify the effect.

4.2

The Rule Synthesis Agent receives the candidate flows surviving benign fingerprinting and decides which contain behavior suitable for rule generation. This is narrower than malicious/benign classification: a flow is useful only if it exposes stable, discriminative features expressible in Suricata’s rule language. Working from the structured flow representation rather than raw packets, the agent looks for two kinds of evidence. The first is a protocol field that names a specific indicator, such as a DNS query, HTTP path, host header, or TLS SNI value. The second is behavior visible only across multiple flows, such as a timestamped connection timeline, per-destination connection counts and inter-arrival times, or the number of distinct destinations a host contacts. A flow lacking either kind of evidence, or carrying only routine activity, unstable artifacts, or indicators too broad to be reliable, is treated as background; ambiguous evidence resolves toward retention. A triage step examines each capture and decides which call to invoke: captures with specific-indicator evidence go to the content call, captures with cross-flow behavioral evidence go to the stateful specialist call, and captures with both kinds of evidence go to each. Deep packet inspection. The content call writes rules that match payload bytes or a parsed protocol buffer with an appropriate protocol and direction. It avoids rules resting only on incidental IP addresses or ports, and prefers protocol-aware conditions specific enough to avoid benign matches. Where several flows share a pattern, it writes one rule capturing their invariant features rather than one per flow; the rule need not reproduce the exact ET Open rule, since behaviorally equivalent rules may differ syntactically. Stateful rule generation. The specialist call carries roughly 600 tokens of stateful-only instruction and none of the contentmatch demonstrations, so the examples that make content rules precise do not pull it back toward writing one. Rules it produces typically carry a content anchor or a specific destination alongside their rate condition, following the same

RULE AUTO P ILOT Architecture

RULE AUTO P ILOT synthesizes and validates Suricata rules from malware network traces. It accepts either a suspicious file, which it executes in a sandbox to capture a PCAP, or an already collected trace. The PCAP is decomposed into flows enriched with Suricata-compatible protocol fields, flows matching recurring benign execution patterns are removed, and the remaining candidate flows are passed to the Rule Synthesis Agent. Rather than forcing a rule per retained flow, RULE AU TO P ILOT leaves it to the model to synthesize rules only where a flow or flow group exposes stable, rule-expressible evidence. Every generated rule is then validated by Suricata execution, and rules that fail are repaired from structured feedback. Four components implement this: Benign Traffic Fingerprinting (§4.1) removes recurring benign execution patterns; the Rule Synthesis Agent (§4.2) separates security-relevant from background flows and writes rules for the former; the Rule Verification Agent (§4.3) checks syntax, malware triggering, and benign false positives by execution; and the Rule Repair Agent (§4.4) revises failures from structured feedback and a curated repair knowledge base. Figure 2 shows the four.

4.1

Rule Synthesis Agent

Benign Traffic Fingerprinting

Malware traces often carry hundreds to thousands of flows, most of them benign background activity: DNS lookups, update checks, time synchronization, certificate validation. To reduce this noise, RULE AUTO P ILOT begins with Benign Traffic Fingerprinting. We take 981 benign samples from a DIKEderived corpus [22], confirm their status with VirusTotal [49], 4

Traffic Preparation Suspicious File

Sandbox

Benign Traffic Fingerprinting

Flow Decomposition

PCAP

Candidate Flows

✓

✓ ✓

Binary

Dynamic Execution

Network Traffic Capture

Match against benign PCAP fingerprints

Extract 5-tuple Deep-packet inspection

Rule Verification Agent

① Syntax Check

suricata -T (Rule syntax check)

valid ✓ invalid ✗

Security-relevant flows Background flows

② Malware Replay

Legend

RuleAutoPilot Generated.rules

Rule must trigger on malware traffic

, Candidate Rules

✓

Validated Rule Output

Rule Synthesis and Verification Rule Synthesis Agent

✓

triggers ✓ = pass no trigger ✗ = fail

③ Benign Replay

Rule must NOT trigger on Benign Traffic

,

no triggers ✓ = pass triggers ✗ = fail

Filtered flows Candidate flows

Rule Repair Agent

Background flows Security-relevant flows

Fixed

Repair

Invalid

Structured Failure Feedback

, Figure 2: RULE AUTO P ILOT architecture for traffic preparation, rule synthesis, verification, and repair. logic as ET Open’s own rate-based signatures, which almost always pair a threshold with a content match rather than firing on volume alone. Every produced rule is combined into a single set of candidate rules, which then goes to the Rule Verification Agent. We evaluate six prompting variants for the content call (§6.2); templates are in Figures 6–8.

4.3

flowbits keyword are an exception, since they belong to the stateful rule class and come in pairs: one sets a named bit when it matches and raises no alert itself, and a second fires only if that bit is already set. Verifying the setter alone would record a trigger failure for a rule that is working correctly, so this check is applied to the pair as a unit, and a rule that tests a bit is rejected unless the rule setting it is also present. Benign false positive. Rules that trigger are replayed against a benign PCAP corpus, where firing means the rule is too broad. The verifier records both the background flows that fired and the malware flows the rule matched, which is what lets the repair agent narrow it while preserving the intended match. A rule passing all three checks is accepted.

Rule Verification Agent

The Rule Verification Agent determines whether a candidate rule is deployable by running it with the Suricata engine. A rule is accepted only if it passes three checks in sequence, and the failure label and execution context of the first failed check become the repair agent’s input. Syntax. Each candidate is loaded with suricata -T, which catches unsupported keywords, invalid headers or directions, a missing sid, misused sticky buffers, and invalid PCRE. The parser error is recorded verbatim. Malware trigger. Survivors are replayed against the source malware PCAP. A trigger does not prove the matched flow malicious; it confirms the rule captures behavior present in the execution. Failures, most often direction mismatches or overly restrictive content matches, are recorded with the flow context the rule was expected to match. Rules that use the

4.4

Rule Repair Agent

When verification rejects a candidate, the Rule Repair Agent attempts repair within a fixed budget (K = 3 throughout). Which strategy applies depends on the failure and on which synthesis call produced the rule. Deterministic rewrites. For failure patterns with a precise programmatic fix, a rewrite is applied before any LLM call, such as swapping source and destination variables when the direction contradicts the header. A rewritten rule that passes 5

Table 1: Composition of the 200-capture cross-system evaluation subset.

verification is accepted with no LLM call; Table 18 lists all 17. Context-aware retrieval-augmented repair. Otherwise the agent uses the verifier’s structured feedback, comprising failure type, verifier output, and flow context, to retrieve a matching entry from a curated knowledge base of 42 error classes and build a failure-specific prompt. Syntax prompts carry the parser error and the retrieved guidance. Trigger prompts carry the flows the rule was expected to match. False-positive prompts carry both the background flows that fired and the malware flows that matched. The knowledge base was curated from documented Suricata constraints, parser-error patterns, and practitioner repair workflows [18,35–37]. Retrieval keeps each prompt narrow rather than exposing the LLM to unrelated advice. Tables 14–19 and Figures 9–12 give the full taxonomy, strategy coverage, and templates. Stateful bundles. When a stateful bundle fails verification, the rule itself is already valid Suricata syntax; the problem is what it is detecting, not how it is written. So the deterministic rewrites used for other rules do not apply here. One rewrite in particular is deliberately withheld: stripping the threshold clause, since that clause is the very condition the rule exists to enforce, and removing it would defeat the rule’s purpose. Instead, the repair agent draws on seven behavioral failure patterns specific to bundles. These include a rate-based rule with no content match to anchor it and a stateful bit that a rule checks for but that no rule ever sets, among others. After each attempt the verifier re-runs the full sequence. Rules still failing after K iterations are excluded from the accepted rule set.

5

Ground-truth type

PCAPs

Families

Content, cleartext Content, mixed Content, encrypted Stateful

60 60 40 40

60 52 15 18

Total

200

104

whether the pipeline wrongly generates rules where it should generate none. Rules-PCAP ground truth. Running Suricata with ET Open on each PCAP yields the ground-truth rules it triggers. Every malware PCAP triggers at least one; no benign PCAP triggers any. The triggered rules span created_at dates from 2010 to 2026 (Table 9), so both older and recent signatures remain relevant here. Flow decomposition. We decompose each PCAP into 5-tuple flows and enrich each with Suricata-compatible protocol fields extracted by tshark, selected per protocol using heuristics from the ET Open ruleset. Fields without Suricata equivalents are dropped and tshark names remapped where necessary; Table 11 lists the full set. A flow is labeled security-relevant if it triggers at least one ET Open rule and background otherwise, which yields a dataset-level mapping from each securityrelevant flow to its triggered rule(s). Stratified subset for cross-system comparison. Running four additional systems over all 1,296 captures is not affordable for the frontier model, so we draw a 200-capture stratified subset spanning 104 families (Table 1). Strata follow what the ground truth requires: a content match on an indicator visible in cleartext, on a partly encrypted trace, on a fully encrypted trace, or a rate or ordered sequence no single-packet rule expresses, which we call stateful. That last stratum is 81 of 1,296 captures (6.2%) and is oversampled to 40 of 200 for statistical power, so per-stratum scores are reported separately rather than pooled. The encrypted stratum draws its 40 captures from only 15 families and is more family-concentrated than the others. Dataset summary. Table 2 summarizes the malware and benign datasets used in our evaluation.

Dataset and Ground Truth Construction

A core challenge in evaluating automated IDS rule generation is the lack of standardized datasets that pair network traffic with rules. We therefore construct malware and benign traffic corpora from real execution artifacts, prioritizing real-world samples over simulated traces. Malware corpus. We collect Windows PE malware from MalwareBazaar [5], ANY.RUN [8], and Triage [42], executing binaries in Cuckoo Sandbox [11] and using Triage PCAPs directly. Each sample’s VirusTotal report [49] supplies first_submission_date, giving a 2014–2026 span (Table 8), and family labels are normalized with ClarAVy [23]. The benchmark holds 1,296 PCAPs across 192 families, with 15 unlabeled PCAPs treated as singletons (Table 10). Benign corpora. We use two benign corpora, kept disjoint so that fingerprinting data never overlaps abstention-evaluation data. The first is 981 DIKE-derived PCAPs [22], used to build the fingerprint library and, in subsets, for benign replay and false-positive testing. The second is a 191-PCAP external benchmark from the AsiaCCS 2021 benign binaries dataset [28], all VirusTotal-verified benign, used to test

6

Evaluation

We evaluate RULE AUTO P ILOT through the three research questions of §1.

6.1

Metrics

Flow classification, syntax pass, and trigger rate. Flow classification is the binary task of separating security-relevant 6

Table 2: Dataset summary for the rule-generation benchmark and benign corpora.

Ground_Truth.rules

Malware PCAP

RuleAutoPilot Generated.rules

Metric Malware corpus Malware benchmark PCAPs Unique malware families Single-PCAP (no-family) PCAPs Security-relevant flows Background flows Unique ET Open rules triggered Benign corpora DIKE-derived benign PCAPs External benign-execution benchmark

Value 1,296 192 15 6,937 124,850 617

Flows Triggered

Flows Triggered

Flow Not Triggered Flow Triggered

981 191

Flow Match Flow Mismatch

...

flows from background, using the labels of §5. Syntax pass rate is the fraction of generated rules that load in Suricata after repair; malware trigger rate is the fraction of accepted rules that fire on the source PCAP, the minimum behavioral requirement for deployment. Flow Alignment Score (FAS). FAS measures how closely a generated rule matches the flow-level alert behavior of the corresponding expert-authored ET Open rule, as illustrated in Figure 3. We execute both rules on the same malware PCAP and compare the alerted flow sets. FAS precision is the fraction of generated-rule alerts that match on groundtruth security-relevant flows; FAS recall is the fraction of ground-truth security-relevant flows covered by the generated rules. Coverage. A system emitting no rules for a capture has produced no detection, so we score it zero there. Because that alone understates deliberate abstention, we also report coverage, the fraction of captures answered. Where coverage falls well short of 100%, we give precision over answered captures alongside the pooled figure, since the two then measure different things. Over-fire ratio. Over-fire is alerted flows divided by groundtruth security-relevant flows, pooled across captures; 1.0 means the system raises as many flow-level alerts as the reference ruleset. Precision cannot separate a rule firing twice too often from one firing fifty times too often, and analyst load tracks the latter. False positive rate (FPR). FPR is the fraction of benign PCAPs triggered by the generated rule set: the number of benign PCAPs with at least one alert divided by the total benign PCAPs tested. Lower FPR means fewer false alerts on benign traffic.

6.2

Flow Alignment

...

Figure 3: Flow-level comparison of ground-truth and generated rules on the same malware PCAP. FAS quantifies how closely generated-rule alerts align with ground-truth alerts at the flow level.

additionally evaluate GPT-5.5 on a stratified 100-PCAP subset. Benign replay for verification and false-positive testing draws on disjoint subsets of the 981-PCAP DIKE-derived corpus, with an external 191-PCAP benchmark [28] reserved for benign-only abstention. For RQ2, we compare RULE AUTO P I LOT against four additional systems on the 200-capture stratified subset of §5 under a uniform benign false-positive control, with all baseline agents receiving a byte-identical prompt (verified by hash) to isolate scaffold and model effects. Full model parameters, prompting details, the comparison-system harness, the false-positive control procedure, token-accounting methodology, and the Suricata validation environment are given in Appendix B.

6.3

RQ1: End-to-End Rule Generation

RQ1 asks how well RULE AUTO P ILOT generates deployable rules from raw traffic. We answer this along three axes: i) rule quality on malware PCAPs, measured by flow-classification F1, syntax pass rate, malware trigger rate, and the Flow Alignment Score (FAS); ii) whether the pipeline correctly abstains on benign traffic rather than manufacturing rules where none are needed; and iii) the generalizability of generated rules to unseen samples within and across malware families, rather than overfitting to their source PCAP. Evaluation on malware PCAPs. We try different in-context learning configurations, including zero-shot, chain-of-thought, and few-shot prompting, and use few-shot-5 as our default throughout. We report macro-averaged scores, weighting each

Experimental Setup

The primary backbone is gpt-oss-120b, evaluated under six prompting strategies (zero-shot, CoT, few-shot-1/3/5, CoT+FS-3); we use few-shot-5 throughout unless noted, and 7

PCAP equally regardless of its flow count, since a corpus dominated by a few flow-heavy captures would otherwise obscure how the pipeline performs on a typical sample. Under this scoring, RULE AUTO P ILOT reaches FAS-F1 0.54 across the 1,296 malware PCAPs, with an 85.3% syntax pass rate, an 81.1% malware trigger rate, and a 0.006% benign FPR. Prompt choice barely changes this: across the three configurations (Table 3), FAS-F1 moves by only 0.018. Performance on malware unseen during training. Because RULE AUTO P ILOT needs no labels or pre-written IoCs, generated rules should not depend on memorized training content from the model. To check this, we evaluate on post-cutoff malware: gpt-oss-120b’s training cutoff is June 2024, and 1,116 of the 1,296 malware PCAPs postdate it. These PCAPs trigger 480 unique ET Open rules, 214 of which were created after the cutoff, so the model cannot simply be recalling them. On this subset, RULE AUTO P ILOT with few-shot-5 reaches a FAS-F1 of 0.533, closely tracking the corpus-wide figure of 0.54, with an 82.1% trigger rate and an 85.1% syntax pass rate. The pipeline leans toward recall over precision here, which is arguably the right trade in a zero-day setting: missing a security-relevant flow is more costly than surfacing an extra candidate for an analyst to dismiss. Benign FPR stays low at 0.0072%, so this recall bias does not flood analysts with rules on benign traffic. These results show that RULE AUTO P ILOT performs well on malware the model could not have memorized, so its rule quality comes from reasoning over observed traffic rather than recall. Evaluation on Benign PCAPs. On the external 191-PCAP benign benchmark, none of which contains a security-relevant flow, fingerprinting removes 98.0% of flows and filters 126 PCAPs entirely, requiring no LLM call. Of the remaining 65, exactly one yields a rule, flagging a deprecated TLS 1.0 connection rather than malware behavior. RULE AUTO P I LOT reaches 99% flow-classification accuracy and emits no malware-style rule across the 191 benign executions. Generalizability of rules A rule that fires on its source capture may be memorizing it. Replaying each capture’s rules against every other and labeling pairs same- or other-family, RULE AUTO P ILOT reaches FAS-F1 0.67 on same-family targets against 0.75 for the expert-authored rules, triggering slightly more often (0.58 versus 0.55). On other-family targets both collapse and RULE AUTO P ILOT falls further (0.06 versus 0.16), so the extra triggers are not indiscriminate. Rules therefore transfer to unseen same-family samples at nearreference accuracy while retaining cross-family specificity. Table 4 summarizes these rates: on same-family targets, fewshot-5 triggers slightly more often than ground truth (0.58 vs. 0.55) but with lower precision (0.83 vs. 1.00) and recall (0.57 vs. 0.60), so generated rules alert on some flows that are not security-relevant while missing a portion of those that are. Ground-truth precision is 1.00 in both relations by construction: a target’s security-relevant flows are defined as the flows ET Open alerts on, so ground-truth rules cannot alert

outside that set. On other-family targets, few-shot-5 triggers more often than ground truth (0.09 vs. 0.07), yet recall drops sharply (0.09 to 0.03), indicating the additional triggers predominantly fire on flows not security-relevant to the target family. A per-family-pair breakdown, including the heatmap of trigger-rate deviations (Figure 13), is in Appendix I.

6.4

RQ2: What Determines Rule Quality

RQ1 establishes that the pipeline works, not why. Rule quality could come from RULE AUTO P ILOT’s own components, from the general strategy of scaffolding an LLM at all, or from the backbone model itself. We isolate these three explanations in turn: (i) an internal ablation removes RULE AUTO P ILOT’s own components (benign filtering, the verification agent, and the repair agent) one at a time to measure each component’s contribution, and separately swaps the backbone under a fixed scaffold to measure how much the backbone itself matters; (ii) a cross-system comparison holds the backbone fixed and swaps RULE AUTO P ILOT’s scaffold for two general-purpose coding agents, testing whether a task-specific scaffold is necessary or a general-purpose one suffices; and (iii) a comparison against a state-of-the-art model and scaffold asks what the best available combination can achieve relative to ours. 6.4.1

Isolating RULE AUTO P ILOT’s Own Components

We isolate the contribution of each component in turn to see how much they contribute in rule quality. Benign traffic filtering. RULE AUTO P ILOT removes background flows before the Rule Synthesis Agent sees them. Across 1,296 PCAPs, the dataset contains 131,787 flows. Of these, 6,937 (5.3%) are security-relevant, meaning they trigger at least one Suricata rule. Benign fingerprinting removes 110,879 flows (88.8%) while discarding only 142 (2.0%) of the security-relevant flows. This cuts the median flows per PCAP from 65 to 12. As Figure 4 shows, filtering halves cost without hurting quality. Mean tokens per PCAP fall from 40.16k to 20.20k, a 1.99× reduction, and benign FPR falls from 0.019% to 0.006%. Flow classification shifts along the precision–recall tradeoff rather than improving outright. This happens because the unfiltered prompt warns that many flows are routine background traffic, making the model more cautious about flagging them. Fewer flags means higher precision (0.46) but lower recall (0.72). Once background flows are removed, that warning is gone, so the model flags more freely. Recall rises to 0.79 but precision falls to 0.44. The LLM classifies flows competently on its own. Filtering simply buys the same overall quality (FAS-F1 0.53 to 0.54, F1 0.51 to 0.52) at half the cost. Verification-agent ablation. We isolate the contribution of each stage in the verification pipeline by holding the synthesis output fixed and enabling stages one at a time. Table 5 shows 8

Table 3: End-to-end performance for zero-shot, CoT, and few-shot-5, macro-averaged over 1,296 malware PCAPs (full verification, gpt-oss-120b). The full six-variant sweep is in Appendix J.1. Flow Classification Variant

P

R

F1

FAS-P

FAS-R

FAS-F1

Syntax

Trigger

FPR

Zero-shot CoT Few-shot-5

0.45 0.44 0.44

0.79 0.77 0.79

0.52 0.51 0.52

0.50 0.50 0.51

0.67 0.64 0.68

0.53 0.52 0.54

86.2% 84.5% 85.3%

80.9% 79.2% 81.1%

0.000% 0.006% 0.006%

(iii) Rule Quality

0.79

0.8

30k

20.20k

20k

0.72

0.6

0.52 0.51 0.44 0.46

0.4

10k

0.2

0k

0.0 With Benign Filtering

Without Benign Filtering

1.0

0.030 0.81 0.82

0.8 0.68

0.66

0.6

0.025 0.019%

0.015

0.4 0.010 0.006%

0.2

0.005

0.0

Precision

Recall

F1

With Benign Filtering

0.020

0.54 0.53

0.51 0.51

Benign FPR (%)

1.0

40.16k

40k

Score

Mean tokens / PCAP

(ii) Flow Classification

1.99× cost reduction

Score / Rate (Full verification level)

(i) LLM Token Cost 50k

Rule Quality

0.000

FAS Prec

FAS Recall

FAS F1

Trigger Rate

Benign FPR

Without Benign Filtering

Figure 4: Impact of benign traffic filtering on (i) LLM token cost, (ii) flow-classification performance, and (iii) rule quality. Table 4: Family generalization on the 1,074-source nonsingleton subset. Metrics are micro-averaged over source– target pairs.

Table 5: Verification-agent ablation on 1,296 PCAPs. Stages are enabled cumulatively. FAS

Relation

Rule Source

Trigger Rate

Same-family Same-family

Ground Truth Few-shot-5

0.55 0.58

1.00 0.83

0.60 0.57

0.75 0.67

Other-family Other-family

Ground Truth Few-shot-5

0.07 0.09

1.00 0.38

0.09 0.03

0.16 0.06

FAS Precision

FAS Recall

FAS F1

Rule Quality

Stage

P

R

F1

Syn.

Trig.

FPR

None + Syntax + Malware + Benign

.435 .445 .506 .507

.531 .549 .679 .680

.443 .456 .537 .539

81.8% 85.2% 85.2% 85.3%

72.1% 71.5% 80.9% 81.1%

0.091% 0.091% 0.128% 0.006%

ately. The repair agent attempts 247 fixes and succeeds on 124, a 50.2% success rate, with malware-trigger repairs landing more often (61%) than syntax repairs (40%). Two residues remain: 37 rules hit Suricata errors with no knowledge-base entry and failed under generic prompting, and 25 had their syntax fixed by keyword substitution while their content predicates stayed on the wrong sticky buffer, so replay still failed. Table 22 breaks outcomes down by failure type. Backbone sensitivity within RULE AUTO P ILOT. We test whether swapping gpt-oss-120b for a frontier model changes how RULE AUTO P ILOT performs. The ablations above hold the backbone fixed and vary the scaffold; this test does the opposite, holding the scaffold fixed and swapping the backbone for GPT-5.5 (Figure 5), on a family-stratified 100PCAP subset sampling up to three PCAPs per malware family until reaching 100. GPT-5.5 outperforms gpt-oss-120b

that each stage catches a different kind of bad rule. Syntax validation catches rules Suricata cannot parse, raising the pass rate from 81.8% to 85.2%. Since these rules were unusable regardless of detection quality, removing them barely moves FAS-F1. Malware replay tests each rule against the source malware traffic and discards rules that never fire, so it removes rules that pass syntax but detect nothing. This is what drives detection quality, raising FAS-F1 from 0.456 to 0.537 through gains in both precision and recall. Benign replay tests each rule against benign traffic and discards rules that misfire on it. Since these rules already detect malware correctly, removing them barely changes FAS-F1, but it cuts benign FPR from 0.128% to 0.006%. Together, the three stages take FAS-F1 from 0.443 to 0.539. Repair agent effectiveness. Across the corpus, 2,530 rules are generated and 2,091 (82.6%) pass every check immedi9

across every metric we track, most notably raising FAS-F1 from 0.561 to 0.701, so a stronger backbone does raise the pipeline’s ceiling. What stays the same is what each verification stage is for. Syntax checking still does nearly all of the work of getting rules to parse, moving the syntax pass rate from 76.7% to 82.0% for gpt-oss-120b and from 85.4% to 90.0% for GPT-5.5, before either model’s rules are ever tested against traffic. Malware replay still drives whether a rule actually fires on the malware it was written for, lifting the trigger rate for both models by a similar margin. Benign replay still does the work of suppressing false positives, taking gpt-oss-120b’s unfiltered benign FPR from 2.09% down to 0.05% and keeping GPT-5.5 close to zero throughout. In short, the backbone changes how high RULE AUTO P ILOT can reach, but not what each stage of the pipeline contributes to getting there. 6.4.2

least one rule, precision rises to 0.662, the highest among the three same-backbone systems. Abstention concentrates in the stateful stratum, where evidence must be aggregated across many flows rather than read from one, and RULE AUTO P ILOT declines rather than emit a weakly supported rule. This selectivity is also cheap. RULE AUTO P ILOT processes 23,836 tokens per capture against 885,734 for OpenClaw on the same open model, a 37× reduction, for 0.144 more precision. Selecting flows before generation and verifying rules after it is, on this corpus, both cheaper and more precise than letting a general-purpose agent explore the capture unguided. 6.4.3

Backbone and Baseline: From Regex to a Frontier Model

The comparison above holds the backbone fixed to isolate scaffold effects. To see what lies at either end of that comparison, we bracket it with two systems outside it: a naive regex extractor with no model at all, showing what selection-free coverage looks like, and Claude Code, a general-purpose coding scaffold paired with a frontier model instead of a weak one, showing what a stronger backbone can add. We then swap scaffolds at that same frontier backbone, pairing our own scaffold with Claude Opus 5 as well. Table 6 reports all three of these systems alongside the three same-backbone systems above. No intelligence, no selection. The naive regex extractor emits one rule per protocol-field indicator it observes anywhere in the capture, with no judgment about whether that indicator is malicious. It reaches the highest recall of any system, 0.809, simply by proposing a rule for everything, but its precision collapses to 0.130 and it alerts on 8.18× the reference flow count, 31 rules per capture on average. FAS-F1 lands at 0.195, the lowest of any system: without selection, coverage is worthless, since the reference ruleset’s own selectivity is exactly what the metric rewards. A frontier model raises the ceiling but is not required for the precision. Claude Code pairs Claude Opus 5 with the same kind of general-purpose scaffold that stalled around FAS-F1 0.40 on gpt-oss-120b for OpenClaw and Hermes (§6.4.2). On the stronger model, that same scaffold class reaches 0.623, the best of any system. That lead comes almost entirely from recall. It answers every capture at 0.967 recall, against RULE AUTO P ILOT’s 0.610. Precision tells a different story. Claude Code’s 0.536 and RULE AUTO P ILOT’s 0.523 are essentially tied, but they get there differently. Claude Code alerts on 10.99× the reference flow count to reach that precision, while RULE AUTO P ILOT alerts on only 1.17×. This is a different operating point, not a better one: ten alerts per real flow is the kind of ratio that drives alert fatigue. RULE AUTO P ILOT’s scaffold at the frontier backbone. We can hold the backbone fixed and swap scaffolds directly, pairing RULE AUTO P ILOT’s own scaffold with the same Claude Opus 5 backbone that powers Claude Code. At equal back-

Our Scaffold vs. Other Scaffolds, Same Backbone

Rule quality could come from RULE AUTO P ILOT’s taskspecific design, or simply from wrapping any capable agent around the same backbone model. To tell these apart, we hold the backbone fixed at gpt-oss-120b and compare RULE AU TO P ILOT against two general-purpose coding agents, OpenClaw and Hermes Agent, given the same task and prompt. Table 6 reports all six systems. The other three, Claude Code, the naive regex extractor, and RULE AUTO P ILOT on the Claude Opus 5 backbone, are analyzed in §6.4.3. This section isolates the three systems sharing the gpt-oss-120b backbone. General-purpose scaffolds are interchangeable; a taskspecific one is not. OpenClaw and Hermes are different agents on the same model and prompt, and they land in the same place: FAS-F1 0.416 versus 0.400, with recall identical at 0.575. Both write rules that stay generic. Neither uses bsize, which pins a matched field to an exact byte length. Neither attaches fast_pattern, which flags the most distinctive byte string in a rule, to more than a fifth of their rules. Both keywords are available to every system we evaluate. Their absence reflects a failure to identify distinguishing patterns, not a lack of access. RULE AUTO P ILOT, at the same backbone, reaches 0.524 FAS-F1 with fast_pattern on 48% of rules. That is ahead of the two agents by 0.109 and 0.125. RULE AUTO P ILOT also emits half as many rules per capture. It alerts on 1.17× the reference flow count, against 2.1–2.5× for OpenClaw and Hermes. Abstention and cost. RULE AUTO P ILOT emits no rules for 42 of 200 captures, 34 of them because the synthesis agent classified every flow as background. When RULE AUTO P ILOT abstains, that capture contributes zero precision and zero recall to the average, since no rule was emitted to be scored. This pulls down RULE AUTO P ILOT’s system-wide precision and recall relative to what it achieves when it does commit to an answer. Averaged across all 200 captures, precision is 0.523. Restricted to the captures where RULE AUTO P ILOT emits at 10

Mean score across PCAPs

(i) Syntax Pass Rate

(ii) Trigger Rate

(iii) FAS-F1

90%

90%

(iv) Benign FPR (%)

0.70

2.00%

0.65

1.50%

85% 85%

80%

0.60

75%

80%

70% 75%

1.00%

0.55

0.50%

65% 0.50

60%

e

Non

ax

Synt

) (Full lware nign +Be

+Ma

e

Non

ax

Synt

) (Full lware nign +Be

0.00%

e

Non

+Ma

ax

Synt

) (Full lware nign +Be

e

Non

+Ma

ax

Synt

) (Full lware nign +Be

+Ma

Verification Levels GPT-OSS With Benign Filtering

GPT-5.5 With Benign Filtering

GPT-OSS No Benign Filtering

GPT-5.5 No Benign Filtering

Figure 5: Cumulative verification progression for GPT-5.5 and GPT-OSS-120b under few-shot-5 prompting, with and without benign-flow filtering. The x-axis shows enabled stages: None uses raw rules; Syntax adds Suricata syntax validation and repair; +Malware adds replay on the source malware PCAP; and +Benign (Full) adds benign replay to suppress false positives. Table 6: Cross-system comparison on 200 malware PCAPs. Cover is the fraction of captures for which the system emits at least one rule. Over-fire is alerted flows divided by ground-truth flows. System (backbone)

Cover

FAS-F1

FAS-P

FAS-R

Rules/PCAP

Over-fire

Tokens/capture

RULE AUTO P ILOT (Claude Opus 5) Claude Code (Claude Opus 5) RULE AUTO P ILOT (gpt-oss-120b) OpenClaw (gpt-oss-120b) Hermes Agent (gpt-oss-120b) Regex, naive (none)

95% 100% 79% 100% 98% 100%

0.656 0.623 0.524 0.416 0.400 0.195

0.618 0.536 0.523 0.379 0.355 0.130

0.830 0.967 0.610 0.575 0.575 0.809

3 3 2 4 4 31

1.49× 10.99× 1.17× 2.07× 2.50× 8.18×

30,910 1,245,898 23,836 885,734 1,192,519 0

bone, RULE AUTO P ILOT reaches FAS-F1 0.656, ahead of Claude Code’s 0.623, with FAS-precision 0.618 against 0.536 and an over-fire ratio of 1.49× against 10.99×. RULE AU TO P ILOT trades some recall for this gain, 0.830 versus 0.967, and answers 95% of captures against Claude Code’s 100%. Our scaffold declines rather than emit a weakly supported rule. Because the backbone is now identical, this result isolates the scaffold as the source of the gap. A task-specific scaffold does not just close the distance to a frontier general-purpose agent; it surpasses it, at a 40× reduction in tokens per capture. Why the cost gap matters. RULE AUTO P ILOT runs on an open-weight backbone that an organization with sufficient compute can self-host, avoiding both per-token billing and the need to send network traffic to a third-party API. Claude Code’s cost and privacy profile are tied to a closed, APIserved model. The token gap above is therefore not just a cost difference; it reflects whether sensitive traffic ever has to leave the organization.

7

TO P ILOT can still do with what remains, and how well the

rules it writes hold up. Performance across the encryption spectrum. We classify each capture by the share of its flow records that Suricata dissects as tls or quic, the two encrypted protocols with enough volume in this corpus to support separate analysis. This classification is based on protocol dissection rather than which port traffic uses, so encrypted traffic on a non-standard port is still correctly identified as encrypted. Based on this share, we sort captures into three bands: cleartext (mostly unencrypted, 331 captures), mixed (a blend of encrypted and unencrypted traffic, 889 captures), and encrypted (mostly encrypted, 76 captures). The bands differ mainly in how much traffic they hide, not in which protocols show up: nearly all mixed and encrypted captures still contain both TLS and HTTP, and every capture in the corpus contains DNS. Table 7 reports RULE AUTO P ILOT’s performance in each band. As encryption increases, FAS-F1 drops from 0.614 on cleartext to 0.355 on encrypted, and precision drops faster than recall. In other words, rules written under encryption still catch a similar share of the ground truth, but they also raise more false alerts along the way. Coverage moves in the opposite direction, rising from 80.4% to 88.2%, and the rules the pipeline does write fire more reliably as encryption increases. RULE AUTO P ILOT does not go quiet as

RQ3: Rule Generation Under Encryption

This section asks how RULE AUTO P ILOT performs once traffic is encrypted. Encryption hides the payload, but leaves some handshake fields visible. We measure what RULE AU 11

Table 7: RULE AUTO P ILOT performance by encryption band across the full 1,296-PCAP corpus (few-shot-5, full verification). Bands are the share of Suricata-dissected flow records carrying tls or quic. Cover is the fraction of captures for which at least one rule survives verification, and Trig. is the trigger rate over those captures. Band Cleartext Mixed Encrypted All

PCAPs

Cover

FAS-P

FAS-R

FAS-F1

Trig.

331 889 76

80.4% 86.8% 88.2%

0.603 0.485 0.349

0.677 0.701 0.452

0.614 0.526 0.355

91.1% 96.3% 96.9%

1,296

85.3%

0.507

0.680

0.539

95.1%

without the anchor they currently key on. In that setting, rule generation would need to lean more heavily on stateful rules that read overall network behavior, such as connection timing and volume, rather than any single field in the traffic. Where FAS goes to zero, and why. Of 1,296 PCAPs, 306 score FAS-F1 of zero: 211 produce no valid rule and 95 produce valid rules that miss every security-relevant flow. Two structural causes account for them. (i) Externally attributed indicators. Signatures keyed on JA3/JA3S fingerprints expose no interpretable pattern in the traffic. Without a live intelligence feed, no generator can decide whether an unseen fingerprint is malicious. The limitation is missing attribution, not rule generation. (ii) Over-generation. Of 1,296 PCAPs, a further 207 PCAPs land in the partial-alignment band (0 < FAS-F1 < 0.5), at mean precision 0.28 and recall 0.77. Rather than one rule aligned with the target, the model emits several for different protocol signals in the same capture. Recall survives when one of these rules matches, but precision suffers when the others hit background flows. What FAS does not measure. FAS scores a generated rule by its flow-level agreement with the ET Open rule that fired on the same capture, treating that rule as ground truth rather than as one valid detection among others. A generated rule that matches on a different but equally valid indicator scores as a false positive, and one that would detect the family more robustly than the expert rule cannot score above it. Agreement alone also cannot separate a rule that fires twice too often from one that fires fifty times too often, which is why we report over-fire alongside FAS. Whether agreement with a curated ruleset tracks operational detection quality is a question this corpus cannot answer without alert data from a live deployment.

payload disappears; it answers more often and more narrowly, at some cost to precision. What RULE AUTO P ILOT writes when the payload is encrypted. Of the 40 samples whose ground truth requires a content match on a fully encrypted trace, listed as Content, encrypted in Table 1, RULE AUTO P ILOT answers 34 and writes 89 rules. Of those, 26 match on tls.sni, the name the session announces in its handshake, and 38 on dns.query, the name it resolves beforehand. Stateful rules are built on the same names: 13 of the 27 written on these samples combine their threshold with a tls.sni match, against 5 of 28 on the cleartext samples. Both forms of detection therefore remain available once the payload is gone, content rules written against the handshake and DNS fields that stay visible, and stateful rules written against behavior that never needed the payload at all.

8

Discussion

Scaffolding is what makes local deployment viable. The choice to run an open-weight model in-house is not just a cost decision. It is often the only option when the traffic itself is the constraint. Malware captures, and the network layout they reveal, can be sensitive operational information an organization cannot send to a third-party API regardless of that API’s quality. Under that constraint, the question this paper answers changes from “which model is best” to “how much of a frontier model’s advantage can a well-designed scaffold recover on a model you are allowed to run.” On the open-weight backbone, our scaffold comes within 0.099 F1 of the frontier system at a fraction of the cost. At equal backbone, it surpasses that system outright, showing the scaffold is not just recovering a frontier model’s advantage but is itself the larger source of it. What happens when more fields are encrypted. Every result in this section rests on a name: the server name in the TLS handshake or the DNS query that precedes it. Wider adoption of Encrypted Client Hello [43] and DNS over HTTPS or TLS [20, 21] would remove both from what a sensor can observe, leaving RULE AUTO P ILOT’s content and rate rules

9

Related Work

We organize related work into three categories: traditional automated signature generation, LLM-based rule generation systems, and LLMs for network traffic analysis.

9.1

Traditional Signature Generation

Honeycomb [25] extracts signatures from honeypot traffic by longest common substring, Autograph [24] clusters gateway payloads for worm signatures, and Polygraph [33] resists polymorphism through conjunctions of invariant substrings. These showed that malicious traffic carries reusable invariants, but they emit byte-level signatures rather than Suricata rules with protocol-aware keywords, flow constraints, and metadata; more recent deep-learning IDS work [53] outputs labels or anomaly scores rather than rule artifacts. Our non-LLM baselines are a modern analogue: they extract every observable protocol-field indicator and emit a rule for each. That they stay competitive on cleartext and degrade sharply under 12

9.3

encryption (§7) suggests the historical difficulty was invariant extraction, whereas under encryption it is invariant selection.

9.2

LLMs for Network Traffic Analysis

Recent work applies LLM-style models to network traffic understanding. TrafficLLM [14] adapts open-source LLMs to raw traffic data, while NetGPT [30] pretrains generative models over packet- and flow-level representations. These systems show that language-model-style architectures can learn useful traffic representations. However, their goals are traffic analysis and synthesis, not deployable IDS rule generation. RULE AUTO P ILOT focuses specifically on transforming raw network PCAPs into validated Suricata signatures.

LLM-Based Rule Generation Systems

RuleMaster+ [27] fine-tunes an LLM to generate IDS rules from proof-of-concept exploit descriptions, showing that LLMs learn rule-writing conventions but assuming curated exploit text as input. FALCON [31] generates Snort and YARA rules from CTI reports using an agentic workflow with explicit validation. Both start from structured intelligence rather than traffic, so neither addresses the challenge that dominates our setting: identifying the security-relevant flows inside a noisy capture before any rule can be written. Moreno et al. [32] are the closest PCAP-based prior work, generating Suricata rules from captures of controlled industrial attack scenarios under zero-shot, few-shot, and chain-ofthought prompting. Their setting isolates anomalous flows by attacker IP address, so the LLM receives traffic already filtered toward the attack, whereas RULE AUTO P ILOT starts from noisy sandbox PCAPs where identifying the relevant flows is part of the task. Their repair loop is parser-error driven; ours adds execution feedback from malware replay and benign testing, which is worth +0.083 FAS-F1 over syntax-only verification in our ablation (0.456 to 0.539). GRIDAI [26] is a multi-agent framework for ruleset evolution, deciding whether an incoming Web-attack sample is a new type or a variant already covered and then generating or repairing accordingly, evaluated on HTTP attack samples. RULE AUTO P ILOT instead targets discovery from noisy multiprotocol sandbox traces, where the relevant flows must be identified before any rule is written. Other recent systems. RuleXploit [40], Hex2Sign [10], GenTI [19], and [29] generate rules from exploit descriptions, hexadecimal payloads, and unseen-attack benchmarks, scoring with BERTScore or an LLM judge in place of execution. The distinction that matters here is whether a rule is ever run: model-judged similarity to a reference does not establish that a rule loads, fires on its source traffic, or stays quiet on benign traffic. Consistent with this, [16] finds practitioners treat execution evidence as a precondition for adoption. General-purpose agents. Agents that operate over a filesystem and shell appear in our evaluation as baselines rather than related systems. They are not built for rule generation and receive our prompt rather than their own. Their role is methodological, establishing what a capable model does with the same instructions and evidence but no task-specific scaffold. To our knowledge no prior work on IDS rule generation holds the backbone fixed while varying the surrounding system, so the scaffold’s contribution has not previously been separated from the model’s.

10

Conclusion

We presented RULE AUTO P ILOT, an agentic framework that synthesizes deployable Suricata rules directly from network traffic. We used it to ask what determines rule quality: the backbone model or the scaffold around it. Evaluated across 1,296 malware PCAPs, execution-grounded verification and benign traffic fingerprinting together raise rule quality while cutting both cost and false positives, showing the pipeline scales to a large, real-world corpus. Under encryption, precision falls as payload disappears, and generation leans more on stateful rules and header-level fields such as the TLS SNI and the DNS query that precedes a connection. Encryption has not eliminated the need for content-based rules, though: ET Open itself continues to publish plaintext-targeting signatures today, a sign that a meaningful share of malware still operates over unencrypted traffic. The central finding is about the scaffold, not the backbone. An open, self-hosted model paired with a task-specific scaffold approaches frontier-model quality in generating deployable Suricata rules directly from network traffic, without sending sensitive traffic to a third-party API. This finding demonstrates that high-quality rule generation is possible without relying on a closed model. Using the same backbone, Claude Opus 5, RULE AUTO P ILOT outperforms Claude Code while consuming 40× fewer tokens per capture. Together, these results show that RULE AUTO P ILOT supports both private, self-hosted deployment and improved detection quality with frontier models, while substantially reducing token consumption and associated costs. Future Work. One natural extension is integrating live threat intelligence to cover the one class of rule our pipeline cannot write on its own: indicators with no interpretable pattern in the traffic itself, such as a JA3, JA3S, or JA4 TLS fingerprint. A fingerprint is just a hash, so no amount of reasoning over the capture can tell us whether an unseen one is malicious. Pairing the synthesis agent with an external feed, such as a fingerprint or hash reputation database built from past incidents, would supply that missing evidence. The model would then focus on deciding whether a matched fingerprint, combined with the rest of the capture’s behavior, is enough reason to write a rule. 13

References

[14] Tianyu Cui, Xinjie Lin, Sijia Li, Miao Chen, Qilei Yin, Qi Li, and Ke Xu. TrafficLLM: Enhancing large language models for network traffic analysis with generic traffic representation. CoRR, abs/2504.04222, 2025. URL: https://arxiv.org/abs/2504.04222, doi:10.48550/arXiv.2504.04222.

[1] Emerging threats ruleset updates. https://communit y.emergingthreats.net/c/ruleset-updates/9. Accessed: 2026-04-26. [2] Feodo tracker: Suricata botnet c2 ip ruleset. https: //feodotracker.abuse.ch/blocklist/. Accessed: 2026-04-26.

CVE-2025-55182 (Re[15] Datadog Security Labs. act2Shell): Remote code execution in React server components and Next.js, December 2025. URL: https: //securitylabs.datadoghq.com/articles/cve-2 025-55182-react2shell-remote-code-executi on-react-server-components/. Accessed 2026-0820.

[3] Suricata user guide: Rules format. https://docs .suricata.io/en/latest/rules/intro.html. Accessed: 2026-04-26. [4] Urlhaus ids ruleset. https://urlhaus.abuse.ch/d ownloads/suricata-ids/. Accessed: 2026-04-26.

[16] Lorenzo di Filippo, Enkeleda Bardhi, Andrea Agiollo, Alessandro Palma, Silvia Bonomi, and Fernando Kuipers. Beyond the syntax: Do security experts trust LLMs for NIDS rule engineering? arXiv preprint arXiv:2607.05916, 2026.

[5] abuse.ch. MalwareBazaar: A resource for sharing malware samples. https://bazaar.abuse.ch/, 2024. [6] Eugene Albin and Neil C. Rowe. A realistic experimental comparison of the Suricata and Snort intrusion-detection systems. In 2012 26th International Conference on Advanced Information Networking and Applications Workshops, pages 122–127, 2012. doi:10.1109/WAINA.2012.29.

[17] Jesús Díaz-Verdejo, Javier Muñoz-Calle, Antonio Estepa Alonso, Rafael Estepa Alonso, and Germán Madinabeitia. On the detection capabilities of signaturebased intrusion detection systems in the context of web attacks. Applied Sciences, 12(2):852, 2022. doi:10.3390/app12020852.

[7] Anthropic. Claude Code. Software, CLI version 2.1.240, 2026. Accessed 2026-08-23. Backbone model claude-opus-5.

[18] Emerging Threats Community. Handling false positive reports as a rule writer. https://community.emergi ngthreats.net/t/handling-false-positive-r eports-as-a-rule-writer-special-guests-p cres-dalton-dalton-s-flowsynth/1031, 2023. Accessed: 2026-04-30.

[8] ANY.RUN. ANY.RUN: Interactive online malware sandbox. https://any.run/, 2024. [9] AV-TEST Institute. Malware statistics and trends report. https://www.av-test.org/en/statistics/malw are/, 2024. Registers over 450,000 new malware and PUA daily.

[19] Hassan Jalil Hadi, Rehana Yasmin, and Ali Shoker. GenTI: Benchmarking LLMs for autonomous IDPS rule generation for unseen attacks. arXiv preprint arXiv:2606.05844, 2026.

[10] Priyanka Balasubramanian, Tarek Ali, Mohammad Salmani, Danial KhoshKholgh, and Panos Kostakos. Hex2Sign: Automatic IDS signature generation from hexadecimal data using LLMs. In Proc. IEEE Int. Conf. on Big Data (BigData), pages 4524–4532, 2024. URL: https://ieeexplore.ieee.org/document/10825 710/.

[20] Paul Hoffman and Patrick McManus. DNS queries over HTTPS (DoH). Technical Report RFC 8484, IETF, 2018. doi:10.17487/RFC8484. [21] Zi Hu, Liang Zhu, John Heidemann, Allison Mankin, Duane Wessels, and Paul Hoffman. Specification for DNS over Transport Layer Security (TLS). Technical Report RFC 7858, IETF, 2016. doi:10.17487/RFC7858.

[11] CERT-EE. Cuckoo Sandbox. https://sandbox.pi kker.ee/. Public hosted Cuckoo Sandbox instance; accessed 2026-04-30.

[22] George-Andrei Iosifescu. DikeDataset: Labeled benign and malicious PE and OLE files. https://github.com /iosifache/DikeDataset, 2021. Dataset with labeled benign and malicious PE and OLE files; accessed 202604-30.

[12] Cisco. Snort: Network intrusion detection and prevention system. https://www.snort.org/, 2024. [13] Carlos Garcia Cordero, Emmanouil Vasilomanolakis, Aidmar Wainakh, Max Mühlhäuser, and Simin NadjmTehrani. On generating network traffic datasets with synthetic attacks for intrusion detection. ACM Transactions on Privacy and Security, 24(2):1–39, 2021. doi:10.1145/3424155.

[23] Robert J. Joyce, Derek Everett, Maya Fuchs, Edward Raff, and James Holt. ClarAVy: A tool for scalable and accurate malware family labeling. CoRR, 14

abs/2502.02759, 2025. URL: https://arxiv.org/ab s/2502.02759, doi:10.48550/arXiv.2502.02759.

[32] Manez Moreno, Xabier Sáez-de Cámara, Aitor Urbieta, and Mikel Iturbe. Leveraging LLMs for automated IDS rule generation: A novel methodology for securing industrial environments. In Proceedings of X Jornadas Nacionales de Investigación en Ciberseguridad (JNIC 2025), pages 113–120, Zaragoza, Spain, June 2025. URL: https://iturbe.info/assets/pdf /moreno2025leveraging.pdf.

[24] Hyang-Ah Kim and Brad Karp. Autograph: Toward automated, distributed worm signature detection. In 13th USENIX Security Symposium (USENIX Security 04), pages 271–286, San Diego, CA, August 2004. USENIX Association. URL: https://www.usenix.org/confe rence/13th-usenix-security-symposium/autog raph-toward-automated-distributed-worm-sig nature.

[33] James Newsome, Brad Karp, and Dawn Song. Polygraph: Automatically generating signatures for polymorphic worms. In Proceedings of the 2005 IEEE Symposium on Security and Privacy, pages 226–241. IEEE Computer Society, 2005. doi:10.1109/SP.2005.15.

[25] Christian Kreibich and Jon Crowcroft. Honeycomb: Creating intrusion detection signatures using honeypots. ACM SIGCOMM Computer Communication Review, 34(1):51–56, 2004. doi:10.1145/972374.972384.

[34] Nous Research. Hermes Agent. Software, version 0.20.1 (2026.8.13), 2026. Accessed 2026-08-23.

[26] Jiarui Li, Yuhan Chai, Lei Du, Chenyun Duan, Hao Yan, and Zhaoquan Gu. GRIDAI: Generating and repairing intrusion detection rules via collaboration among multiple LLM-based agents. CoRR, abs/2510.13257, 2025. URL: https://arxiv.org/abs/2510.13257, doi:10.48550/arXiv.2510.13257.

[35] Open Information Security Foundation. Suricata user guide: Dns keywords. https://docs.suricata.io /en/latest/rules/dns-keywords.html. Accessed: 2026-04-30. [36] Open Information Security Foundation. Suricata user guide: Payload keywords. https://docs.suricat a.io/en/latest/rules/payload-keywords.html. Accessed: 2026-04-30.

[27] Wenjuan Lian, Chengxin Zhang, Hongbao Zhang, Bin Jia, and Baihang Liu. RuleMaster+: LLM-based automated rule generation framework for intrusion detection systems. Chinese Journal of Electronics, 34(5):1402– 1415, 2025. doi:10.23919/cje.2024.00.342.

[37] Open Information Security Foundation. Suricata user guide: Rules format. https://docs.suricata.io/e n/latest/rules/intro.html. Accessed: 2026-0430.

[28] Keane Lucas, Mahmood Sharif, Lujo Bauer, Michael K. Reiter, and Saurabh Shintre. Malware makeover: Breaking ML-based static analysis by modifying executable bytes. In Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security, ASIA CCS ’21, 2021. doi:10.1145/3433210.3453086.

[38] Open Information Security Foundation (OISF). Suricata: Open source IDS/IPS/NSM engine. https://surica ta.io/, 2024. [39] OpenClaw. OpenClaw: An open-source coding agent. Software, version 2026.7.1-2 (commit 0790d9f), 2026. Accessed 2026-08-23.

[29] Cheng Meng, Wenxin Le, Xinyi Li, Qiuyun Wang, Fangli Ren, Zhengwei Jiang, and Baoxu Liu. From context to rules: Toward unified detection rule generation. arXiv preprint arXiv:2604.11078, 2026. Substitutes an LLM judge for execution-based evaluation; see §3.2 and §6.3.

[40] Angelos Papoutsis, Athanasios Dimitriadis, Ilias Koritsas, Dimitrios Kavallieros, Theodora Tsikrika, Stefanos Vrochidis, and Ioannis Kompatsiaris. RuleXploit: A framework for generating Suricata rules from exploits using generative AI. In Proc. IEEE Int. Conf. on Cyber Security and Resilience (CSR), 2025. URL: https: //ieeexplore.ieee.org/document/11130010/.

[30] Xuying Meng, Chungang Lin, Yequan Wang, and Yujun Zhang. NetGPT: Generative pretrained transformer for network traffic. CoRR, abs/2304.09513, 2023. URL: https://arxiv.org/abs/2304.09513, doi:10.48550/arXiv.2304.09513.

[41] Proofpoint. Emerging Threats Open Ruleset (ET Open). https://rules.emergingthreats.net/open/, 2024.

[31] Shaswata Mitra, Azim Bazarov, Martin Duclos, Sudip Mittal, Aritran Piplai, Md Rayhanur Rahman, Edward Zieglar, and Shahram Rahimi. FALCON: Autonomous cyber threat intelligence mining with LLMs for IDS rule generation. CoRR, abs/2508.18684, 2025. URL: https://arxiv.org/abs/2508.18684, doi:10.48550/arXiv.2508.18684.

[42] Recorded Future. Triage: Malware analysis sandbox. https://tria.ge/, 2024. [43] Eric Rescorla, Kazuho Oku, Nick Sullivan, and Christopher A. Wood. TLS encrypted client hello. InternetDraft draft-ietf-tls-esni, IETF, 2024. 15

Ethical Considerations

[44] Shrit Shah and Fatemeh Khoda Parast. AI-driven cyber threat intelligence automation. CoRR, abs/2410.20287, 2024. URL: https://arxiv.org/abs/2410.20287, doi:10.48550/arXiv.2410.20287.

Malware samples. Malware artifacts used in this study were sourced from MalwareBazaar [5], ANY.RUN, and Triage. Where binaries were collected (MalwareBazaar and ANY.RUN), samples were executed solely within isolated Cuckoo Sandbox [11] environments with no network connectivity to production systems; Triage-provided PCAPs were ingested directly as traffic traces. No malware binary files were distributed, deployed, or tested outside the sandboxed environment. Responsible use. While RULE AUTO P ILOT could in principle reveal information about malware detection capabilities (e.g., which network patterns are being monitored), the generated rules are standard Suricata signatures comparable to those already publicly available in the ET Open ruleset. We do not believe this work introduces novel dual-use risks beyond those inherent to existing public rulesets.

[45] Teodor Sommestad, Hannes Holm, and Daniel Steinvall. Variables influencing the effectiveness of signaturebased network intrusion detection systems. Information Security Journal: A Global Perspective, 31(6):711–728, 2022. doi:10.1080/19393555.2021.1975853. [46] Yakuta Tayyebi and D. S. Bhilare. A comparative study of open source network based intrusion detection systems. International Journal of Computer Science and Information Technologies, 9(2):23–26, 2018. URL: https://www.ijcsit.com/docs/Volume%209/vol 9issue2/ijcsit2018090201.pdf. [47] Koen T. W. Teuwen, Tom Mulders, Emmanuele Zambon, and Luca Allodi. Ruling the unruly: Designing effective, low-noise network intrusion detection rules for security operations centers. In Proceedings of the 20th ACM Asia Conference on Computer and Communications Security, ASIA CCS ’25, pages 1428–1441, 2025. doi:10.1145/3708821.3710823.

Open Science 1. Source code: https://anonymous.4open.science/ r/rulepilot-artifact-DB05. Includes the RULE AU TO P ILOT pipeline, evaluation scripts, and the baseline harness used for the cross-system.

[48] Unit 42. Exploitation of critical vulnerability in React server components (CVE-2025-55182). Palo Alto Networks Unit 42 Threat Brief, December 2025. Accessed 2026-08-20.

2. Dataset. We do not release the malware corpus or binaries for the reasons described in the Ethical Considerations section.

[49] VirusTotal. VirusTotal. https://www.virustotal.c om/, 2024. [50] VulnCheck. Critical vulnerability in React and Next.js (CVE-2025-55182), December 2025. URL: https: //www.vulncheck.com/blog/cve-2025-55182-r eact-nextjs. Accessed 2026-08-20. Ships a proof-ofconcept packet capture with Suricata and Snort rules.

A

Dataset and Ground-Truth Details

This appendix provides additional details on the malware corpus, the ET Open rules used to construct ground-truth labels, and the normalized malware-family distribution used in our evaluation.

[51] Abdul Waleed, Abdul Fareed Jamali, and Ammar Masood. Which open-source IDS? Snort, Suricata or Zeek. Computer Networks, 213:109116, 2022. doi:10.1016/j.comnet.2022.109116.

A.1

Temporal Coverage of Malware PCAPs

[52] Joshua S. White, Thomas Fitzsimmons, and Jeanna N. Matthews. Quantitative analysis of intrusion detection systems: Snort and suricata. In Cyber Sensing 2013, volume 8757 of Proceedings of SPIE, page 875704, 2013. doi:10.1117/12.2015616.

We report the temporal coverage of the malware corpus to show that the benchmark includes both older and recently submitted samples. Table 8 groups malware PCAPs by VirusTotal first_submission_date.

[53] Zhiwei Xu, Yujuan Wu, Shiheng Wang, Jiabao Gao, Tian Qiu, Ziqi Wang, Hai Wan, and Xibin Zhao. Deep learning-based intrusion detection systems: A survey. CoRR, abs/2504.07839, 2025. URL: h t t p s : / / a r x i v . o r g / a b s / 2 5 0 4 . 0 7 8 39, doi:10.48550/arXiv.2504.07839.

A.2 Temporal Coverage of Triggered ET Open Rules We also examine the age of the ground-truth ET Open signatures triggered by the malware corpus. Table 9 groups unique triggered ET Open rules by their created_at metadata field. 16

Table 8: Temporal coverage of malware PCAPs. Malware PCAPs are grouped by VirusTotal first_submission_date. Period

Malware PCAPs

2014–2015 2016–2017 2018–2019 2020–2021 2022–2023 2024–2025 2026 Total

SMB, plus benign boundary examples carrying no rule; N denotes malicious examples per protocol. Source PCAPs of the examples are excluded from evaluation. Templates are in Appendix D. Benign corpus. The 981 DIKE-derived benign PCAPs build the fingerprint library; a 600-PCAP subset serves benign replay during verification and a disjoint 200-PCAP subset final false-positive testing. An external 191-PCAP benchmark [28], used for none of these, evaluates benign-only abstention. Comparison systems. For RQ2 we run four additional systems over the 200-capture subset of §5. Two are generalpurpose coding agents driving the same gpt-oss-120b backbone as RULE AUTO P ILOT: OpenClaw and Hermes Agent. One is Claude Opus 5 under Claude Code. One is a non-LLM field extractor, 270 lines of Python over tshark output, emitting one rule per indicator observed anywhere in the capture with no benign filtering or model judgment. The three baseline agents receive a byte-identical prompt, verified to one MD5 across all 603 sandbox copies (Appendix E), the same task description, and a script that replays Suricata against its own capture only, in its own sandbox tree with the capture hardlinked across trees so all five systems provably read the same bytes. That prompt states the task, the evidence constraint, and how rules are scored, but it is not RULE AU TO P ILOT ’s own pipeline prompt, which cannot be shared: the pipeline prompt carries 31 few-shot examples quoting real ET Open rules, so handing it to a baseline would leak the ground truth. §6.3 bounds the residual asymmetry, where moving from zero-shot to few-shot-5 is worth 0.018 FAS-F1 against the 0.109–0.125 gaps measured here. The design isolates one variable at a time: OpenClaw versus Hermes swaps one general-purpose scaffold for another at a fixed model and prompt, the open-model agents versus Claude Code vary the model, and RULE AUTO P ILOT versus OpenClaw and Hermes replaces a general-purpose scaffold with a task-specific one at a fixed model. Uniform benign false-positive control. RULE AUTO P ILOT discards every candidate rule that fires on a 600-PCAP benign corpus, and the baselines have no such stage, which would confound any precision comparison. We therefore apply the identical stage to every system after the fact, replaying each system’s emitted rules against the same 600 benign PCAPs and re-scoring what survives. All RQ2 results are reported under this control, with the uncontrolled scores alongside because the difference is itself a result. Token accounting. All gpt-oss-120b systems are measured at the API level through a logging proxy with no prompt caching. Claude Code is measured from client transcripts and does use caching, so we report both tokens processed and a billed-equivalent figure charging cache reads at 10% of input rate. Comparisons needing no such correction, such as RULE AUTO P ILOT against OpenClaw on the same model, are noted where they occur. Suricata configuration. All rule validation uses

7 64 32 44 24 126 999 1,296

Table 9: Temporal coverage of unique triggered ET Open rules. Rules are grouped by the ET Open created_at metadata field. Period

A.3

Unique Triggered Rules

2010–2011 2012–2013 2014–2015 2016–2017 2018–2019 2020–2021 2022–2023 2024–2025 2026

21 22 28 26 53 42 72 336 17

Total

617

Malware Family Distribution

To characterize family diversity, we group normalized malware-family labels by the number of PCAPs assigned to each family. Table 10 summarizes this distribution and reports PCAPs without an assigned family label as SINGLETON instances.

B

Experimental Setup Details

Model. The primary backbone is gpt-oss-120b, served through an OpenAI-compatible institutional endpoint. We use an open model because network flows may carry sensitive operational information that cannot be sent to a third-party API, and because proprietary models become costly at corpus scale. We additionally evaluate GPT-5.5 on a stratified 100-PCAP subset (§6.4). All runs use temperature 0, a 65,536token completion budget, and up to three repair iterations per rule. Prompting strategies. We evaluate six strategies under identical generation parameters: zero-shot, chain-of-thought (CoT), few-shot with 1, 3, or 5 examples, and CoT+FS-3. Few-shot examples are (flow, rule) pairs drawn from real dataset PCAPs with their ET Open rules across HTTP, DNS, TLS, TCP, and 17

Table 10: Distribution of normalized malware-family sizes. PCAPs without an assigned family label are treated as SINGLETON instances and reported separately. Family names are grouped by the number of PCAPs assigned to each family. Samples / family

# Families

# PCAPs

Families

100+ 41–100 11–40 3–10

1 6 12 64

271 341 181 335

1–2

109

153

Subtotal: labeled families SINGLETON / no family

192 –

1,281 15

– –

Total

192

1,296

–

asyncrat vidar, remcos, sfuzuan, xworm, corewarrior, futu autoinject, wacatac, cryptnot, agenttesla, salat, quasar, blank, connectwise, lummastealer, virlock, gandcrab, stealc dorshel, formbook, salatstealer, minix, bookworm, heracles, shifu, lolbot, zorex, sagent, teslacrypt, disco, maskgramstealer, bumblebee, crypt0l0cker, birddog, netwire, socgholish, njrat, purelogs, locky, upatre, babar, taskun, atbot, wannacry, buterat, redline, nuitka, havoc, bladabindinet, shade, ravartar, blankgrabber, houdini, acsogenixx, lummac, dbadur, vnfraye, disfa, egairtigado, fragtor, metastealer, amadey, kasidet, fabookie, xmrig, mikey, cobaltstrike, nanocore, qwexlafiba, lodarat, trickbot, azorult, raccoon, venomrat, netsupport, nemesis, raccoonstealer, morphine, black, emotet, alevaul, wanker allaple, telebot, injuke, cryptowall, midie, injectornett, snakekeylogger, hlux, multiplug, znyonm, lethic, nemucod, ctblocker, warezov, cleanuploader, mydoom, darkcrystal, boxter, fugrafa, ligooc, strab, guloader, cerbu, oyster, gootkit, blind, hangover, jalapeno, zmutzy, obfdldr, hijackloader, socstealer, valleyrat, neutrino, lummac2, clipbanker, artifact, smokeloader, vobfus, ammyy, darkcomet, downeks, kelios, discordbot, windigo, denes, inoci, mars, amatera, discordrat, hiddentear, neshta, sdrop, risepro, backdoornjrat, brresmon, kepavll, blamon, makoob, lumma, cozyduke, barys, lockfile, privateloader, dridex, mofksys, crysan, tinynuke, sphinx, umbralstealer, bobik, stelpak, crypren, alvaro, aurastealer, spymax, mami, icedid, danabot, koobface, zegost, zombie, khalesi, dapato, medusalocker, obfsobjdat, bazloader, rhadamanthys, fatbeehive, sectoprat, nitol, kazuar, bladabindidldr, titanstealer, andromeda, banload, strelastealer, darkgate, cryxos, amnesia, polazert, expiro, lokibot, solmyr, ghostrat, multiverze, qqrob, hpqakbot, gcleaner

Suricata 8.0.3 on a local server equipped with dual AMD EPYC 7313 processors (32 physical cores, 64 logical CPUs, up to 3.7 GHz) and 1 TiB RAM.

C

Table 13 reports the effect of benign traffic fingerprinting on the 1,296 malware PCAPs used in the main evaluation. The filter removes most background flows while preserving nearly all security-relevant flows.

Flow Representation and Benign Fingerprinting

D

Prompt Templates

This appendix describes the flow representation given to the synthesis agent and the benign fingerprint library used to remove recurring background traffic before LLM processing.

This appendix provides the prompt templates used by the Rule Synthesis Agent. Each template is filled at runtime with flow context, direction hints, and the serialized flow records.

C.1

D.1

Protocol-Aware Flow Representation

The synthesis agent receives protocol-aware flow records rather than raw packets. Table 11 lists the per-protocol fields included in the flow JSON after removing fields without Suricata equivalents and remapping tshark names to Suricatacompatible keywords.

Zero-Shot Prompt

The zero-shot setting gives the model only the task instructions and flow records, without any worked examples. Figure 6 shows the zero-shot user prompt template. User Prompt -- Zero-Shot Total flows: {num_flows}

C.2

Benign Fingerprint Categories

{flow_review_context} Analyze the following network flows. Classify each flow and generate Suricata rules as described in the system instructions. Use only evidence present in the provided flow records.

The benign fingerprint library captures recurring identifiers observed in benign executions, such as DNS names, HTTP hosts, destination IPs, and local name resolution artifacts. Table 12 summarizes the fingerprint categories and representative filtering patterns.

C.3

{direction_hints} OUTPUT FORMAT: Return ONLY valid JSON -- no markdown, no prose. {"classification":[{"flow_name":"flow_N","tag":"imp"}], "rules":["alert dns ... (msg:\"Rule\"; sid:...; ...)"]} IMPORTANT: Only list flows you label "imp" in the classification array. Flows you omit are treated as "unimp" automatically.

Filtering Impact

Flows JSON: {flows_json}

We measure the effect of benign fingerprinting on the malware corpus before LLM processing. Table 13 reports how many background and security-relevant flows are removed, along with the reduction in flows passed to the LLM.

Figure 6: Zero-shot user prompt template tokens are filled at runtime.

18

Table 11: Per-protocol fields included in network flows. #

Protocol

Fields in flow JSON

1

HTTP

2

DNS

3 4

TLS TCP

5 6

SMB NBDGM

7

LLMNR

http.request.method|http.method, http.host, http.request.full_uri, http.request.line, http.request.version, http.user_agent, http.accept, http.connection, http.content_type, http.content_length, http.response.code|http.stat_code, http.response.line, http.server, http.file_data, http.request.headers parsed from request_line. dns.qry.name, dns.qry.type, dns.qry.class, dns.flags.response, dns.flags.rcode|dns.rcode, dns.count.queries, dns.count.answers, dns.count.auth_rr, dns.count.add_rr, dns.resp.name, dns.a, dns.cname. tls.handshake.extensions_server_name|tls.sni, tls.handshake.type, tls.record.length, tls.version. 5-tuple fields, tcp.flags, tcp.flags.str, tcp.window_size_value, tcp.options.mss_val, tcp.seq, tcp.ack, tcp.len, tcp.checksum, tcp.checksum.status, and tcp.payload.strings for printable ASCII runs ≥ 8 characters, domains, and URIlike patterns. nbss.type, smb.cmd, smb.tid, smb.uid, smb.mid, smb.nt_status, smb.path, smb.trans_name. nbdgm.src.ip, nbdgm.src.port, nbdgm.dgram_len, nbdgm.source_name, nbdgm.destination_name, nbdgm.type, smb.server_component, smb.cmd, smb.error_class, smb.trans_name, mailslot.opcode, mailslot.class, mailslot.name, browser.command, browser.period, browser.windows_version, browser.os_major, browser.server_type. dns.flags.response, dns.count.queries, dns.count.answers, dns.count.auth_rr, dns.count.add_rr, dns.qry.name, dns.qry.type, dns.qry.class, dns.retransmit_request, _ws.expert, _ws.expert.severity.

Note: All flows include the 5-tuple, transport protocol, and packet counts (packets_total, packets_src_to_dst, and packets_dst_to_src).

Table 12: Benign fingerprint categories and common filtering patterns. User Prompt -- Chain-of-Thought Category

Count

Filtering Examples

DNS names

159

HTTP hosts

37

Destination IPs

49

LLMNR names NBNS names NBDGM mailslots

27 57 5

SSDP fingerprints

N/A

Reverse-DNS lookups such as in-addr.arpa and ip6.arpa; routine Windows and system lookups. Common update and infrastructure hosts, including Microsoft, Apple, DigiCert, update, and time-service domains. Repeated DNS, NTP, and operating-system service endpoints. Local multicast name-resolution queries. Windows NetBIOS and file-sharing broadcasts. NetBIOS datagram destinations and mailslot names. UPnP and service-discovery multicast traffic.

D.2

Total flows: {num_flows} {flow_review_context} Analyze the following network flows using chain-of-thought reasoning. {direction_hints} STEP 1 -- CLASSIFICATION REASONING: For each flow, think step-by-step: a) What protocol/application layer is this? b) What specific indicators are present? c) Are indicators suspicious in isolation or only combined? d) Does traffic show bidirectional exchange or one-sided probe? e) Final classification: imp or unimp? Why?

Chain-of-Thought Prompt

STEP 2 -- RULE GENERATION: For each imp flow, use the mandatory rule headers from the system prompt, then add detection keywords for the behavioral pattern. Focus on generalizable patterns, not specific IOCs.

The chain-of-thought setting adds structured reasoning steps before the final JSON output, asking the model to identify protocol evidence, suspicious indicators, traffic direction, and rule anchors. Figure 7 shows the chain-of-thought prompt template.

OUTPUT FORMAT: Return ONLY valid JSON -- no markdown, no prose. {

D.3

"classification": [ { "flow_name": "flow_N", "tag": "imp", "reason": "one-sentence summary of key evidence" } ], "rules": [ "alert dns ... (msg:\"Rule\"; sid:...; ...)" ], "reasoning": "If zero rules, explain why."

Few-Shot Prompt

The few-shot setting prepends worked flow-to-rule examples before the task instructions. Figure 8 shows a representative example that pairs a remapped flow record with its corresponding ET Open rule. }

E

Baseline Harness Prompt

Flows JSON: {flows_json}

This appendix reproduces the prompt given to the three agentic baselines of §6.2. OpenClaw and Hermes Agent read it from AGENTS.md. Claude Code receives it via -append-system-prompt. All 603 sandbox copies hash to one MD5, so the baselines differ only in scaffold and back-

Figure 7: Chain-of-thought user prompt template.

19

Table 13: Benign traffic fingerprinting. Metric

User Prompt -- Few-Shot Example --- EXAMPLE (HTTP Tier 1) --Flow: { "app_protocol": "http", "src_ip": "192.168.168.206", "dest_ip": "178.32.110.193", "dest_port": 80, "packets_src_to_dst": 5, "packets_dst_to_src": 3, "data": { "http.method": "POST", "http.host": "prets-immobiliers.org", "http.uri": "/dbconnect.php", "http.user_agent": "Mozilla/5.0 (Windows NT 6.3; ...", "http.request_body": "data=EB7ED7C4C2C538751D5E01..." } }

Value

Flow composition (1,296 PCAPs) Total raw flows Security-relevant flows Background flows

131,787 6,937 (5.3%) 124,850 (94.7%)

After benign fingerprinting Security-relevant flows removed Background flows removed Flows passed to LLM Median flows/PCAP (raw) Median flows/PCAP (filtered)

142 (2.0% of security-relevant) 110,879 (88.8% of background) 20,766 (15.8% of total) 65 12

bone, not instructions. What is held fixed is the task, evidence constraint, rule budget, and scoring rule.

E.1

Task Description

Each sandbox also carries the following TASK.md, together with run_suricata.sh, which replays Suricata over that sandbox’s capture only, and empty rules/ and work/ directories. TASK.md

Classification: imp

# Task ‘capture.pcap‘ in this directory is a network capture from a single host. Write Suricata rules that detect the malicious/noteworthy activity in it, into ‘rules/generated.rules‘. Everything you need is in this directory. See the system prompt for the method, the constraints, and what counts as a good rule. ## What is here | path | what it is | |---|---| | ‘capture.pcap‘ | the capture -- your only evidence | | ‘run_suricata.sh‘ | ‘./run_suricata.sh <rules_file> [outdir]‘ -- test your rules against the capture | | ‘rules/‘ | write ‘generated.rules‘ here | | ‘work/‘ | scratch space; anything you like | ‘tshark‘ is available for inspecting the capture.

Rule: alert http $HOME_NET any -> $EXTERNAL_NET any (msg:"ET MALWARE Alphacrypt/TeslaCrypt Ransomware CnC Beacon"; flow:established,to_server; http.method; content:"POST"; http.uri; content:".php"; endswith; http.user_agent; content:"Mozilla"; startswith; http.request_body; content:"data="; startswith; fast_pattern; pcre:"/^[A-F0-9]{100,}$/R"; classtype:trojan-activity; sid:2022504; rev:1;)

F

--- [examples 2-5 follow the same format] ---

Comprehensive Error Patterns

This appendix lists the verifier failure patterns used by the Repair Agent. Failures are grouped into syntax errors, notrigger failures, and benign false positives.

Figure 8: Representative few-shot example, 1 of 5. Flow is shown post-remapping; rule is the matching ET Open signature.

F.1

Syntax-Error Patterns

Syntax-error patterns correspond to rules that Suricata cannot parse or load. Table 14 lists these errors, their root causes, and the associated repair strategies.

F.2

No-Trigger Patterns

No-trigger patterns correspond to rules that pass syntax validation but fail to alert on the source malware PCAP. Table 15 summarizes the main causes and repair strategies for these failures. 20

Table 14: Syntax-error patterns. #

Error Type

Root Cause

Fix Strategy

1

UNKNOWN_KEYWORD

2 3 4

MALFORMED_HEADER PORT_SYNTAX MISSING_REQUIRED_FIELD

Replace with the correct Suricata keyword, e.g., http.request.method → http.method. Use a valid Suricata protocol and the direction token ->. Use valid formats such as numeric ports, ranges, port lists, or any. Add ; after every rule option.

5 6

MISSING_SEMICOLON CONTENT_QUOTE_ERROR

Keyword does not exist, is misspelled, or a tshark field is used directly Invalid protocol or direction token in the rule header Invalid port format or reversed port tokens Missing semicolon causes the parser to consume the next keyword Missing or misplaced semicolon in rule options Unterminated quote or malformed content string

7

PCRE_ERROR

Invalid PCRE syntax or unsupported modifiers

8

EMPTY_STICKY_BUFFER

Sticky buffer appears without a following content: or pcre:

9

ENDSWITH_MISUSE

endswith; is used without preceding content:

10

APP_LAYER_MISMATCH

Rule header protocol conflicts with the sticky-buffer family

11 12 13

INVALID_SID INVALID_REV INVALID_FLOW

sid is missing, non-numeric, or zero-valued rev is missing or non-numeric flow: contains unsupported tokens

14 15 16 17 18 19 20

DUPLICATE_SID DUPLICATE_REV THRESHOLD_SYNTAX DETECTION_FILTER_SYNTAX BSIZE_TOO_LARGE TLS_VERSION_MISUSE FAST_PATTERN_NO_CONTENT

Rule contains multiple sid fields Rule contains multiple rev fields Malformed threshold: keyword Malformed detection_filter: keyword Content length exceeds declared bsize value tls.version uses invalid decimal or quoted format fast_pattern; appears without preceding content:

21

LEGACY_HTTP_KEYWORD

Old Suricata 2.x HTTP post-match modifiers are used

22 23 24 25

NOCASE_NO_CONTENT DNS_RCODE_MISUSE SSL_VERSION_INVALID DIRECTION_CONFLICT

nocase; appears without preceding content: or pcre: dns.rcode uses invalid string or malformed value Legacy SSL/TLS version field uses malformed value Rule mixes request-direction and response-direction buffers

F.3

False-Positive Patterns

ization, direction fixes, and sticky-buffer cleanup. Table 18 lists the rewrite rules applied before the LLM repair loop.

False-positive patterns correspond to rules that also trigger on benign traffic. Table 16 shows the narrowing strategy used for this failure class.

G

G.3

Repair Knowledge Base

H

Fixer Agent Feedback Templates

This appendix shows the prompts used when a generated rule fails verification. The Repair Agent uses one shared system prompt and three failure-specific feedback templates for syntax failures, no-trigger failures, and benign false positives.

Repair Strategy Coverage

Table 17 summarizes how the 42 error patterns are split across deterministic rewrites and LLM-guided repair.

G.2

LLM-Guided Repairs

LLM-guided repairs handle failures that require contextual reasoning over the failed rule, verifier output, and relevant malware or benign flows. Table 19 lists the error patterns handled through these context-aware repair prompts.

This appendix summarizes how the Repair Agent uses the error taxonomy from Appendix F. Some failures are corrected with deterministic rewrites before any LLM call, while the remaining failures use LLM-guided repair with verifier feedback and flow context.

G.1

Ensure every keyword option ends with ;. Use double quotes, escape internal quotes, or hex-escape special characters. Fix regex syntax, balance parentheses/classes, and avoid unsupported modifiers. Ensure every sticky buffer is followed by at least one match condition. Move the content match into the sticky buffer before applying endswith;. Align protocol header and buffer family, e.g., alert dns with DNS buffers. Set sid to a positive integer. Set rev to a positive integer. Use valid tokens such as established, to_server, or to_client. Preserve one valid sid and remove duplicates. Preserve one valid rev and remove duplicates. Remove the entire threshold:...; clause. Remove the entire detection_filter:...; clause. Remove bsize:N; or adjust it to a valid length. Use canonical constants such as tls12 or tls13. Remove the stray fast_pattern; keyword or place it after a content match. Convert to modern sticky-buffer syntax, e.g., http_uri → http.uri. Place nocase; immediately after the content or PCRE it modifies. Use numeric values or uppercase constants such as NXDOMAIN. Prefer tls.version with constants such as tls12. Separate into two rules or align buffers with to_server/to_client.

H.1

Fixer Agent System Prompt

The shared system prompt defines the Repair Agent’s role and constrains the output to a corrected Suricata rule string. Figure 9 shows the system prompt.

Deterministic Rewrites

Deterministic rewrites handle structural failures that can be corrected without LLM reasoning, such as keyword normal21

Table 15: No-trigger error patterns. #

Error Type

1

Align C2 traffic as $HOME_NET → $EXTERNAL_NET with to_server. DIRECTION_MISMATCH_EXT_SRC $EXTERNAL_NET is source while flow uses to_server Swap network variables so the client side is the source. DIRECTION_MISMATCH_INBOUND Outbound protocol is written as traffic going to $HOME_NET Reverse the direction so $HOME_NET is the source. PROTOCOL_BUFFER_MISMATCH Alert protocol does not match the observed application layer Use the correct alert protocol, e.g., HTTP → alert http. DUPLICATE_STICKY_BUFFER Same sticky buffer appears multiple times Merge all content or PCRE matches under one sticky-buffer instance. PCRE_R_ON_DNS_QUERY PCRE uses the relative /R flag with dns.query Remove the /R modifier from DNS query PCREs. GET_WITH_REQUEST_BODY Rule matches GET method with http.request_body Use http.uri for GET traffic or change the method to POST if appropriate. SMB_ASCII_ENCODING ASCII strings are used for SMB2 content matching Convert SMB strings to UTF-16LE hex encoding. SMB_SHARE_BUFFER smb.share buffer is empty for some tree-connect operations Use raw UTF-16LE hex content matching instead of smb.share. ENDSWITH_ON_HTTP_URI endswith; on http.uri fails when URI has query parame- Remove endswith; and match the stable URI substring. ters OVERLY_ANCHORED_DNS_PCRE DNS PCRE uses both start and end anchors too strictly Remove the leading anchor or match a stable domain suffix. DETECTION_FILTER_PRESENT detection_filter or threshold requires multiple Remove the clause and write a single-packet detection rule. matches OVER_SPECIFIC_CONTENT Content string is too long or sample-specific Replace with the shortest stable substring that appears across target flows. REQUEST_BODY_NO_METHOD http.request_body is used without checking HTTP Add http.method; content:"POST"; or remove body matchmethod ing for GET traffic. HTTP_RESPONSE_BUFFER_TO_ Response buffers are used with flow:to_server Flip to to_client or replace response buffers with request-side SERVER fields. TLS_CERT_SUBJECT_DIRECTION Certificate fields are matched with client-to-server direction Flip to to_client when matching certificate subject or issuer fields.

2 3 4 5 6 7 8 9 10 11 12 13 14 15 16

DIRECTION_MISMATCH

Root Cause

Fix Strategy

Rule direction does not match actual traffic direction

Table 16: False-positive error patterns #

Error Type

Root Cause

Fix Strategy

1

FALSE_POSITIVE

Detection pattern is too broad and also matches benign traffic

Tighten content with malware-specific anchors, add negated benign patterns, or narrow protocol/flow scope. Context-guided repair uses both benign and malware flow exemplars to narrow the pattern.

Fixer Agent System Prompt You are a Suricata rule expert. Fix the provided broken rule based on the error details. Preserve all existing metadata fields. If metadata is missing, add required ET metadata fields: affected_product, attack_target, created_at, deployment, signature_severity, tag, updated_at. Return ONLY the corrected rule string -- no markdown, no explanation.

Figure 9: Fixer-agent system prompt template.

22

Table 17: Error pattern distribution and repair strategy coverage. Category

Total

Deterministic

LLM-Guided

SYNTAX_ERROR

25

7

18

NO_TRIGGER

16

10

6

FALSE_POSITIVE

1

0

1

TOTAL

42

17

25

Primary Mode Keyword substitution, option-format normalization, delimiter repair, and metadata-preserving syntax fixes. Direction and buffer alignment, protocol correction, anchor relaxation, and context-grounded content revision. Context-guided narrowing using benign-triggering flows and malware flows that should remain matched. Hybrid deterministic and LLM-guided repair.

Table 18: Deterministic rewrite rules applied by the Repair Agent. #

Error Type

Root Cause

Rewrite Action

1

ENDSWITH_MISUSE

endswith or startswith is used incorrectly

2

DIRECTION_CONFLICT

Rewrite as valid Suricata syntax, e.g., content:"x"; endswith;. Fix the rule direction so the header and flow option agree.

8 9 10

Rule header direction conflicts with flow:to_server or flow:to_client DETECTION_FILTER_SYNTAX detection_filter clause is malformed BSIZE_TOO_LARGE Declared bsize value is inconsistent with matched content length FAST_PATTERN_NO_CONTENT fast_pattern appears without a preceding content match LEGACY_HTTP_KEYWORD Legacy Suricata HTTP modifiers are used instead of sticky buffers UNKNOWN_KEYWORD Wireshark or tshark field name is used as a Suricata keyword DIRECTION_MISMATCH_EXT_SRC $EXTERNAL_NET is source while flow uses to_server DIRECTION_MISMATCH_INBOUND Outbound protocol is written as traffic going to $HOME_NET DUPLICATE_STICKY_BUFFER Same sticky buffer appears multiple times

11 12 13

PCRE_R_ON_DNS_QUERY GET_WITH_REQUEST_BODY ENDSWITH_ON_HTTP_URI

14

OVERLY_ANCHORED_DNS_PCRE

15

DETECTION_FILTER_PRESENT

16

HTTP_RESPONSE_BUFFER_TO_ SERVER TLS_CERT_SUBJECT_DIRECTION

3 4 5 6 7

17

Strip the malformed detection_filter clause. Adjust or remove the invalid bsize option. Remove the orphaned fast_pattern keyword. Replace legacy modifiers with modern sticky-buffer syntax, e.g., http_uri → http.uri;. Substitute with the corresponding valid Suricata keyword.

Swap network variables so the client side is the source. Reverse the direction so $HOME_NET is the source. Merge all content or PCRE matches under one sticky-buffer instance. PCRE uses the relative /R flag with dns.query Remove the /R modifier from DNS query PCREs. Rule matches GET method with http.request_body Remove http.request_body matching for GET-only traffic. endswith; on http.uri fails when URI has query parame- Remove endswith; and match the stable URI substring. ters DNS or TLS PCRE uses strict anchors that prevent matching Remove overly strict anchors or match a stable domain suffix. observed traffic detection_filter requires multiple matches and prevents Remove the clause and write a single-packet detection rule. single-flow detection Response buffers are used with flow:to_server Flip to to_client or replace response buffers with request-side fields. Certificate fields are matched with client-to-server direction Flip to to_client when matching certificate subject or issuer fields.

23

Table 19: LLM-guided repair error patterns. #

Error Type

Root Cause

Fix Strategy

1

MALFORMED_HEADER

2

PORT_SYNTAX

Rule header contains an invalid protocol, direction token, or address/port structure Port value or port range is invalid or reversed

3 4

MISSING_SEMICOLON CONTENT_QUOTE_ERROR

5

PCRE_ERROR

6

EMPTY_STICKY_BUFFER

7

APP_LAYER_MISMATCH

8 9 10

INVALID_SID INVALID_REV INVALID_FLOW

11

DUPLICATE_SID

12 13

DUPLICATE_REV THRESHOLD_SYNTAX

14

OTHER_SYNTAX

15 16

NOCASE_NO_CONTENT DNS_RCODE_MISUSE

17

TLS_VERSION_MISUSE

18

SSL_VERSION_INVALID

19

DIRECTION_MISMATCH

20 21 22 23

PROTOCOL_BUFFER_MISMATCH SMB_ASCII_ENCODING SMB_SHARE_BUFFER OVER_SPECIFIC_CONTENT

24

REQUEST_BODY_NO_METHOD

25

FALSE_POSITIVE

Rewrite the header using a valid Suricata action, protocol, address, port, and direction. Replace with a valid port, range, or any when the observed flow does not justify a fixed port. Missing ; causes the parser to merge or lose later rule options Add missing semicolons between rule options. Content string has an unterminated or malformed quote Rewrite the content option with balanced quotes and escaped special characters. PCRE contains invalid syntax, unbalanced groups, or unsup- Simplify or rewrite the PCRE using valid Suricata-compatible ported modifiers syntax. Sticky buffer appears without a following content or pcre Add a valid match under the buffer or remove the unused sticky buffer. match HTTP, DNS, TLS, or SMB buffer is used in a rule with the Change the alert protocol or replace the buffer with one matching wrong application protocol the observed application layer. sid is missing, non-numeric, duplicated, or zero-valued Replace with a single valid numeric sid. rev is missing or non-numeric Replace with a valid numeric rev. flow: contains an unrecognized or incompatible token Rewrite flow: using valid options such as to_server, to_client, and established. Two rules share the same sid, or one rule contains multiple Keep one sid per rule and assign unique identifiers across gensid fields erated rules. rev appears more than once in one rule Keep a single rev field. threshold clause is malformed Rewrite the clause using valid threshold syntax or remove it if not required for detection. Suricata parser error does not match a specific known pattern Use the parser feedback to check option ordering, semicolons, parentheses, and unsupported keywords. nocase appears without a preceding content option Move nocase after a valid content match or remove it. Replace with a valid DNS response code value or remove the dns.rcode uses an invalid format or out-of-range value condition. tls.version uses an invalid decimal or unsupported version Use canonical Suricata constants such as tls1.0 or tls1.3. string Legacy ssl_version uses an unrecognized version constant Replace with a valid TLS/SSL version keyword or use the modern Suricata TLS version field. Rule direction or flow: does not match the observed traffic Align the rule direction with the observed client/server flow. direction Alert protocol does not match the observed application layer Use the correct alert protocol, e.g., HTTP → alert http. ASCII content is used for SMB2 traffic Convert SMB strings to UTF-16LE hex encoding. smb.share buffer is empty for tree-connect operations Use raw UTF-16LE hex content matching instead of smb.share. Content string is too long or contains session-specific tokens Replace with a shorter stable substring that appears in the target flows. http.request_body is used without checking the HTTP Add http.method; content:"POST"; or remove body matchmethod ing for GET traffic. Rule also triggers on benign traffic Tighten the match with malware-specific content, add negated benign patterns, or narrow protocol and flow scope.

24

H.2

Syntax-Invalid Feedback Template

matched. Figure 12 shows the template used to narrow the rule while preserving malware detection.

For syntax failures, the verifier provides the failed rule, the Suricata parser error, and the matching knowledge-base guidance. Figure 10 shows the syntax-invalid feedback template.

FALSE_POSITIVE Feedback Template The following Suricata rule triggers on benign traffic: {rule_text}

SYNTAX_INVALID Feedback Template

It triggered on {n_benign} benign PCAP(s).

Rule to fix: {rule_text}

Knowledge base: {knowledge_base_section}

Error/Problem: {suricata_parser_error}

Benign traffic that caused false positives: {benign_context}

Knowledge base: {knowledge_base_section}

Malware traffic the rule correctly matched: {malware_flow_context}

Return ONLY the corrected rule string.

Task: make this rule MORE SPECIFIC so it still triggers on malware traffic but stops triggering on benign traffic.

Figure 10: Syntax-invalid feedback prompt template. Runtime placeholders are filled with the failed rule, parser error, and matching syntax-error knowledge-base entry.

H.3

Strategies: - Add content present in malware but absent from benign flows. - Use negated content to exclude benign patterns. - Narrow PCRE anchors to avoid common benign strings. - Add bsize constraints if malware queries have distinctive length.

No-Trigger Feedback Template

Return ONLY the corrected rule string.

For no-trigger failures, the rule passes syntax validation but does not alert on the source malware PCAP. Figure 11 shows the template used to inject verifier feedback and malware flow context.

Figure 12: False-positive feedback prompt template. Runtime placeholders are filled with the failed rule, benign-triggering flows, malware-matching flows, and false-positive repair guidance.

NO_TRIGGER Feedback Template The following Suricata rule passed syntax check but did NOT trigger on the malware PCAP: {rule_text}

I

{sub_type_hint} Knowledge base: {knowledge_base_section}

A rule that triggers on its source PCAP does not by itself show how specific or general it is across the broader dataset. To measure this, we perform a source–target generalization analysis. For each source PCAP, we replay its associated rule set against every other PCAP in the dataset, producing one source– target pair per target. We apply the same procedure to both the ground-truth ET Open rules and the RULE AUTO P ILOTgenerated rules, allowing us to compare their cross-PCAP behavior. We label each source–target pair using the malware-family labels. A pair is same-family if the source and target share the same family label, and other-family otherwise. We exclude PCAPs with a SINGLETON family label, since singleton sources have no family-labeled targets. This leaves 1,281 non-singleton PCAPs. In the few-shot-5 run, RULE AU TO P ILOT produced valid rules for 1,074 out of 1,281 nonsingleton source PCAPs. For a fair comparison, we restrict both ET Open and RULE AUTO P ILOT’s generated rules to this matched 1,074-source subset and replay each source rule set against the remaining 1,280 non-singleton target PCAPs, excluding the source PCAP itself.

Rule-vs-traffic analysis: {flow_context} Rewrite this rule so it triggers on the target malware traffic. Constraints: - Compare current anchors against observed traffic. - Remove anchors absent from target traffic. - Prefer anchors repeated across target flows. - Avoid exact full URLs or brittle regexes unless supported. Return ONLY the corrected rule string.

Figure 11: No-trigger feedback prompt template. Runtime placeholders are filled with the failed rule, no-trigger subtype, knowledge-base guidance, and malware flow context.

H.4

Rule Family Generalization

False-Positive Feedback Template

For benign false positives, the verifier provides both benigntriggering flows and malware flows that should remain 25

triggering across multiple targets. For example, (wacatac → connectwise) increases by +0.21, while (connectwise → wacatac) increases by +0.15. This suggests that cross-family leakage is concentrated in specific family pairs rather than uniformly distributed across all pairs. Overall, RULE AUTO P ILOT captures useful family-level patterns while largely maintaining cross-family specificity.

Trigger-Rate Difference: RulePilot vs. Ground Truth 0.70

blank

+0.13 +0.13

+0.11 +0.13

+0.10

sfuzuan

+0.22 +0.21

Source family

0.35

vidar

+0.09

-0.66

asyncrat

+0.47

+0.05

connectwise

+0.15 +0.15 +0.15

agenttesla

+0.06 +0.06

-0.07

-0.09

+0.11

+0.13 -0.16 +0.14

-0.08

+0.06 +0.15 +0.11

0.00 +0.06

autoinject quasar

-0.13

-0.11

−0.35

wacatac

+0.05 +0.21 +0.21

+0.15 +0.21 +0.12

+0.07 +0.12 +0.14

Delta trigger rate = FS-5 minus Ground Truth

corewarrior

J

This appendix reports supplementary evaluation results that support the main evaluation but are omitted from the main text for space.

remcos xworm

-0.11

nk

an

ri u ar uz sf ew or

a

bl

or

c

-0.11

t s m at ar ac ise sla ec co cr or as at w te nj m ct yn ac oi xw qu nt re w as ut ne ge a n a co

ar

d

vi

Additional Evaluation Results

−0.70

Target family Over-trigger

Under-trigger

J.1

Per-Variant Prompt Comparison

J.2

Micro-Averaged Prompt Comparison

The main text reports macro-averaged metrics, where each PCAP contributes equally. Table 21 reports the corresponding micro-averaged results, where true positives, false positives, and false negatives are pooled across all 1,296 PCAPs before computing precision, recall, and F1.

Figure 13: Trigger-rate difference between RULE AUTO P ILOT (few-shot-5) and ground-truth rules per source–target family pair.

J.3

For each pair, trigger rate measures whether the source rule set fires at least once on the target PCAP. Because triggering alone does not imply correct security coverage, we also compute FAS precision, recall, and F1 by comparing the flows alerted by the source rule set on the target PCAP with that target’s ground-truth alerted flows. All metrics are reported as micro-averages, pooling raw flow counts across all pairs within each relation type, so that families with more source– target pairs contribute proportionally more to the aggregate results. Table 4 in §6.3 reports the aggregate same-family and other-family rates; overall, RULE AUTO P ILOT-generated rules exhibit generalization behavior comparable to the groundtruth ET Open rules, transferring to unseen same-family targets with near-ground-truth fidelity (FAS-F1: 0.67 vs. 0.75), consistent with the hypothesis that rules should align more strongly with samples from their own malware family than with samples from other families. Figure 13 below breaks this down further at the family-pair level. Figure 13 visualizes trigger-rate differences across the top12 families. Red cells indicate target families where RULE AU TO P ILOT ’s generated rules over-trigger relative to ground truth; blue cells indicate under-triggering. White cells indicate that RULE AUTO P ILOT matches ground-truth triggering behavior on those family pairs. The largest deviations are concentrated in a small number of source–target pairs: asyncrat → sfuzuan shows strong over-triggering (+0.47), while vidar → asyncrat shows strong under-triggering (−0.66). Several families, such as wacatac and connectwise, show mild over-

Repair Agent Outcome Details

Table 22 breaks down Repair Agent outcomes by initial failure status and knowledge-base section. It shows which failure types were most common and which were most often repaired successfully.

26

Table 20: Per-variant end-to-end performance across 1,296 malware PCAPs (full verification, gpt-oss-120b), macro-averaged. Flow Classification

Rule Quality

Variant

P

R

F1

FAS-P

FAS-R

FAS-F1

Syntax

Trigger

FPR

Zero-shot CoT Few-shot-1 Few-shot-3 Few-shot-5 CoT+FS-3

0.45 0.44 0.43 0.43 0.44 0.43

0.79 0.77 0.78 0.78 0.79 0.78

0.52 0.51 0.50 0.51 0.52 0.51

0.50 0.50 0.50 0.51 0.51 0.51

0.67 0.64 0.67 0.68 0.68 0.66

0.53 0.52 0.53 0.53 0.54 0.53

86.2% 84.5% 85.1% 86.1% 85.3% 85.2%

80.9% 79.2% 80.5% 82.0% 81.1% 80.0%

0.000% 0.006% 0.006% 0.000% 0.006% 0.010%

Table 21: Micro-averaged end-to-end performance across 1,296 malware PCAPs. Flow Classification

Rule Quality

Variant

P

R

F1

FAS-P

FAS-R

FAS-F1

Syntax

Trigger

FPR (%)

Zero-shot CoT Few-shot-1 Few-shot-3 Few-shot-5 CoT+FS-3

0.422 0.424 0.410 0.405 0.417 0.411

0.870 0.830 0.869 0.859 0.868 0.840

0.568 0.561 0.557 0.551 0.563 0.552

0.502 0.514 0.503 0.493 0.505 0.514

0.749 0.692 0.766 0.754 0.771 0.735

0.601 0.590 0.608 0.596 0.611 0.605

94.40% 93.80% 93.90% 95.40% 95.10% 93.70%

83.80% 82.80% 84.30% 84.80% 83.70% 83.30%

0.0799 0.0721 0.0815 0.0819 0.0817 0.0766

Table 22: Repair Agent (Agent 3) knowledge-base sections accessed, GPT-OSS-120b full run. Init. Status

KB Section

Rules

Fixed

Fix%

SYNTAX_INVALID

Unknown keyword Other syntax error Direction conflict bsize value too large TLS version keyword misuse Content quote error DNS rcode misuse Port syntax error Subtotal

40 37 35 6 3 2 1 1 125

15 0 27 6 0 1 1 0 50

38% 0% 77% 100% 0% 50% 100% 0% 40%

NO_TRIGGER

No specific match (generic KB) endswith on http.uri GET with request body Direction mismatch TLS cert.subject + direction HTTP resp. buffer on to-server Request body w/o method check SMB ASCII encoding Subtotal

78 10 10 9 6 3 2 1 119

52 4 8 8 0 0 0 1 73

67% 40% 80% 89% 0% 0% 0% 100% 61%

FALSE_POSITIVE

Rule triggers on benign traffic Subtotal

3 3

1 1

33% 33%

27

Record · ID 919267 · SHA-256 b404aa40bebcbd36
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.