Conceptio › Archive › arXiv CS
arXiv CSopen access

Autonomy in Check: Governor-Mediated Adaptive Security at the Edge

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Autonomy in Check: Governor-Mediated Adaptive Security at the Edge Ijaz Ahmad∗ , Ijaz Ahmad† , Flavio Esposito‡ , and Erkki Harjula∗ ∗ Centre for Wireless Communications, University of Oulu, Oulu, Finland

Email: [email protected], [email protected] † VTT Technical Research Centre of Finland, Espoo, Finland

Email: [email protected] ‡ Department of Computer Science, Saint Louis University, St. Louis, Missouri, USA

arXiv:2609.18338v1 [cs.CR] 16 Sep 2026

Email: [email protected]

Abstract—Adaptive security at the network edge increasingly relies on automated planners, including rule-based controllers, learned policies, and LLM-assisted agents, that translate observations into enforcement actions. Once such a planner can influence live policy state, syntactic validity is not enough. A semantically wrong action, produced from incomplete or manipulated observations, can be faithfully executed by an enforcement substrate that cannot judge mission context. We address this problem by treating the boundary between planner output and kernel enforcement input as the primary security object. We propose a split-control architecture in which an untrusted planner emits typed security intents, a deterministic governor checks each intent against safety, resource, temporal-stability, and proportionality invariants, and only admitted actions are bound to signed receipts and compiled into pre-installed eBPF map updates. The paper formalizes this trust-boundary problem, defines three threat classes, develops the governor admission predicate, and reports an end-to-end prototype. Across rule-based and LLM-assisted planners on a Raspberry Pi 5 testbed connected to the university’s 5G Test Network, the governor admits, rejects, and bounds intents at microsecond cost without disrupting protected-flow regularity. The contribution is conceptual as much as empirical: adaptive security does not need to trust the author of an action. It needs a mediation boundary that decides whether the action is admissible. Index Terms—Adaptive security, healthcare, agentic AI, edge computing, eBPF, policy governor, 5G, 6G, resilience, intentbased security, split-control architecture, reference-monitor, policy oscillation.

I. I NTRODUCTION Modern edge systems increasingly rely on adaptive security mechanisms that adjust enforcement policies in response to changing network conditions, resource constraints, and observed threats. These mechanisms are often driven by automated planners, ranging from rule-based controllers to learning-based and LLM-assisted agents, that translate observations into actions such as monitoring, rate limiting, tier escalation, or isolation [1], [2]. This shift improves responsiveness, but it also changes the security problem. The adaptation process itself becomes part of the attack surface. Current adaptive security systems often trust planner outputs once they are syntactically valid, even though semantically unsafe actions can emerge from manipulated observations or imperfect reasoning. A planner operating on incomplete,

noisy, or adversarially shaped inputs may issue an action that is locally consistent with policy but unsafe for the mission context. The enforcement substrate then executes that action faithfully. The risk is not only that a threat is missed. It is also that a controller may throttle protected traffic, under-react to an active attack, or oscillate rapidly between enforcement states because its inputs have been shaped by an adversary [3]–[5]. This problem is not tied to a single application domain. In industrial control systems, security mechanisms must account for control-loop timing and survivability requirements [6]. In vehicular and emergency-response edges, adaptation must support delay-critical decisions under changing mobility and load [7]. Private 5G deployments face similar pressure because mixed-criticality traffic classes share the same access layer [8]. O-RAN xApp/rApp ecosystems add another instance of the same problem, since control-plane intelligence can influence operational state and may itself be compromised [9]. Healthcare ward networks provide a concrete example of the same pattern. Clinical alarms and routine telemetry may traverse the same access network [10], [11]. An adaptive controller that is too restrictive can delay or suppress urgent flows, while one that is too permissive leaves the same path open to a compromised endpoint. Across these domains, the key question is no longer only how to detect malicious behavior, but how to constrain what an adaptive controller is allowed to do when its inputs cannot be fully trusted. Existing closed-loop and intent-based control frameworks automate policy refinement and actuation, while programmable substrates such as extended Berkeley Packet Filter (eBPF) provide efficient in-kernel observability and enforcement [12]–[16]. These foundations are important, but efficiency and well-typed actuation do not by themselves establish semantic safety. The kernel can enforce a rate limit or update a policy map, but it cannot decide whether the requested update is appropriate for the current mission context. Likewise, model-level defenses for LLM and agent security reduce planner-side risk, but they do not remove the need to mediate unsafe outputs that may still be produced [17], [18]. This paper treats the planner-enforcement boundary as the primary security object. We propose a split-control architecture

grounded in the reference-monitor principle [19]. The planner may use rules, learned logic, or LLM-assisted reasoning, but it cannot directly modify enforcement state. Instead, it emits typed security intents. A deterministic governor mediates each intent against safety, resource, temporal-stability, and proportionality invariants before any action can be compiled into the eBPF enforcement substrate. This paper makes the following contributions: • Trust-boundary formulation. We frame adaptive edge security as a trust-boundary problem in which policy-valid actions can be semantically unsafe under manipulated observations or imperfect reasoning. We define three concrete threat classes at the planner-enforcement interface: context manipulation, policy oscillation, and residual-window exploitation (Section III). • Governor-mediated architecture. We design a deterministic admission predicate over a restricted intent vocabulary. The predicate checks syntactic validity, protected-flow safety, resource headroom, temporal stability, and proportionality before enforcement. This separates intent mediation from planner-specific mechanisms such as action masking in safe reinforcement learning or per-call runtime verification (Section IV). • Bounded enforcement with integrity. We bind every admitted intent to an HMAC-protected receipt and confine the compiler to a pre-installed set of enforcement state space, so neither the planner nor a compromised compiler can synthesize new policy logic. Every approved action lands in an append-only audit log (Section IV). • Prototype evaluation under structured threats. We implement the design on a Raspberry Pi 5 attached to the University’s 5G Test Network with three ESP32 endpoints. We evaluate rule-based and LLM-assisted planners under the three threat classes, characterize governor overhead, and compare governed behavior against static and no-governor baselines (Section VI). The results support a simple design principle: autonomy belongs in planning, but actuation should remain mediated, bounded, and auditable. Section II situates the work against closed-loop control, eBPF enforcement, AI-driven control ecosystems, LLM security, and runtime policy mediation. Section III develops the system and threat model. Section IV presents the governor-centered design. Section V describes the prototype and attack emulation methodology. Section VI reports the evaluation results and discusses deployment implications. II. R ELATED W ORK Closed-loop and intent-based control. Closed-loop network control has developed through autonomic networking, zerotouch service management, intent-based networking, and edgecloud orchestration. ETSI ZSM models closed-loop automation as an observe, decide, and actuate process with minimal human intervention [12]. Intent-based networking similarly raises the interface from low-level configuration to desired

outcomes, leaving the system to refine those outcomes into device-level actions [13]. Recent work on adaptive trust and edge-cloud orchestration applies related ideas to 6G and IoT settings, where service quality, resource headroom, and security posture must be adjusted at runtime [1], [20]. These systems are important foundations for adaptive security, but they typically treat the controller, planner, or intent source as trusted once its output is well formed. Our work focuses on the complementary problem. A planner output may be syntactically valid and still semantically unsafe when observations are manipulated or reasoning is imperfect. Programmable enforcement with eBPF. eBPF has become a practical substrate for programmable observability, traffic control, and cloud-native security because it offers stable kernel hooks, verifier-checked programs, and low-overhead access to packet and system events [14], [16]. Recent systems also use eBPF to observe AI and agent workloads at the system level [21]. At the same time, work on BPF hardening and kernel-extension safety shows that verifier acceptance and efficient execution are not the same as end-to-end security correctness [22]–[24]. eBPF can enforce a rate limit, select an inspection path, or update a policy map efficiently. It does not decide whether the requested action is appropriate for the mission context. Open and AI-driven control ecosystems. Open and programmable control ecosystems make the planner–enforcement boundary more important. O-RAN introduces xApps and rApps that can optimize radio and service behavior through near-real-time and non-real-time control loops, but this openness also creates risks from compromised or malicious control functions [9], [25]. Recent agentic-AI proposals extend this direction by using LLM-assisted reasoning for closed-loop RAN or network-security tasks [17], [18]. LLM-integrated systems also face prompt injection, indirect prompt injection, and adversarially shaped inputs that can cause an agent to infer or follow unintended context [3]. Guidance on deployed AI systems further emphasizes that monitoring remains difficult once AI components interact with tools, external data, and operational environments [4], [5]. These works reduce plannerside risk, but they do not remove the need to mediate unsafe outputs that may still be produced. Runtime safety and policy mediation. The closest conceptual anchor for our work is runtime policy mediation. The reference-monitor principle requires security-sensitive operations to be mediated by a mechanism that is always invoked, tamper-resistant, and small enough to reason about [19]. Schneider’s work on enforceable security policies formalizes policies that can be enforced by monitoring execution and suppressing disallowed actions [26]. Runtime verification checks execution traces against formal specifications while a system runs [27]. In learning-enabled control, shielding prevents unsafe actions from a learner before they affect the environment [28]. Our governor follows the same safetyoriented spirit, but it mediates a different object. A shield typically masks individual learner actions against a verified safety

model. Runtime verification observes program traces against a specification. Intent-based frameworks refine desired outcomes into actions while usually assuming that the intent source is trusted. Our governor instead mediates typed security intents before they can modify live enforcement state. The planner is explicitly outside the trusted computing base, and each proposed intent must clear the predicate Γ over deploymentspecific invariants such as Fcrit , rmin , ∆min , and ϕ. Identified gap. Prior work has advanced adaptive control, programmable enforcement, open control ecosystems, AIassisted security, and runtime safety. These lines of work make critical edge systems more flexible, but they often assume that a well-formed controller output is safe to execute. We take a different position. In critical edge deployments, the planner’s output is part of the attack surface. A semantically wrong but syntactically valid action can delay protected traffic, weaken containment, or drive unstable policy changes. Our architecture therefore separates planning from actuation: the planner proposes, the governor mediates, and the eBPF substrate enforces only actions that pass explicit safety and stability checks. III. T RUST B OUNDARY AND T HREAT M ODEL A. Deployment Context and System Model We consider a critical edge network carrying mixedcriticality traffic over a 5G access layer, with processing tiers at the local edge, MEC, and the cloud. Some flows are urgent and delay-sensitive, while others are routine and tolerant of inspection or queuing delays. A security controller observes the communication state and decides whether to continue passive monitoring, rate-limit a traffic class, escalate inspection to a higher-capacity tier, or isolate a flow subset. Each decision affects both security posture and the service quality of the flows under management. The proposed defense mechanism is structured as three logical modules. A planner receives the current observation state and proposes a typed security intent. A governor checks whether that intent is admissible under the current mission and resource constraints. A compiler translates approved intents into bounded updates over a pre-installed eBPF enforcement substrate. The planner may use rule-based logic, a learned policy, or an agentic reasoning component. Its internal mechanism is outside the trust boundary argument. What matters is that it cannot modify enforcement state without governor approval.

capabilities, three threat classes emerge, summarized in Table I and illustrated in Fig. 1. T1 – Context manipulation. The adversary manipulates observable inputs, such as inflated anomaly signals, masked malicious flows, or workloads shaped to mislead, to induce an action that is syntactically valid but semantically wrong. Suppressing the measurement of an ongoing attack, for instance, causes the planner to perceive low threat and propose passive monitoring when isolation is warranted. This class maps directly onto indirect injection attacks studied in LLMintegrated systems [3], where the attack targets the input environment rather than the model itself. T2 – Policy oscillation. Rather than inducing a single large policy error, the adversary alternates observable behavior at a rate fadv chosen to keep the planner near a decision boundary. If fadv > 1/∆min , where ∆min is the governor’s minimum inter-action cooldown, each enforcement transition creates a brief window during which the active policy is mismatched with the actual traffic condition. Repeated over time, this degrades both protection coverage and service quality without requiring any single catastrophic decision. T3 – Residual-window exploitation. Every enforcement state transition involves a brief window before the new policy takes full stable effect. An adversary who can trigger policy transitions, whether through context manipulation or oscillation, can time traffic delivery to coincide with this window. The primary mitigation is bounding the time from governor approval to stable kernel enforcement, addressed in Section IV-C. TABLE I T HREAT C LASSES AT THE P LANNER –G OVERNOR I NTERFACE

Class Adversary action

Governor defense

T1

Shape observations to deceive the planner Alternate behavior to drive rapid enforcement switching Time delivery to coincide with transition gaps

Proportionality invariant (I5); uncertainty gating Temporal stability invariant (I4) Bounded compiler latency; receipt sequencing

telemetry

at

T2 T3

Trusted Base

eBPF Observer

T1 ⃝

State Builder

xt

Planner

ρt

Governor

T2 ⃝

reject

eBPF Enforcer

T3 ⃝

Adversary

B. Threat Model The governor and enforcement substrate are part of the trusted computing base, whereas the planner is not. The planner therefore cannot load arbitrary eBPF programs, hold file descriptors to enforcement maps, or directly invoke the compiler. This boundary is what we study, not something we abstract away. The adversary may inject malicious traffic [29], compromise a non-critical endpoint device, or influence the telemetry and context presented to the planner. From these

Fig. 1. Threat-annotated pipeline overview. The adversary targets three interfaces in the adaptation path.

C. Problem Formulation: Governor-Mediated Actuation To formalize how the governor constrains planner actions under the threats described above, we model each decision as an intent proposed from the current system state and independently admitted before enforcement. In particular, at decision epoch t, the security controller observes xt = [ot , ct , rt , qt , ht , mt , ut ],

(1)

where ot denotes traffic observations, ct flow criticality labels, rt resource headroom per enforcement tier, qt queue and path state, ht a bounded history of recent policy decisions, mt mission-context indicators such as active alarm sessions and protected service classes, and ut ∈ [0, 1] an uncertainty signal derived from measurement freshness and cross-source consistency. Let A denote the typed intent vocabulary and let L = {local, MEC, cloud} denote the available execution tiers. A planner proposes an intent at ∈ A and a tier ℓt ∈ L, which can be abstracted as the following planner-side problem: min. Rt (at , ℓt ) + λCt (at , ℓt ) + µOt (at , ht ) at ∈A ℓt ∈L

s.t.

Dtprot ≤ Dmax ,

(2)

Ut ≤ Umax , at ∈ Asafe (xt ). Here, Rt denotes the planner’s estimate of residual security risk, Ct denotes resource and coordination cost, and Ot penalizes unstable policy changes such as reversals on the same scope. Dtprot is the protected-flow service bound and Ut is a resource-utilization bound. The state xt is supplied by the state builder at each epoch. The safe action set Asafe (xt ) is not computed by the planner. It is the output of the governor admission predicate, evaluated independently of how the proposed intent was generated. A planner operating on manipulated observations may still produce an intent in A; the governor determines whether that intent belongs in Asafe (xt ) under the actual system state. This separation between proposal and admissibility is the central structural property of the design. Prototype instantiation. Equation (2) is a conceptual abstraction of the planner-side objective. The prototype does not implement a numerical optimizer for Rt , Ct , or Ot , and the governor does not trust the planner’s estimates of these terms. Instead, the governor evaluates only the invariant predicate in Section IV-B. The state builder provides the a bounded severity signal from smoothed drop and queuepressure indicators, the uncertainty signal ut , and the intent impact ordinal shown in Table II. The proportionality bound is ϕ(s, u) = max(1, 4s·c(u)), where c(u) = 1 for u ≤ 0.5 and decreases linearly to 0 at u = 1. The floor of 1 keeps ROLL BACK admissible at any uncertainty. The cooldown ∆min and reserve rmin instantiate the temporal-stability and protectedheadroom checks used by I4 and I3. IV. G OVERNOR -C ENTERED D ESIGN A. Bounded Intent Vocabulary Restricting the planner to a fixed vocabulary is the first step in bounding its actuation authority. Rather than accepting free-form policy expressions or arbitrary kernel operations, the governor sees only the typed intents in Table II. The vocabulary covers passive observation, bounded rate control, tier escalation, non-critical isolation, human-approval requests, and state rollback. High-impact actions are named explicitly so

the governor can apply proportionality checks independently of planner intent. TABLE II T YPED SECURITY INTENT VOCABULARY WITH IMPACT ORDINAL . Intent (ord)

Parameters and effect

M ONITOR (0)

scope, duration; passive observation, no traffic effect

R ATE L IMIT (2)

scope, rate tier, duration; bounded throttling with mandatory exemptions for flows in Fcrit E SCALATE (3) scope, target tier, timeout; requests higher-capacity inspection without modifying the local rate policy I SOLATE -NC (4) scope, duration; isolates only non-critical flows or devices; critical flows are unaffected by construction R EQ A PPROVAL reason, suggested action, urgency; suspends direct (0) actuation and requests operator review ROLLBACK (1) receipt-id, target state; reverts to the last governorapproved stable enforcement state

B. Admission Predicate The governor admits at if and only if Γ(at , xt ) = 1, where Γ = I1 ∧ I2 ∧ I3 ∧ I4 ∧ I5 . The safe action set follows as Asafe (xt ) = {a ∈ A | Γ(a, xt ) = 1}. Each invariant addresses a distinct failure mode. I1 – Syntactic validity. at belongs to the declared vocabulary with valid parameter types and in-range values. Syntactic validity is necessary but not sufficient; the remaining invariants check semantic correctness. I2 – Flow safety. A restrictive action may not affect flows in Fcrit without an explicit policy exemption: scope(at ) ∩ Fcrit = ∅ (3) ∨ type(at ) ∈ {R EQ A PPROVAL, ROLLBACK}. Regardless of what the planner proposes, this invariant prevents urgent flows from being throttled or isolated without explicit authorization. I3 – Resource bound. Applying at must leave sufficient bandwidth headroom for protected services: headroom(rt , at ) ≥ rmin . (4) I4 – Temporal stability. The last approved action on scope(at ) must have occurred at least ∆min epochs prior:  t − tlast scope(at ) ≥ ∆min . (5) An adversary driving oscillation at fadv > 1/∆min cannot force enforcement changes at that rate. The governor rejects each intent until the cooldown elapses, bounding the achievable oscillation rate above by 1/∆min regardless of how quickly the planner proposes new intents. I5 – Proportionality. The impact of at must not exceed what observed threat severity and current uncertainty jointly support:  impact(at ) ≤ ϕ severity(ot ), ut , (6) where ϕ decreases monotonically in ut . As uncertainty rises, ϕ contracts toward Low, directing the planner to M ONITOR, R EQ A PPROVAL, or ROLLBACK rather than direct mitigation. On rejection, the governor returns a typed reason identifying which invariant failed and, where applicable, a lower-impact alternative. This allows the planner to revise its proposal

without a full state reset, preserving adaptability while maintaining the admission boundary. Each invariant removes a different unsafe path from the planner to enforcement. I1 prevents malformed or out-of-range intents from entering the compiler path. I2 protects flows in Fcrit from restrictive actions unless an explicit policy exemption exists. I3 prevents an admitted action from consuming the headroom reserved for protected services. I4 bounds the rate of policy transitions on a scope and therefore limits adversary-induced oscillation. I5 prevents high-impact actions when the observed severity and uncertainty do not justify them. Section VI shows how these invariants fire under the three attack classes and under nominal LLM-assisted planning. C. Integrity Path to Kernel Enforcement An intent that passes the admission check is not written to enforcement maps directly by the planner or governor. Instead, the governor emits a signed approval record:  ρt = t, at , H(xt ), HMACk t ∥ at ∥ H(xt ) , (7) where H(xt ) is a hash of the decision-context snapshot and k is held exclusively by the governor process. A privileged compiler service, which is the only component holding eBPF map file descriptors and the relevant Linux capabilities, verifies ρt before performing any update. It then translates the approved intent into a bounded set of kernel-facing operations. These include selecting a pre-defined rate tier, enabling a predefined inspection path, or writing a schema-validated entry into a policy map. D. Proposed Split-Control Architecture The proposed framework has six logical components arranged as a unidirectional enforcement pipeline, shown in Fig. 2. Autonomy is confined to the planning stage. All live actuation authority rests within the trusted base. Observation plane. Kernel-resident eBPF programs attached at tc and XDP hooks export per-class packet and byte rates, drop counts, queue pressure, protected-flow activity indicators, and policy-map access timestamps through a ring buffer. These programs are pre-compiled and pre-verified at deployment. No user-space component can load or modify them. State builder. A user-space daemon assembles xt from ringbuffer data, computes ut as a confidence-weighted function of measurement freshness and cross-source consistency, and maintains ht as a bounded sliding window of past governor decisions and observed flow-quality outcomes. Planner. The planner receives xt and returns a structured intent from the vocabulary in Table II. The implementation may be an LLM-based agent over a typed tool interface, a learned policy, or a rule-based controller. The architecture is independent of this choice. Governor. The governor evaluates Γ(at , xt ) per Section IV-B, emits ρt on admission, and returns a typed rejection with an optional lower-impact suggestion otherwise.

Observation ring buffer

telemetry

State Builder state

Planner

reject

intent

Flow Score Assets0.82 A

Governor Enforcement rate map

policy map

Cls Tier Assets2 A

Flow Act Assets B shape

Compiler map update

Trusted base

I1 syntax I2 flow safety receipt I3 headroom I4 cooldown I5 proportionality

eBPF zone

Fig. 2. Compact split-control architecture. The planner proposes typed security intents from the built state, but only the governor may admit them for enforcement. Approved actions become bounded eBPF map updates after compiler-side receipt verification. Rejected actions return typed feedback for replanning.

ESP32 (alarm)

WiFi (wlan0)

RPi 5 enforcement node

(wwan0) 5G uplink

(n78 / 3.5 GHz)

UPF

eBPF hooks (tc/wlan0)

ESP32 (alarm)

5GTN Core

AMF/SMF

State Builder Governor Planner

ESP32 (routine) ■ Untrusted endpoint

□ Trusted base (RPi5)

■ 5GTN core

Fig. 3. Testbed topology. ESP32 nodes connect to the RPi 5 enforcement node over WiFi (wlan0). The RPi 5 runs the state builder, governor, and planner, with eBPF hooks on wlan0 for enforcement. The wwan0 interface provides 5G connectivity to the 5GTN core.

Compiler. After verifying the HMAC in ρt , the compiler translates the approved intent into bounded map updates and appends the receipt to the audit log. Enforcer. Approved actions become bounded eBPF map updates after compiler-side receipt verification. V. P ROTOTYPE I MPLEMENTATION A. 5G Edge Testbed The prototype runs on the University’s standalone 5G Test Network (5GTN) [30], an academic research and experimentation platform operating on the n78 band (3.5 GHz) with an on-premise 5G core. A Raspberry Pi 5 (4-core CortexA76, 8 GB RAM, kernel 6.12.79) serves as the enforcement node and gateway. Its wlan0 interface connects to the ESP32S3 endpoint devices over a local WiFi access point, and its wwan0 interface provides 5G connectivity to the 5GTN core over the cellular uplink. eBPF programs compiled against a BTF-generated vmlinux.h attach at tc ingress on wlan0 and use three BPF maps: an LRU hash for per-source epoch rate counters, a hash for compiler-written rate limits, and a 128 kB ring buffer for telemetry export. All user-space components run on the RPi5, including a Go state builder that assembles xt every 100 ms epoch, a Go governor, a privileged Go compiler, and a planner. Fig. 3 shows the testbed topology. Two traffic classes originate from the ESP32-S3 firmware. The protected alarm class generates 20-byte UDP datagrams at

B. Adversarial Workloads The three threat classes in Table I are executed independently with a 30 ms cooldown to avoid overlap. For T1, the context-manipulation script suppresses alarm-source reports from two of the three ESP32 nodes before they reach the state builder, causing the planner’s view of protected-flow activity to diverge from the receiver-side traffic record. For T2, a test process injects intents directly to the governor’s Unix socket at fadv = 2/∆min . This models a co-located adversarial process and represents a conservative worst case, since a remote attacker would face additional transit latency that reduces effective injection frequency. For T3, hping3 on the RPi5 injects UDP bursts timed to epoch boundaries, immediately after a governor-approved enforcement transition, to probe the stabilisation window. VI. E VALUATION This section evaluates the proposed split-control design in the 5G test network. The goal is not only to measure enforcement performance but also determine whether the governor preserves the security boundary between planner output and kernel actuation under realistic and adversarial conditions. We therefore examine four aspects: protected-flow cadence, blocking of unsafe actions, control-path overhead, and robustness to planner-side manipulation. A. Security and Performance Results The evaluation uses four configuration labels, listed in Table III. We use the friendly labels in figures and prose, while the internal run names are kept in the artifact CSV files for reproducibility. Static is the no-adaptation control, Rule is the deterministic planner with governor, Rule, no gov. is the ablation that removes the admission boundary, and LLM is the proposed planner-governor configuration. R EQ A PPROVAL and ROLLBACK are part of the vocabulary but are not triggered by the attack workload in this evaluation. Their inclusion lets the

TABLE III E XPERIMENTAL CONFIGURATIONS AND RECEIVER - SIDE PROTECTED - FLOW CADENCE . Config.

n p99 p99.9 gaps > 150 intervals excess (ms) count ms

Role

Static Rule

No adaptation control Rule planner with governor Rule, no gov. Governor ablation LLM LLM planner with governor

54032 5.84 18.31 54092 7.27 40.40

0 0

3 18

54061 7.02 41.02 54091 6.51 33.78

0 0

23 19

Note: Excess is measured over the nominal 100 ms alarm period using only RPi receive timestamps.

proportionality bound contract toward operator-mediated states rather than direct mitigation when uncertainty is high. M1 – Receiver-side protected-flow cadence (Fig. 4, Table III). The earlier version of the experiment reported ESP32-to-RPi one-way latency. We do not use that metric here because the long runs revealed cross-device clock drift and invalid negative samples. M1 therefore asks a narrower and more reliable question: does the protected alarm stream remain regular at the receiver? Across all four 30-minute baseline runs, the answer is yes. The p99 excess over the nominal 100 ms alarm period remains below 7.3 ms, and no sequence gaps are observed in any configuration. The p99.9 excess remains below the 50 ms service slack for all four baselines. Rare intervals above 150 ms are reported explicitly in Table III rather than hidden inside a latency CDF. Cadence excess over 100 ms (ms)

10 Hz, corresponding to a nominal 100 ms inter-arrival period, while the routine class generates 1 kB telemetry bursts at 1 Hz. We do not measure the ESP32-to-RPi one-way latency as an evaluation metric because the long runs exposed cross-device clock drift and occasional invalid negative latency samples. Instead, protected-flow service is evaluated using receiverside cadence. Inter-arrival times are computed directly from RPi receive timestamps, grouped by source IP and sequence number. Baseline runs last approximately 30 minutes, and each attack run lasts approximately five minutes. The LLM planner uses llama3.1:8b in Q4_K_M quantized form, served through Ollama [31] on a co-located GPU node (RTX 4000 Ada, 20 GB VRAM). The planner wrapper submits only structured tool-call intents to the governor. Non-tool or no-op outputs from the LLM are logged separately and do not update enforcement state. The prototype uses ∆min = 5 epochs, corresponding to 500 ms, and rmin = 6.4 kbps, a 2× reserve over the nominal protected-alarm rate.

50

p95

p99.9

p99

50 ms slack

40

30

20

10

0 Static

Rule

Rule, no gov.

LLM

Fig. 4. Protected-flow delivery cadence computed from receiver-side timestamps only. The figure shows excess inter-arrival time over the nominal 100 ms alarm period. The 50 ms reference corresponds to the service slack used in the original alarm budget. Sequence-gap counts are reported in Table III.

M2 – Attack containment (Table IV). The attack experiments exercise the governor boundary in three different ways. T1 activates the context-manipulation path in the three-ESP32 testbed: suppressing two alarm sources raises ut to 0.600 during the attack window. This is not reported as a traffic-latency result. It is evidence that the uncertainty channel responds when the planner’s context is manipulated. In T2, I4 rejects 49.7% of injected oscillation intents, which is the expected alternate-intent clipping pattern at fadv = 2/∆min . T3 remains partially contained as the residual window is bounded but non-

TABLE IV ATTACK - CONTAINMENT SUMMARY.

T1 T2

max ut during attack I4 rejection rate

T3

avg window / burst

Observed Result

Scenario

0.600 YES 49.67% YES 2.616 ms 1279 B PARTIAL

Rule+Gov, nominal Rule+Gov, T1 Rule+Gov, T2 Rule+Gov, T3 LLM+Gov, nominal LLM+Gov, T1 LLM+Gov, T2 LLM+Gov, T3

mean

103

p99

1 ms

mean (µs)

p99 (µs)

54.34 55.11 41.33 54.08 27.86 27.86 30.00 32.03

76.24 82.56 78.52 102.22 129.09 100.78 79.20 126.26

(a) Security Attack-intent admit rate (%)

Governor decision overhead (us, log)

T1: context manipulation; T2: policy oscillation; T3: residual-window attack.

n 1801 301 601 321 51 10 309 25

102

Rule+Gov

101 Nom

T1

T2

LLM+Gov T3

Nom

T1

T2

T3

Fig. 5. Governor decision overhead across nominal and attack scenarios. For each scenario, the open circle is the mean and the filled square is the p99; the connecting line is the spread. All p99 values remain below 1 ms, so the governor is sub-millisecond even under active attack.

zero, with an average window of 2.616 ms and an average admitted burst of 1279.2 B. We report this as a bounded residual leak rather than claiming complete elimination of transition exposure. The proportionality contraction engages only when severity also rises; under T1 with nominal traffic, severity stays low and I5 fires only when the LLM proposes high-impact actions in this elevated-uncertainty state. The rule planner under B2 remained at low impact, so its I5 rate stays near zero — see Fig. 7. M3 – Governor overhead (Fig. 5, Table V). Governor interposition remains sub-millisecond in every scenario. Across nominal and attack runs, p99 decision overhead ranges from 76.24 µs to 129.09 µs, well below the 1 ms reference used in Fig. 5. The result separates the safety boundary from the planner’s complexity. The governor check is fast even when the planner is stochastic or the input stream is adversarial. The original 5 ms engineering guardrail is still comfortably met, but a more impactful result for critical-edge systems is that admission control stays below one millisecond at p99. M4 – Governance tradeoff across configurations (Fig. 6). Fig. 6 compares the proposed governed configurations with two natural alternatives: Static, which has no adaptation channel, and Rule, no gov., where the rule planner writes to the compiler without the admission predicate. Panel (a) reports the fraction of attack-induced structured intents that reached enforcement. Static is marked N/A because it emits no adaptive intents. Rule, no gov. admits all planner-emitted

100%

100

87%

80 60

48%

40 20 N/A

0 tic ov. Rule LLM Sta no g , le Ru

Nominal p99 cadence excess (ms)

Attack Measure

TABLE V G OVERNOR DECISION OVERHEAD .

(b) Utility 8

6

7.02 7.27

6.51

5.84

4

2

0 tic ov. Rule LLM Sta no g , le Ru

Fig. 6. Governance tradeoff across four configurations. (a) Attack-intent admit rate. Static has no adaptation channel and is marked N/A. Rule, no gov. admits all planner-emitted intents because no admission predicate is present. Rule and LLM with governor admit 86.9% and 48.0% of attack-induced structured intents, respectively. (b) Nominal p99 cadence excess over the 100 ms alarm period. All four configurations remain within a 1.5 ms band. Static is hatched to mark the absence of adaptation.

intents by construction, while the governed Rule and LLM configurations admit 86.9% and 48.0%, respectively, because rejected intents do not reach the compiler. Panel (b) shows the utility side: nominal p99 cadence excess remains within a 1.5 ms band across all four configurations. The governor therefore changes what can be actuated under attack without measurably degrading receiver-side protected-flow cadence. M5 – Planner abuse robustness (Fig. 7, Table VI). The governor’s value is clearest when the planner is imperfect. In the nominal LLM run, the wrapper logged 473 LLM decisions. Only 51 were structured intents submitted to the governor; the remaining 422 were non-actuating outputs, mostly prose or JSON-like text converted to do_nothing. Among the 51 structured intents, 13 were admitted and 38 were rejected before enforcement: 19 by I1, five by I2, and 14 by I5. Table VI gives representative examples. Under T2, both planner families show the same I4 signature, with 149 cooldown rejections. Under T3, both paths show ten I2 flow-safety rejections. The same admission boundary applies regardless of how the intent was generated. Invariant I3 reserves rmin = 6.4 kbps of bandwidth headroom for the protected class, a 2× margin over the nominal alarm-class

Admitted

I1

I2

I4

I5

Rule+Gov Nom

n=1801

T1

n=301

T2

n=601

T3

n=321

LLM+Gov Nom

n=51

T1

n=10

T2

n=309

T3

n=25

0

20

40

60

80

100

Intent disposition (%) Fig. 7. Intent disposition by governor invariant. T1 represents context manipulation, T2 is policy oscillation, and T3 is residual-window exploitation. TABLE VI E XAMPLES OF STRUCTURED LLM INTENTS STOPPED BY THE GOVERNOR .

Proposed intent

Inv.

RATELIMIT with I1 zero rate RATELIMIT I2 protected source ESCALATE global I5 ISOLATE-NC impact scope

high- I5

Governor response Invalid parameter; RATELIMIT requires a positive rate. Restrictive action targets a source carrying alarm-class traffic. Action impact exceeds the proportionality bound. Isolation impact is not justified by observed severity and uncertainty.

load. In the measured workload, the alarm and routine traffic remain below this floor, so I3 does not activate. This is expected: I3 is a structural guardrail for saturation conditions, while I1, I2, I4, and I5 exercise the planner-governor boundary in the attacks evaluated here. B. Deployment Implications The results apply broadly to mixed-criticality scenarios in edge-could continuum. The protected flow set, the service budget for protected flows, and the cooldown ∆min are the main deployment-specific parameters. In a healthcare ward, Fcrit is the alarm class and a manipulated planner that throttles alarms threatens patient safety [10], [11]. In industrial control, Fcrit becomes the control-plane messages and the service budget is the process cycle time. In vehicular edge, ∆min shortens to match faster control decisions. The invariant logic, the HMAC receipt chain, and the eBPF substrate carry over unchanged. The prototype uses ∆min = 5 epochs (500 ms) and rmin = 6.4 kbps, corresponding to a 2× reserve over the nominal protected-alarm rate. These parameters expose the expected

safety–agility tradeoff. A larger ∆min makes oscillation attacks harder by enforcing a longer cooldown between successive policy changes, but also slows adaptation to benign changes in operating conditions. A smaller ∆min improves responsiveness, but increases susceptibility to oscillatory manipulation. Likewise, a larger rmin preserves more headroom for protected traffic and tightens the governor’s admission boundary, while a smaller rmin increases enforcement flexibility at the cost of a weaker safety margin. We leave a systematic parametersensitivity study to future work. The T2 result reflects why temporal stability matters in clinical and industrial edge settings. An adversary does not need to defeat the enforcement mechanism outright. Repeatedly forcing policy transitions is enough to create short periods of mismatch between the active policy and the traffic condition. I4 clips this behaviour at the admission boundary. In our experiment, 149 of the 300 injected oscillation intents fail the cooldown check, producing the expected near-50% clipping pattern for fadv = 2/∆min . This gives operators a simple deployment knob. ∆min can be chosen with the protected-flow period and acceptable adaptation rate in mind. The prototype evaluates the planner-governor boundary at the unit-cell scale. Three endpoints and a single enforcement node is the right granularity for the structural claims (I2 flow safety, I4 cooldown bounding, HMAC integrity). Larger-scale evidence possibly with synthetic flow replay across thousands of sources, multi-gateway coordination of cooldown state, and saturation testing for I3 is left for the extended version with a packet-level emulator. C. Scope and Implications Our prototype evaluates a deliberately small deployment: one RPi5 enforcement node, three endpoints, and a 5G uplink. An open problem is the coordination of governor state across multiple enforcement points in larger deployments. As importantly, the results suggest that adaptive security does not require trusting the component that makes the decision. The planner can be rule-based, learned, or LLM-assisted, while the same governor independently controls what is allowed to reach the enforcement layer. This separation is useful as planners become more capable but also harder to predict: new planning mechanisms can be introduced without granting them direct control over critical network state. The governor therefore provides a stable security boundary between evolving autonomous intelligence and the enforcement substrate. The experiments also expose a practical limit of this approach. Admission control can reject unsafe actions and bound how frequently policies change, but it cannot eliminate the short transition window after an action is approved. Applications with stricter timing requirements will therefore need faster enforcement activation in addition to governor-mediated admission. VII. C ONCLUSION Adaptive security in critical edge systems fails not only when a planner misses a threat, but also when enforcement

faithfully executes an unsafe decision based on manipulated observations. This paper treats the planner–enforcement boundary as the primary security object and protects it through a governor with five admission invariants. A prototype on the University’s 5GTN testbed with in-kernel eBPF enforcement shows that protected alarm delivery remains regular, with p99 cadence excess below 7.3 ms and no sequence gaps across baseline runs. Governor decisions remain sub-millisecond, with p99 below 130 µs across nominal and attack scenarios. Under policy oscillation, the governor rejects 49.7% of injected intents, while under nominal LLM planning it rejects 38 of 51 structured intents before enforcement. These results show that adaptive planners can retain flexibility without being granted direct authority over critical enforcement state. ACKNOWLEDGMENTS This research is supported by the Finnish Doctoral Program Network in Artificial Intelligence, AI-DOC (decision number VN/3137/2024-OKM-6), Business Finland funded projects TOMOHEAD (8095/31/2022) and SUNSET6G (8682/31/2022), and by the Research Council of Finland funded projects 6G Flagship (369116) and Profi6 (336449). The work of Dr. Flavio Esposito is supported by USA NSF Awards CNS #2133407 and OAC #2530896. R EFERENCES [1] I. Ahmad, S. Gimhana, I. Ahmad, and E. Harjula, “Adaptive trust architecture for secure IoT communication in 6G,” IEEE Networking Letters, vol. 7, no. 2, pp. 113–116, 2025. [2] A. Sangiorgi, A. Pinto, R. Tourani, and F. Esposito, “Mitigating deauthentication DoS attacks in 802.11 via eBPF and XDP,” in Proceedings of the IEEE Conference on Network Softwarization (NetSoft), Budapest, Hungary, June 2025. [3] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world LLMintegrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec ’23), 2023. [4] National Institute of Standards and Technology, “CAISI issues request for information about securing AI agent systems,” Jan. 2026, nIST news release, January 12, 2026. [Online]. Available: https://www.nist.gov/news-events/news/2026/01/ caisi-issues-request-information-about-securing-ai-agent-systems [5] ——, “Challenges to the monitoring of deployed AI systems,” NIST AI 800-4, Mar. 2026. [Online]. Available: https://nvlpubs.nist.gov/nistpubs/ ai/NIST.AI.800-4.pdf [6] A. A. Cárdenas, S. Amin, and S. Sastry, “Research challenges for the security of control systems.” HotSec, vol. 5, no. 15, p. 1158, 2008. [7] L. Liu, C. Chen, Q. Pei, S. Maharjan, and Y. Zhang, “Vehicular edge computing and networking: A survey,” Mobile networks and applications, vol. 26, no. 3, pp. 1145–1168, 2021. [8] J. Fue, J. A. Gutierrez, and Y. Donoso, “Understanding security vulnerabilities in private 5g networks: Insights from a literature review,” Future Internet, vol. 17, no. 11, p. 485, 2025. [9] O-RAN ALLIANCE, “O-RAN security threat modeling and risk assessment,” O-RAN ALLIANCE, Technical Report O-RAN.WG11.ThreatModel, 2024. [Online]. Available: https://www.o-ran.org/specifications [10] The Joint Commission, “Medical device alarm safety in hospitals,” Sentinel Event Alert, no. 50, pp. 1–3, Apr. 2013, pMID: 23767076. [Online]. Available: https://www.jointcommission.org/en-us/ knowledge-library/newsletters/sentinel-event-alert/issue-50 [11] S. Sendelbach and M. Funk, “Alarm fatigue: A patient safety concern,” AACN Advanced Critical Care, vol. 24, no. 4, pp. 378–386, 2013.

[12] ETSI, “Zero-touch network and service management (ZSM); closed-loop automation,” European Telecommunications Standards Institute, Tech. Rep. GS ZSM 009-1 v1.1.1, 2021. [Online]. Available: https://www.etsi.org/deliver/etsi_gs/ZSM/001_099/00901/01. 01.01_60/gs_zsm00901v010101p.pdf [13] A. Clemm, L. Ciavaglia, L. Z. Granville, and J. Tantsura, “Intent-based networking—concepts and definitions,” RFC 9315, 2022. [14] D. Soldani, P. Nahi, H. Bour, S. Jafarizadeh, M. F. Soliman, L. Di Giovanna, F. Monaco, G. Ognibene, and F. Risso, “ebpf: A new approach to cloud-native observability, networking and security for current (5g) and future mobile networks (6g and beyond),” IEEE Access, vol. 11, pp. 57 174–57 202, 2023. [15] I. Ahmad, I. Ahmad, and E. Harjula, “Adaptive security at the edge for 6G-enabled healthcare IoT,” in 2026 Joint European Conference on Networks and Communications & 6G Summit (EuCNC/6G Summit), 2026. [16] Cilium Authors, “Introduction to cilium and hubble,” 2026, project documentation. [Online]. Available: https://docs.cilium.io/en/stable/ overview/intro/ [17] S. Chatzimiltis, M. B. Mashhadi, M. Shojafar, M. Debbah, and R. Tafazolli, “Agentic AI for 6G: A new paradigm for autonomous ran security compliance,” arXiv preprint arXiv:2512.12400, 2025. [18] H. Wen, P. Sharma, V. Yegneswaran, A. Gehani, P. Porras, and Z. Lin, “Mobillm: An agentic ai framework for closed-loop threat mitigation in 6g open rans,” in MILCOM 2025 - 2025 IEEE Military Communications Conference (MILCOM), 2025, pp. 1–6. [19] J. P. Anderson, “Computer security technology planning study,” Electronic Systems Division, Air Force Systems Command, Tech. Rep. ESDTR-73-51, 1972, the foundational reference monitor concept. [20] I. Ahmad, F. Shahid, I. Ahmad, J. Islam, K. N. Haque, and E. Harjula, “Adaptive lightweight security for performance efficiency in critical healthcare monitoring,” in 2024 18th International Symposium on Medical Information and Communication Technology (ISMICT), 2024, pp. 78–83. [21] Y. Zheng, Y. Hu, T. Yu, and A. Quinn, “AgentSight: System-level observability for AI agents using eBPF,” in Proceedings of SOSP, 2025. [22] D. Jin, A. J. Gaidis, and V. P. Kemerlis, “BeeBox: Hardening BPF against transient execution attacks,” in 33rd USENIX Security Symposium (USENIX Security ’24), 2024. [23] H. Sun and Z. Su, “Approximation enforced execution of untrusted Linux kernel extensions,” in 34th USENIX Security Symposium (USENIX Security ’25), 2025. [24] J. Jia, R. Qin, M. Craun, E. Lukiyanov, A. Bansal, M. Phan, and T. Xu, “Rex: Closing the language-verifier gap with safe and usable kernel extensions,” in 2025 USENIX Annual Technical Conference (USENIX ATC ’25), 2025. [25] K. Alam, M. A. Habibi, M. Tammen, D. Krummacker, W. Saad, M. D. Renzo, T. Melodia, X. Costa-Pérez, M. Debbah, A. Dutta, and H. D. Schotten, “A comprehensive tutorial and survey of o-ran: Exploring slicing-aware architecture, deployment options, use cases, and challenges,” IEEE Communications Surveys & Tutorials, vol. 28, pp. 1637–1678, 2026. [26] F. B. Schneider, “Enforceable security policies,” ACM Trans. Inf. Syst. Secur., vol. 3, no. 1, p. 30–50, Feb. 2000. [Online]. Available: https://doi.org/10.1145/353323.353382 [27] C. Sánchez, G. Schneider, W. Ahrendt, E. Bartocci, D. Bianculli, C. Colombo, Y. Falcone, A. Francalanza, S. Krstić, J. M. Lourenço et al., “A survey of challenges for runtime verification from advanced application domains (beyond software),” Formal Methods in System Design, vol. 54, no. 3, pp. 279–335, 2019. [28] M. Alshiekh, R. Bloem, R. Ehlers, B. Könighofer, S. Niekum, and U. Topcu, “Safe reinforcement learning via shielding,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018. [29] S. Gimhana, I. Ahmad, P. Porambage, and E. Harjula, “Mitigating DoS attacks in mMTC: An energy efficiency perspective,” in 2025 Joint European Conference on Networks and Communications & 6G Summit (EuCNC/6G Summit), 2025, pp. 199–204. [30] University of Oulu, “5GTN: 5g test network,” University of Oulu research infrastructure, 2024. [Online]. Available: https://5gtn.fi [31] Ollama, “Ollama,” 2026, local large language model serving framework. Accessed 2026-05-16. [Online]. Available: https://github.com/ollama/ ollama

Record · ID 965345 · SHA-256 08b31ffb2d57e5ea
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.