1
Type-Safe Decision Frameworks for Agentic 5G Control: A Theory-Driven Testbed Characterization of Where They Can Be Applied
arXiv:2609.33689v1 [cs.NI] 27 Sep 2026
Michail-Alexandros Kourtis and George Xilouris
Abstract—This paper presents a theory-driven characterization of type-safe decision frameworks for the agentic control of 5G networks, where every decision must be an element of a declared option set rather than free text. Three design points are evaluated on an Open5GS/UERANSIM testbed with a closed core-policy loop, namely a hosted typed model (Jev), an open fine-tunable typed encoder (Laya), and a zero-label retrofit of a general language model (AnyJev). The proposed theoretical framework turns timeliness, type conformance, certification cost and cardinality into checkable applicability predicates, supported by an optimal act/escalate/abstain gate, an escalation-feasibility floor, a co-location stability condition, per-type conformal risk control with a certification label floor, and a type-mismatch bound. Measuring every predicate yields an applicability map from framework to 5G decision class. Type safety removes format failures but not the question: the fine-tuned typed encoder returned its training answer for 98–99.5% of changed questions, and its calibrated gate then acted wrongly on up to 80% of them, whereas the question-reading frameworks acted wrongly on at most 0.143 (Jev) and 0.137 (AnyJev) of any changed question, but were either hosted and 11–29 times slower (Jev) or reliant on an 8B language model (AnyJev). Index Terms—Agentic AI, type safety, conformal risk control, O-RAN, 5G core network
I. I NTRODUCTION HE convergence of agentic Artificial Intelligence (AI) with the programmable control surfaces of 5G and 6G networks, namely the RAN Intelligent Controllers (RICs) of the Open Radio Access Network (O-RAN) and the servicebased 5G Core (5GC), has created the need for new requirements in the design and deployment of network control loops in terms of correctness, timeliness, auditability and energy efficiency [1], [2]. An agentic controller turns observations into actions: it selects a handover target, decides whether a slice’s Service-Level Agreement (SLA) is breached, grades an anomaly, or chooses which slice holds a priority flow. The network, however, can only execute an action that belongs to the set of options it exposed, which imposes a stringent contract on the controller’s output. Large Language Models (LLMs) answer in free text, and therefore fail this contract by format, by parsing, or by answering a different question than the one asked. Constraining their decoding to a grammar restores the output domain but reduces accuracy, since it removes the free-text reasoning pass on which the model relies [3], even when
T
M.-A. Kourtis and G. Xilouris are with the National Centre for Scientific Research “Demokritos” (NCSRD), Athens, Greece.
the constraint itself is enforced efficiently [4]. Recent agentic O-RAN designs address the same concern at the system level: A1gent compiles operator goals into typed A1 policy instances enforced by deterministic near-real-time (near-RT) applications [5], the Contract-based Agentic Intent Framework (CAIF) audits intents against formal RAN constraints before actuation [6], while xTRUCE and AURA arbitrate conflicting agent proposals with safety and stability guarantees [7], [8]. These designs govern actions after an agent has proposed them; they do not characterize the decision function that produces each proposal, which is where type safety, calibration and latency are decided. In this respect, a new class of type-safe decision frameworks has emerged, which returns, by construction, a probability distribution over the declared options instead of text. Jev is offered by TypeSafe AI as a hosted “System One” model with a stated end-to-end response time of 70–500 ms [9]; Laya is an open 421M-parameter typed encoder built on ModernBERT-large and trained with strictly proper scoring rules [10], [11]; and AnyJev is a layer from Nokia Applied Research that reads an open LLM as a typed decider, without training at its zero-label level, through the probabilities of the option letters [12]. Each framework guarantees the output type, yet none has been characterized against the constraints that 5G control actually imposes: the deadlines of the O-RAN control loops [13], questions whose type is declared only at runtime, the number of labels needed before an action can be trusted, and the co-location of several models on a single edge Graphics Processing Unit (GPU). This paper presents a theory-driven characterization of the three frameworks, which aims to answer what each one offers, what each one costs, and where each one can be applied in a 5G control stack. The contribution of the presented work is four-fold: (i) a theoretical framework with eight evaluation axes, whose results compose into applicability predicates for timeliness, type, coverage and cardinality, and which contributes a co-location stability condition, a certification label floor, and a type-mismatch bound that is attained by questionblind models; (ii) an experimental characterization of all three frameworks on an Open5GS/UERANSIM platform, which integrates a core-policy loop from the Network Data Analytics Function (NWDAF) through the Policy Control Function (PCF) and the Session Management Function (SMF) to the User Plane Function (UPF), with every endpoint pre-registered and reported whether it held or failed; (iii) an applicability map from framework to 5G decision class, where every cell is
2
either a measurement or a named violated predicate; and (iv) a set of findings that the map makes precise, specifically that type safety removes format failures but not the question, that the design points form a trade-off rather than a ranking, that a zero-shot framework is not zero-label for certified action, that free-text escalation to an 8B-class on-edge LLM is infeasible at or below 1 s on the evaluated edge GPU, and that a colocated LLM breaks a 20 Hz batch-1 decision loop at any escalation rate. The paper is organized as follows. Section II presents the type-safe frameworks and the related technologies. Section III defines the system model, and Section IV presents the theoretical framework. Section V describes the evaluation architecture and methodology, Section VI reports the characterization per axis, and Section VII derives where each framework applies. Section VIII discusses the architectural, methodological and deployment implications, Section IX states the limitations, and finally Section X concludes the paper. II. T YPE -S AFE D ECISION F RAMEWORKS AND R ELATED T ECHNOLOGIES A. Type-safe decision frameworks Jev is TypeSafe AI’s first “System One” model, described by its vendor as a function call that maps unstructured state to typed probabilistic decisions; it samples in parallel instead of generating tokens sequentially, it is trained with a method termed Reinforcement Learning for Calibrated Decisions (RLCD), and it is stated to answer in 70–500 ms end to end, 40–200× faster than LLMs on comparable tasks [9]. The model is proprietary and is accessed through a hosted Application Programming Interface (API), so the evaluated version (jev-1.13.0) is fixed only by its identifier. Laya is a non-autoregressive decision model with 421M parameters, composed of a ModernBERT-large backbone [11] and a decision head with an option-marker scorer and an act/escalate output; it is trained with reinforcement learning under strictly proper scoring rules, it is released under Apache-2.0, and it is stated to answer in about 33 ms per forward pass [10]. AnyJev turns an open LLM into a typed decider: its zero-label level L0 removes the position bias of the option letters by averaging over cyclic option orders and divides out a label-free prior, L1 adds temperature scaling on 100–500 labels, and L2 fits a closed-form head on a mid-depth hidden state from 100–300 labels per question layout, with heads shipped for the Qwen3 family [12], [14]. The three frameworks therefore span three distinct design points, namely a hosted specialist, an open finetunable specialist and an open retrofit, and each is type-safe by construction. None of them, however, states under which deadlines, question types and label budgets its guarantees hold in a network. B. Agentic control in O-RAN and the 5G Core The agentic view of the RAN organizes the task landscape around slice life-cycle management, Radio Resource Management (RRM) closed loops and cross-cutting security, and introduces planning, tool use and self-management gates
into long-lived control loops [2]. A1gent separates non-realtime (non-RT) agentic reasoning from near-RT deterministic execution through typed A1 policy instances and a fixedpriority action merger [5]; CAIF decouples probabilistic intent extraction from strictly governed policy execution and evaluates the design on an O-RAN platform for network slicing [6]. When several agents act on shared resources, xTRUCE places a provably safe arbiter in the near-RT RIC with a three-layer constraint hierarchy [7], and AURA shows on a live O-RAN system that two agents with individually correct objectives drive recurring opposing excursions, which an arbitration layer with feasibility invariants and dwell times removes [8]. OpenTwin certifies a digital twin by re-simulation and evaluates each xApp action before it executes, after observing that an E2 control request may be reported as successful while the base station never applies it [15]. A recent tutorial maps agentic capabilities onto the 5G and 6G control planes and standardization [1]. These works secure the composition of agent actions; the presented work is complementary and characterizes the decision function that produces each action, which is where type safety, calibration and latency are decided. C. Certified abstention and conformal methods in wireless networks Chow’s reject rule trades error against the rate of refusal [16], and selective classification extends the trade to deep networks [17]. Split conformal prediction [18] and Conformal Risk Control (CRC) [19] make such trades valid in finite samples under exchangeability, and role-stratified CRC assigns a separate threshold and risk budget to each argument role of an LLM tool call, so that a rare high-risk field is not averaged away [20]. In wireless networks, meta-learned context-dependent conformal prediction calibrates O-RAN applications when calibration and runtime contexts differ [21], post-hoc conformal prediction quantifies, after the fact, the miscoverage of prediction sets whose size is fixed by an operational constraint rather than by a prescribed level [22], and confounding-valid conformal inference bounds counterfactual Key Performance Indicators (KPIs) from logged telemetry complemented by scarce randomized telemetry [23]. Cascades that defer to a stronger model on low confidence expose a further attack surface, since an adversary can lower the weak model’s confidence and force deferral [24]. The presented work applies CRC to the joint event {ACT and wrong} per decision type, and shows why its guarantee must be indexed by the declared question type and not only by the output domain. D. Control timescales, serving and emulation O-RAN places near-RT control loops between 10 ms and 1 s and non-RT loops above 1 s [13], while 5GC analytics through the NWDAF [25] and policy authorization over the N5 interface [26] operate at the core-policy timescale. Stochastic network calculus has been used to bound the probability that the delay of Ultra-Reliable Low-Latency Communications (URLLC) services exceeds its budget in the near-RT loop [27]. An LLM acting as a second-stage decider is typically served with batched decoding [28] and quantized weights [29], which
3
fixes its time per output token and couples its latency to the load. Finally, open-source 5G platforms expose similar interfaces under different timing fidelity, so that functional compatibility does not imply timing fidelity [30]; the presented platform therefore scopes its RAN-side timing by a real-time factor. III. S YSTEM M ODEL AND T YPE S AFETY The proposed system model separates the declared decision from the framework that answers it, in order to express every property of the characterization as a property of a (framework, decision class) pair. Definition 1 (Decision class). A decision class d is a question template td with option schema Od , |Od | = Kd , a deadline Dd , a deadline-miss budget εd , a risk budget αd , a coverage floor cd , a network latency Lnet,d and a host. A decision type is the pair (td , Od ). Definition 2 (Type-safe framework). A System-1 F is typesafe for d if, for every state s, it returns a distribution pF (· | s,td ) on Od ; its action is arg max pF and its confidence m = max pF . A gate maps (m, σ ), with σ the slack left after System-1, to one of three actions: (i) ACT, which executes System-1’s option; (ii) ESC, which escalates the decision to a System-2 LLM; and (iii) ABS, which applies the class’s safe default. The loss of the default is ℓ0 , and β := 1 − ℓ0 is the accuracy an action needs in order to improve on it. A generative model can also be made type-safe, by masking at every decoding step the tokens that do not continue some option; a type-safe framework instead scores the options directly and returns a law over Od without any decoding. IV. T HEORETICAL F RAMEWORK The proposed framework organizes the characterization into eight axes, listed in Table I; the axes that carry the main results are backed by a formal result and measured in Section VI, and the proofs are given in the Appendix. A. The optimal gate and the near-RT collapse (A5) Let System-1 be calibrated within the class, P(correct | m) = m. Relative to the default, VACT (m) = m − β , VABS = 0 and VESC (m, σ ) = E[1{L2 ≤ σ }(C2 − β ) | m] − µP(L2 > σ | m) − κ, where C2 is System-2’s correctness, L2 its latency, µ ≥ 0 the multiplier on deadline misses, κ ≥ 0 the escalation cost, and F2 (σ ) := P(L2 ≤ σ | m) the conditional latency distribution, which under the independence assumed below does not depend on m; the argument σ is omitted where clear. These values arise from the following program: the gate minimizes the expected escalation cost E[c(a)] over randomized rules (m, σ ) 7→ a ∈ {ACT, ESC, ABS}, subject to a per-class risk constraint E[ℓ | d] ≤ αd on the loss ℓ ∈ [0, 1] of the executed option and a per-class deadline-miss constraint P(miss | d) ≤ εd . With multipliers λ , µ ≥ 0 on the two constraints, dividing the Lagrangian by λ expresses every term in loss units, so that
TABLE I E VALUATION AXES , THE FORMAL RESULT BEHIND EACH , AND WHERE IT IS MEASURED . Axis
Formal object
Measured by
A0 type safety A1 decision quality A2 certified action A3 label floor A4 type conformance A5 timeliness A6 cardinality
output ∈ Od (Def. 2) proper-score decomposition
typed interface known-type accuracy
per-type CRC gate (Thm. 2)
Table III
Lemma 2 Prop. 2
label curves (Table IV) Fig. 2
Thm. 1, Lem. 1, Prop. 1 option budget KFmax
A7 rewording
robustness vs. insensitivity
A8 cost, sovereignty
J/decision; data leaves the domain
Figs. 3, 4 handover k ≤ 64 (Fig. 6) rewordings; counterfactual questions L4 energy; hosting
κ = c2 /λ is the escalation cost and µ the rescaled price of a missed deadline. Theorem 1 (Threshold gate). For fixed multipliers, the optimal gate picks arg max{VACT ,VESC ,VABS } pointwise. If L2 is independent of (C2 , m) and π2 (m) := P(C2 = 1 | m) ≤ m for all m, the ACT region is exactly {m ≥ β } for every slack. Corollary 1 (Collapse to act-or-abstain). The ESC region is empty at slack σ if F2 (σ ) = 0, i.e., System-2 cannot answer in time, or if L2 is independent of (C2 , m) and π2 (m) ≤ max(m, β ) for all m. The optimal gate is then Chow’s rule with the safe default [16]: ACT iff m ≥ τ, else ABS. It is worth noting that the corollary turns the escalation question into two measurable conditions, namely whether System-2 can answer within the slack and whether it is more accurate than System-1 where System-1 is unsure. Two remarks qualify the theorem in practice. First, the independence of L2 from (C2 , m) is an idealization: on the presented platform the System-2 answers that arrive early are the short ones, and short answers are more often correct, so that the value of escalating is not monotone in the slack unless a missed deadline is priced explicitly. The multiplier µ therefore has to reflect how the operator counts a miss, and a replay with µ = 0 escalated decisions whose answers then expired (Section VI). Second, the theorem is stated per decision class, and the classes in which π2 > m holds at all are an empirical property of the System-1/System-2 pair; the characterization measures them rather than assuming them. B. Escalation feasibility and co-location (A5) Lemma 1 (Service-time floor). A System-2 answer of n tokens takes at least TTFT+(n−1)·TPOT, with TTFT the time to the first token and TPOT the per-token decode time at concurrency 1, which is the shortest the server attains since batching only lengthens each decode step. Escalation is feasible within D only for answers with n ≤ 1 + (D − TTFT)/TPOT. Proposition 1 (Co-location stability). Let System-1 decisions arrive every T and be served in order, taking si while a co-
4
located System-2 is idle and sb while it decodes (si < T < sb ), and let System-2 be busy a fraction u2 of the time, alternating slowly relativeto sb . System-1 is stable iff u2 < u∗ := sb (T − si )/ T (sb − si ) , and within a System-2 busy period of length B its lateness grows to B(sb − T )/T . The proposition makes explicit a cost that a per-request latency figure hides: a System-1 that meets its deadline in isolation can still fall behind its period whenever it shares the GPU with a decoding System-2.
Open5GS core, Plane B (built) UEs, gNB (UERANSIM)
per-slice UPF
per-slice SMF
PFCP
N7
PCF N5
NWDAF
loop notify agent choice
System-2 Qwen3-8B (vLLM)
gate: act / abstain (escalate)
System-1 service: Jev | Laya | AnyJev (typed /v1/judge)
C. Certified action per type (A2, A3, A4) Let Z = (m, 1{wrong}) and calibrate on n labelled decisions of type t, restricting the threshold to the candidate set Tn = {m1 , . . . , mn , ∞}: τ̂ = min{τ ∈ Tn : (∑i 1{mi ≥ τ, wrongi } + 1)/(n + 1) ≤ α}, with τ̂ = ∞ if no finite threshold qualifies. Theorem 2 (Per-type gate [19]). If calibration and test decisions of type t are exchangeable, Pt (ACT ∧ wrong) ≤ α. For i.i.d. calibration from P and an independent test decision drawn from Q, the risk is at most α + dTV (PZ , QZ ). Lemma 2 (Certification label floor). The gate can act (τ̂ < ∞) only if n ≥ ⌈1/α⌉ − 1, i.e., 19 labels at α = 0.05 and 9 at α = 0.10, per decision type. Zero-shot frameworks therefore pay only this floor for a new type, while specialists pay their training labels on top: zero-shot is not zero-label for certified action. The floor is a necessary condition only: the ACT rate that the gate reaches at a given α grows with n, since the finite-sample term 1/(n + 1) forces a conservative threshold on small calibration sets, and a framework whose confidence separates correct from wrong decisions poorly may still act on almost nothing after calibration. The label cost of a new decision type is thus the sum of three terms, namely the training labels of a specialist, the calibration floor of Lemma 2, and the additional calibration labels that bring the ACT rate above the coverage floor cd . Proposition 2 (Type mismatch). If the gate of type t answers a ′ query of type t ′ ̸= t, then Pt ′ (ACT ∧ wrong) ≤ αt + dTV (PZt , PZt ), and no better bound holds in general: a System-1 that ignores the question has the same law of m under t and t ′ while its correctness can flip, so the risk approaches the ACT rate. For such a question-blind System-1 no test on m can detect the mismatch; only a check of the declared type can. The guarantee of Theorem 2 is thus a property of the (framework, type) pair, not of the framework.
Fig. 1. The presented evaluation architecture and the framework plug-in point. Each framework serves the same typed endpoint, and the core-policy loop (Plane B) runs end to end; Plane A (E2/near-RT RIC) is not closed on this platform.
E. Applicability predicates Proposition 3 (Applicability). Under the exchangeability of Theorem 2, F serves class d with P(ACT ∧ wrong) ≤ αd , P(miss) ≤ εd and ACT rate ≥ cd if (T) q1−εd (Lnet,d + LF (Kd )) ≤ Dd , where LF (Kd ) is F’s per-decision latency at Kd options and q1−ε the (1 − ε)-quantile, with the latency of the group-and-final reduction when Kd > KFmax and Proposition 1’s condition on a shared GPU; (Y) td is in F’s calibrated domain with at least ⌈1/αd ⌉ − 1 labels; (C) the gate’s ACT rate at αd is at least cd ; and (K) Kd ≤ KFmax , directly or through the group-and-final reduction. (Y) is necessary in the worst case (Prop. 2) and (T) by definition. The proposition is shallow by design: it turns the question of where F can be applied into a predicate that a measurement settles, which is the purpose of the characterization that follows. V. E VALUATION A RCHITECTURE AND M ETHODOLOGY The presented evaluation architecture is designed to provide a controlled and reproducible 5G environment, integrating an open-source 5G Core, an emulated RAN, a typed decision service and a GPU-hosted System-2. The system is structured into multiple layers in order to ensure that every framework is evaluated behind the same typed interface, on the same decisions and under the same timing conditions, as depicted in Fig. 1. The primary ambition of the architecture is to attribute every difference in the results to the framework, and not to the path around it. We implemented every layer with open-source components, except for the hosted framework, as detailed below.
D. Cardinality (A6)
A. Architecture layers
Let KFmax be the largest option count that F scores in one call. A decision over K > KFmax options can be served by a round of groups of g ≤ KFmax options followed by a final among
Core network layer. This layer is realized by Open5GS 2.8.0 [31] with a Session Management Function (SMF) and UPF pair per slice, namely enhanced Mobile Broadband (eMBB), URLLC and massive Machine-Type Communications (mMTC). User Equipment (UE) and the gNodeB are emulated by UERANSIM [32] with two local patches, specifically the PDU session resource modify procedure and uplink QoS-flow classification, so that a
the ⌈K/g⌉ group winners, at the price of more than one call per decision and of an accuracy that is the product of the round accuracies; the analysis of this reduction is reported in an extended version of this work, and the present paper uses only its measured accuracy and latency where Kd > KFmax .
5
policy change reaches the UPF data path in both directions (UERANSIM commit 48554b7 with patches 0001 and 0002). An srsRAN [33] gNodeB over ZeroMQ is available on the platform, but no RAN-side timing result is reported in this paper. There is no real radio; RAN-side timing is scoped by a real-time factor [30] and lies outside the presented results. The decision host and the core run on different hypervisor nodes, and since the platform’s switch does not trunk the dedicated decision-path VLANs between them, the loop agent’s traffic to the System-1 endpoint and to the PCF traverses the shared 1 GbE management VLAN. This affects only transport latency, not the signalling path: every policy is installed through N5 at the PCF. Analytics and loop layer. An Open5GS-targeted NWDAF [34] observes logs and counters and notifies a loop agent, which turns each notification into a typed threeslice choice, sends it to the decision layer, and installs the chosen policy through N5 policy authorization at the Policy Control Function (PCF), from where it propagates to the SMF over N7 and to the UPF over the Packet Forwarding Control Protocol (PFCP). No 3GPP conformance of the NWDAF is claimed. Typed decision layer. This layer exposes one typed endpoint (/v1/judge) behind which each framework is plugged in turn. Laya is used zero-shot (“base”) and after full finetuning on the v2 training split (5 seeds); a plain encoder, namely ModernBERT-base [11] with one linear head per fixedcardinality class, is trained on the same split (5 seeds) as the task-specific reference. AnyJev reads Qwen3-8B (bf16) [14] zero-shot through cyclic-order marginalization of the optionletter softmax with a label-free batch prior (L0); its L2 variant fits a closed-form head per question layout on the training split, which the presented work pools per option count (†, an extension to the shipped levels). Jev (jev-1.13.0) is queried zero-shot over the Internet. Software versions: Laya 0.3.5 and AnyJev commit 3cd8c6f on PyTorch 2.6.0 (CUDA 12.4); Qwen3-8B (bf16) and Qwen3-8B-AWQ on vLLM 0.30.0. Every gate in this paper operates on each framework’s topoption probability max pF , not on the confidence field that some frameworks report next to it. Gate and escalation layer. The gate implements the conformal risk-control threshold of Theorem 2 on each framework’s top-option probability, fitted per decision class on the heldout calibration partition (Laya, and the encoder on its direct classes) or, where a framework has no calibration run, by 2-fold cross-fitting on the test split (Jev, AnyJev, and the encoder’s cells at k ≥ 16), i.e., each fold’s threshold is fitted on the other half of the test items; every such choice is logged as a pre-registration deviation. Cross-fitting halves the effective calibration size, so at k ≥ 16 the finite-sample term 1/(n + 1) of Theorem 2 is about 0.015, i.e., 30% of α = 0.05, and the corresponding cells of Table III should be read with that granularity. The escalation target is Qwen3-8B served by vLLM [28] with 4-bit Activation-aware Weight Quantization (AWQ) [29], co-located with System-1 on the edge GPU. Measurement layer. We measured System-1 latency on an NVIDIA L4 (the edge GPU) with the Streaming Multiprocessor (SM) clock locked at 1050 MHz after a 20 s settle, at
batch 1 from raw per-call samples, reported as p50 and p99, and energy per decision is computed as mean power times mean latency from 20 Hz power samples; we ran training and accuracy runs on an NVIDIA RTX 4090 whose timings are never reported. Jev’s latency includes the wide-area network and is reported for cold (a new TLS session per call) and warm connections, while its energy cannot be measured locally. The co-location and catch-up-batching runs (Section VI) serve Laya’s base checkpoint at k = 8 and 20 decisions/s, since latency does not depend on the weights, against Qwen3-8BAWQ under vLLM on the same locked L4. The trace-driven gate replay uses fine-tuned Laya with isotonic-calibrated confidence as System-1, System-2 correctness from the bf16 Qwen3-8B on the RTX 4090 and System-2 latency from the AWQ deployment on the L4, with π2 and F2 2-fold crossfitted on the test split; the AWQ deployment is less accurate on handover than the bf16 model, so bf16 correctness biases the replay toward escalation, and the set of escalated decisions was unchanged when AWQ correctness was substituted. B. Decisions, data and protocol The frozen v2 corpus comprises 3,645 test items over eight known decision classes, namely handover target selection with k ∈ {4, . . . , 64} neighbours, SLA breach, urgency and anomaly severity, with labels generated from simulator rules. Per class, the test partition holds 540 items for handover k = 4 and k = 8, urgency and anomaly severity, 135 items for each of handover k = 16, 32 and 64, and 1,080 items for SLA breach; the calibration and shift partitions have the same perclass sizes. Each item also carries a slice criticality, high (URLLC), medium (eMBB) or low (mMTC), which assigns it a declared deadline of 100, 500 or 1,000 ms in the gate replay, where the class risk budget is scaled by 0.5, 1 and 2 respectively. The five seeds vary the training run of the fine-tuned arms and share one calibration and test split. Two derived corpora probe generality: (i) eight runtime-declared templates (2,400 test items), which keep a v2 item’s state and options and ask a different question; and (ii) the v2 test items under four rewordings. Each v2 item carries a network state, i.e., slice-level Key Performance Measurements (KPMs) such as the 95th-percentile latency, the slice’s SLA latency target and per-neighbour load and Reference Signal Received Power (RSRP), together with a typed question of one of three judgment types: (i) choice, which selects one of k options, e.g., the handover target among k neighbour cells; (ii) yes/no, which returns a probability, e.g., whether the slice’s SLA is breached; and (iii) score, which selects an ordered level, e.g., the urgency of an intervention or the severity of an anomaly. The corpus is split into training, calibration, test and shift partitions, and the v2 corpus has been frozen since the first result. The runtime-declared templates ask, over the same states and options, among others for the most loaded neighbour, for the least loaded neighbour subject to an RSRP constraint, for the neighbour with the strongest signal, whether the jitter exceeds a fraction of the SLA, for a congestion level, for the root-cause KPI of an anomaly, and for the known handover question under renamed option keys. Four of them
6
TABLE II T HE DESIGN POINTS , WITH MEASURED VALUES : KNOWN - TYPE ACCURACY ( MEAN OVER HANDOVER k=4, 8, SLA, URGENCY, ANOMALY; TEST ), QUESTION - READING ( MEAN ACCURACY ON THE FOUR COUNTERFACTUAL TEMPLATES ), P 99 LATENCY AND ENERGY OF ONE k=4 DECISION ON THE L4 (J EV: END TO END , WARM ), AND LABELS NEEDED BEFORE CERTIFIED ACTION ON A NEW TYPE . known acc. question-reading p99 (ms) J/dec. Laya FT Laya base Encoder AnyJev-8B L0 AnyJev-8B L2† Jev
0.90 0.27 0.92 0.44 0.73 0.67
0.41 0.27 no readout 0.81 no readout 0.99
25.8 25.8 23.5 283 94 596
TABLE III C ERTIFIED GATE PER FRAMEWORK AND KNOWN CLASS (α = 0.05, TEST ): P(ACT ∧ WRONG) AND ACT RATE . L AYA AND THE ENCODER ’ S DIRECT CLASSES : 5- SEED MEAN , CALIBRATED ON THE HELD - OUT CALIBRATION PARTITION ; J EV, A NY J EV AND THE ENCODER ’ S CELLS AT k ≥ 16: 2- FOLD CROSS - FITTING ON THE TEST ITEMS ; n = 540 PER CLASS (135 AT k ≥ 16, 1,080 FOR SLA BREACH ). B OLD : ABOVE α , ALL WITHIN FINITE - SAMPLE NOISE ( ONE ITEM IS 0.001–0.007).
labels for a new type
1.10 training + 19 1.10 19 (gate rarely acts) 0.55 training (new head) + 19 19.6 19 6.6 head per layout + 19 n/m 19; data leaves the domain
change only the question while keeping the state and options of a v2 item, which isolates whether a framework reads the question at all. The rewordings comprise (i) the rule stated in words, (ii) the rule stated through the state’s field names, and (iii) two paraphrases written from a template rather than generated by any framework under evaluation. Shift penalties use a sample-split (Scheffé-set) estimate of total variation on Z, reported next to a null estimate on two halves of one sample. We pre-registered every endpoint before its run, with a noninferiority margin of 3 pts, paired bootstrap over items (10,000 resamples) and 5 seeds wherever training is involved, and every departure from a registration is logged. Byte-identical items reach every framework, and each cell states its label use. The applicability map evaluates Proposition 3 on the deadline grid {100 ms, 500 ms, 1 s, 5 s, 60 s}, which spans the near-RT, core-policy and non-RT timescales of Section II, with εd = 0.01 (so that (T) uses the p99 latency), αd = 0.05 and a coverage floor cd = 0.5, i.e., a framework counts as applicable only if its certified gate acts on at least half of the decisions; these parameters were fixed in the pre-registration, and the map is reported for a dedicated GPU, with the shared-GPU condition of Proposition 1 treated separately in Section VI. Table II summarizes the design points with the measurements of Section VI. VI. C HARACTERIZATION R ESULTS A. A0/A1: type safety and known-type quality Each framework returns a distribution over the declared options by construction, so format and parsing failures do not arise at the typed interface, and every framework is compared on the same options. On known types, the fine-tuned specialists reach specialist accuracy (Table II: encoder 0.92, Laya FT 0.90), while the zero-shot frameworks remain substantially lower on the rule-defined classes (Jev 0.67, AnyJev L0 0.44, Laya base 0.27); the exception is SLA breach, which Jev answers at 0.963. B. A2/A3: certified action and its label cost Table III shows that the risk promise holds for every framework in distribution; what differs is how often a framework can act at that risk, which is the practical utility of the gate. The specialists act on 43–99% of decisions, zero-shot AnyJev on at most 29% and Jev on 8–78%. Under the shift split, AnyJev
class
Laya FT
Encoder
risk ACT
risk ACT
handover k=4 0.043 handover k=8 0.042 handover k=16 0.024 handover k=32 0.047 handover k=64 0.040 SLA breached (yes/no) 0.052 urgency (score) 0.056 anomaly (score) 0.051
0.91 0.039 0.69 0.046 0.76 0.035 0.63 0.052 0.43 0.042 0.99 0.057 0.79 0.060 0.81 0.043
AnyJev-8B L0 risk
0.99 0.042 0.96 0.046 0.86 0.022 0.74 0.037 0.70 0.037 0.99 0.044 0.85 0.044 0.84 0.043
ACT
Jev risk ACT
0.29 0.032 0.12 0.044 0.09 0.070 0.09 0.038 0.07 0.037 0.05 0.037 0.08 0.026 0.06 0.042
0.35 0.23 0.34 0.24 0.22 0.78 0.21 0.08
TABLE IV L ABELS TO CERTIFIED ACTION FOR A NEW FAMILY, HELD OUT OF BOTH ARMS ’ TRAINING ( LEAVE - ONE - FAMILY- OUT, LOFO): LABELS n AT WHICH THE ARM FIRST COMES WITHIN 3 PTS OF ITS OWN FULL - DATA ACCURACY (5 SEEDS ), PLUS THE L EMMA 2 FLOOR . family SLA breached urgency handover anomaly
Laya (LOFO+FT)
encoder
ratio
100 300 1000 300
300 1000 >1000 100
3.00 3.33 > 1.00 0.33
zero-shot (Jev, AnyJev L0): 19 labels/type at α = 0.05
L0’s SLA gate is the one cell that leaves its budget: its joint risk rises to 0.199, about four times α, while its ACT rate rises from 0.05 to 0.26; the sample-split shift on that framework’s own Z is 0.35, so the cell stays within the bound α + dTV of Theorem 2, but the bound is loose there, and the cell shows that a zero-label confidence can be shift-sensitive even where the state distribution barely moves. It should be noted that a usable shift penalty requires a debiased estimate: the plug-in total variation on Z read 0.09–0.16 at high k, whereas the sample-split estimate read 0.007–0.034. Table IV presents the label axis. Both arms follow the same leave-one-family-out (LOFO) protocol: each is first trained on the other three v2 families, and then continued on n labels of the held-out family with the whole model updated, under the same small-budget schedule; the arms differ in their backbone, ModernBERT-large (421M) for Laya against ModernBERTbase for the encoder, which is a confound that this comparison does not separate. The pre-registered claim that the typed encoder needs three times fewer labels than a plain encoder held in 2 of 4 families: Laya is far more label-efficient on handover at 100–300 labels (0.726/0.828 against 0.225/0.321), but reaches its own parity only at 1,000, and on anomaly severity the plain encoder converges faster. At zero labels the held-out family is not served (e.g., handover 0.141). A new question type, however, is learnable: with 100 labels per runtime-declared template, continued full fine-tuning lifts Laya to 0.78–0.97 on all eight templates (“most loaded” 0.009 → 0.809), whereas 10 labels do not suffice (0.14–0.76) and
7
TABLE V Q UESTION - READING ON THE RUNTIME - DECLARED TEMPLATES ( ACCURACY, TEST, ZERO LABELS ; L AYA FT: 5- SEED MEAN ). L AST COLUMN : SHARE OF ITEM - SEED PAIRS ON WHICH L AYA FT RETURNED THE ANSWER IT GIVES TO THE SOURCE V 2 QUESTION . template most loaded neighbour least loaded, RSRP ≥ bound strongest signal jitter above SLA share weak neighbour congestion level root-cause KPI renamed keys
Laya base Laya FT AnyJev L0 0.197 0.143 0.153 0.603 0.570 0.287 0.343 0.193
0.009 0.691 0.437 0.493 0.461 0.351 0.329 0.825
0.990 0.713 0.927 0.597 0.743 0.573 0.167 0.690
Jev
same answer
0.987 0.987 1.000 1.000 1.000 1.000 0.803 0.677
0.993 0.991 0.995 0.982 – – – –
neither does training the head alone (0.123 on “most loaded” at 100). It can be deduced that the specialist must update its encoder, and not only its head, in order to read a new question.
C. A4: type conformance On templates that keep a known item’s state and options and change only the question, fine-tuned Laya returned its known-question answer for 0.982–0.995 of item-seed pairs; asked for the most loaded neighbour, it scored 0.009. Its calibrated gate kept acting (0.808 of those items), and 0.801 of all items ended as act-and-wrong, as depicted in Fig. 2. The pre-registered retention endpoint nonetheless passed (+0.138 [+0.114, +0.163]), since templates whose answer correlates with the trained rule dominate the pool; we report the endpoint as registered, with this measurement beside it. In contrast, Jev read the questions (act-and-wrong ≤ 0.032 on 6 of 7 templates), and AnyJev L0 did so on most templates. The mismatch evaluation covers the seven templates whose answer type has a v2 gate to apply, namely the handover k = 4 and k = 8 gates for the choice templates and the SLA-breach gate for the yes/no templates; the congestion-level template, a fourlevel score, has no v2 class of its answer type and is therefore not gated. The pre-registered outcome counts (safe / within the Proposition 2 bound / violating) at α = 0.05 were 6/1/0 for Jev, 4/3/0 for AnyJev L0 and 1/6/0 for Laya FT. Laya’s losses sit within the bound because its measured shift (0.18– 0.90) is as large as the loss itself: this is the worst case of Proposition 2, and the bound, although correct, requires labels of the new type in order to be evaluated. Only a declared-type check prevents the failure. Table V decomposes the result per template. The finetuned specialist gains exactly where the new question’s answer coincides with the trained rule, e.g., under renamed keys (0.193 → 0.825) and for the RSRP-constrained handover (0.143 → 0.691), and loses exactly where it inverts the rule, which is the signature of a decision function of (state, options, family) rather than of the question. The base checkpoint does not read the new questions either, since its accuracy stays near the chance level on most templates, whereas Jev reached at least 0.98 on six of the eight templates and AnyJev L0 reached 0.990 on the most-loaded question that defeats the specialist.
D. A5: timeliness System-1 timeliness separates the design points before any escalation is considered. The local specialists stay within about 30 ms at p99 up to k = 8 (Table II lists k = 4), whereas the zero-label AnyJev L0 grows with k, since it marginalizes over k cyclic option orders. Jev’s median latency is flat in k, as expected of a network-bound service whose TCP connection alone takes 199 ms at the median from the platform’s site, and its warm-connection p50 (281 ms at k = 4) already sits above the tightest near-RT classes. At concurrency 1, the co-located System-2 decodes at 22.1 ms/token after a 94 ms first token, so a 1 s class leaves a budget of 41 tokens, while the shortest System-2 answer was 47 tokens (45 quantized). Free-text escalation to this System2 at or below 1 s is therefore infeasible (Lemma 1), and by Corollary 1 the optimal gate in that regime is act-or-abstain. Co-location, in turn, costs more than the escalation itself (Fig. 3b): while System-2 decodes, each System-1 decision takes 67 ms at p99, above the 50 ms period, so that at one escalation every 50 s, 20% of decisions are already late. Proposition 1 places the stability boundary at a System-2 busy fraction u∗ ≈ 0.66 (with T = 50 ms, si = 36.2 ms and sb = 62.0 ms). Between 0.1 and 0.2 requests/s System-2’s offered load rises from 0.64 to 1.88, and its busy fraction, u2 ≈ 1 − e−ρ under the light-load approximation, from 0.47 to 0.85, i.e., across u∗ ; the System-1 median switches in the same interval from near its idle value (37.4 ms at 0.1 requests/s) to its busy value (62.0 ms), as the proposition predicts. The same proposition indicates the remedy, namely bringing the busy-period cost per decision below T . Serving every already-due decision in one forward pass (catch-up batching, Fig. 3c) achieves this: during decode bursts, batches of 2 form at 41.5 ms per decision (60.6 ms unbatched), and the loop-time p99 stays at 142–146 ms at every rate up to 1 request/s, against 2.0–29 s without batching. No decision misses a 250 ms deadline, but 12–56% miss 100 ms. Where escalation becomes feasible, Theorem 1 determines which decisions to escalate. In a trace-driven replay with measured System-1 and System-2 latencies and accuracies, we replaced the heuristic rule “escalate if System-2’s p99 fits” with the rule of Theorem 1 (termed R1 hereafter; pricing a miss like an error); this changed nothing at up to 2× the class deadlines, and at 60× reduced SLA violations from 0.068 to 0.044, escalating only SLA-breach items (14.2 per 3,645 decisions), which is the one class where System-2 is more accurate than System-1 (Fig. 4). A third gate, termed R1optimal, applies Theorem 1 with a tuned-threshold ACT set; it records the fewest SLA violations at every deadline scale, but it does so by abstaining more (coverage 0.629 at 60×), which is the coverage side of the same trade-off; the likefor-like comparison is therefore between the two gates that share the conformal ACT set and differ only in the escalation rule (conformal+R1 and conformal-deadline in Fig. 4). With µ = 0, i.e., when a missed deadline is not priced, the rule also escalated a few handover decisions whose answers then expired, as the first remark on Theorem 1 anticipates.
8
Laya FT
P(ACT ∧ wrong)
0.8
Laya base
AnyJev-8B L0
Jev
0.6 0.4 0.2 α = 0.05
0.0
ded
t loa
mos
max
P
RSR
r
dove
an tr. h cons
k wea
r
r
jitte
hbou
neig
root
caus
e
med
rena
keys
Fig. 2. Type mismatch: P(ACT ∧ wrong) when each framework’s known-class gate (α = 0.05) answers the seven gated runtime-declared templates (“max RSRP” is the strongest-signal template and “constr. handover” the RSRP-constrained one of Table V). Dashed: the α each gate promises on its own class.
TPOT p99
60
60
0.6
40
0.4 S1 late fraction
40
0.2
20
0.0
S2 busy ρ (scaled)
20
S1 p50 (ms)
1
2
4
8
16
32
0 10−1
64
System-2 concurrency c
S1 loop time p99 (ms)
0.8
fraction
TPOT p50
80
(c) catch-up batching
80
1.0
1 s class at c = 1: budget 41 tokens shortest D17 answer 47 tokens → 0% of answers fit
100
ms / token
(b) System-1 at 20 Hz under co-location
S1 p50 (ms)
(a) service curve
batch 1
10
4
batch ≤ 8 batch ≤ 16
103 250 ms
102
100
escalation rate re (req/s)
100 ms
10−1
100
escalation rate re (req/s)
Fig. 3. Timeliness on the shared L4 (System-1: Laya at k = 8, 20 decisions/s; System-2: Qwen3-8B-AWQ under vLLM). (a) System-2 service curve; the “shortest answer” annotation refers to the shortest free-text System-2 answer on the test split (47 tokens). (b) A 20 Hz System-1 under co-located escalations; the busy-fraction curve ρ in this panel is the System-2 busy fraction u2 of Proposition 1. (c) The same loop with catch-up batching of System-1 (loop time = completion − due time).
E. A6: cardinality Laya’s option budget truncates each option to a few tokens at k ≥ 16, so that direct choice there is at chance level, and AnyJev labels options with letters, which limits it to k ≤ 26. Above the option budget, the map of Section VII serves the local specialists through the group-and-final reduction of Section IV (groups of 8 followed by a final), with its measured accuracy, latency and certified gate (Table III); the analysis of where this reduction pays off is reported in an extended version of this work. F. A7: rewording Under four rewordings of the known questions, fine-tuned Laya’s accuracy moved by at most 0.4 pts (handover k ≤ 8: 0.908 at the original wording), and the pre-registered rewording endpoint held in 4 of 4 families. Given A4, this is insensitivity rather than robustness: an invariance test cannot tell the two apart, whereas a counterfactual question can. The plain encoder was stable under paraphrase, but fell from 0.948 to 0.263 when the handover rule was written into the question; AnyJev L2’s trained heads dropped 24–45 pts; and Jev gained when the rule was stated (anomaly 0.347 → 0.880). Certified
robustness to word edits by randomized smoothing [35], [36] did not separate the models within this budget: over five seeds, 75–84% of certified-correct decisions sit at the largest lower bound that 100 samples can give, so the certified radius reflects the sampling budget and not the model, and we make no radius claim. G. A8: energy and sovereignty In terms of energy efficiency, a k = 4 decision on the L4 costs 0.55 J with the encoder, 1.10 J with Laya and 19.6 J with AnyJev-8B L0, i.e., more than an order of magnitude more for the LLM-based question-reader, and among the LLM-based variants only the smaller AnyJev-1.7B with trained L2 heads comes close to the specialists, at the price of markedly lower accuracy on the known classes. Jev’s energy is not measurable locally, and its inputs, i.e., network state and KPIs, leave the operator’s domain, which turns sovereignty into a deployment criterion alongside latency and energy. Each framework was also placed behind the same typed endpoint of the corepolicy loop of Fig. 1, and every chosen policy change reached the SMF and the UPF data path; since those decisions were zero-shot and have no outcome labels, the loop establishes
9
0.09
up batching (Section VI), and the first applicable class is then about 250 ms. No zero-shot framework reaches the coverage floor on a known rule-defined class, except Jev on SLA breach. For new types, the question-reading frameworks are applicable, and only from 0.5–1 s: Jev 6/8 (warm connection) and AnyJev L0 4/8 at 1 s. The specialists fail (Y) or (C), with the exception of fine-tuned Laya on the renamed-keys template (1/8), which is its known question under new option names. Fig. 5 presents the same trade-off as a design space.
0.08
SLA violations
0.07 0.06 0.05 0.04 0.03 0.02 0.01
near-RT: ESC region empty (Cor. 1)
0.00 ×0.5
×1
×2
×10
×60
deadline scale (× 100 / 500 / 1000 ms per criticality) conformal-deadline (heuristic) conformal ACT set + R1 escalation (μ = 1) R1-optimal gate (μ = 1)
Fig. 4. Theorem 1 in a trace-driven replay: SLA violations vs. deadline scale (α = 0.05). The R1 escalation rule equals the heuristic deadline check wherever escalation is infeasible and improves on it where it opens. The R1optimal gate abstains more (coverage 0.629 vs. 0.861).
question-reading accuracy (counterfactual templates)
1.0
0.8
0.6
0.4
0.2
0.0 102
103
p99 latency of one k = 4 decision (ms, log) marker size = known-type accuracy; hollow = no readout for a new question
Laya FT: known-type acc. 0.90
AnyJev-1.7B L2†: known-type acc. 0.63
Laya base: known-type acc. 0.27
AnyJev-8B L0: known-type acc. 0.44
Encoder: known-type acc. 0.92
Jev (warm, WAN): known-type acc. 0.67
AnyJev-8B L2†: known-type acc. 0.73
Fig. 5. The trade-off space: latency, question-reading and known-type accuracy per framework variant.
integration rather than policy quality, and its per-framework analysis is reported in an extended version of this work. VII. W HERE E ACH F RAMEWORK A PPLIES Fig. 6 evaluates Proposition 3 for every framework, class and deadline from the measurements alone. For known types, the fine-tuned specialists are applicable from 100 ms on a dedicated GPU (encoder 8/8, Laya 7/8); with a decoding System-2 co-located on the same GPU the 100 ms column does not hold, since 12–56% of decisions are late even with catch-
Per framework. Jev offers question-reading at zero training labels: it answered the counterfactual templates at 0.99, gained on anomaly severity when the rule was stated in the question, and its gate was safe on six of seven new templates. Its cost is a network-bound latency that never reached 100 ms (warm p50 281 ms, cold 726 ms), low coverage on rule-defined known classes, and the loss of sovereignty over the decision inputs. Laya offers a fast, local and fine-tunable typed head (p99 25.8 ms; 1.10 J), which reaches specialist accuracy on its trained questions, although its certified ACT rate at handover k = 64 (0.43, Table III) falls below the coverage floor, which makes handover k = 64 the one known class it does not serve in Fig. 6; after fine-tuning it does not read the question, so it must only serve declared types on which it was calibrated, and a new type costs about 100 labels and a full fine-tuning pass (0.78–0.97 on the new templates). Where its label efficiency exceeds that of a plain encoder (handover at 100–300 labels), Laya is the natural specialist to stand up quickly; otherwise a plain encoder performs as well or better, at approximately half the energy per decision. AnyJev offers an open question-reader: L0 read the counterfactual templates at 0.81, but it is slow (p99 873 ms at k = 8), and acts rarely at a strict risk budget, while its trained L2 heads are fast but inherit the layout-bound failure of any per-type head. Trade-off. The characterization highlights a practical tradeoff between speed and generality that no single design point resolves. The local specialists meet near-RT deadlines at 0.55– 1.10 J per decision but answer only the questions on which they were trained, whereas the question-reading frameworks generalize to runtime-declared questions at the price of a latency that excludes the classes below about 500 ms, an order of magnitude more energy (AnyJev L0) or data that leave the operator’s domain (Jev). Type safety is common to all three and settles none of these costs. Deployment guidance. The results translate into four rules: (i) dispatch on the declared decision type and not on confidence, sending known types to a local specialist behind a pertype gate and new types to a question-reading model or to the safe default; (ii) treat near-RT classes as act-or-abstain, do not provision free-text escalation to an 8B-class on-edge LLM at or below 1 s, and never co-locate a decoding System-2 with a batch-1 System-1 that serves a periodic loop; with catch-up batching, co-location is safe for deadlines of about 250 ms and above, but not for 100 ms classes; (iii) budget ⌈1/α⌉−1 labels per new type even for zero-shot frameworks, and hundreds to a thousand for a specialist; and (iv) weigh hosted questionreading against sovereignty, since the data leave the domain.
10
Known types (v2 classes, 8)
New types (runtime-declared templates, 8)
Encoder
8/8
8/8
8/8
8/8
8/8
0 (Y)
0 (Y)
0 (Y)
0 (Y)
0 (Y)
Laya FT
7/8
7/8
7/8
7/8
7/8
1/8
1/8
1/8
1/8
1/8
Laya base
0 (C)
0 (C)
0 (C)
0 (C)
0 (C)
0 (C)
0 (C)
0 (C)
0 (C)
0 (C)
AnyJev-8B L2†
2/8
2/8
2/8
2/8
2/8
0 (Y)
0 (Y)
0 (Y)
0 (Y)
0 (Y)
AnyJev-1.7B L2†
1/8
1/8
1/8
1/8
1/8
0 (Y)
0 (Y)
0 (Y)
0 (Y)
0 (Y)
AnyJev-8B L0
0 (T)
0 (C)
0 (C)
0 (C)
0 (C)
0 (T)
2/8
4/8
4/8
4/8
Jev (warm)
0 (T)
1/8
1/8
1/8
1/8
0 (T)
2/8
6/8
6/8
6/8
Jev (cold)
0 (T)
0 (T)
0 (C)
1/8
1/8
0 (T)
0 (T)
3/8
6/8
6/8
Qwen3-8B gen. (ref.)
0 (T)
0 (T)
0 (T)
0 (T)
0 (Y)
0 (T)
0 (T)
0 (T)
0 (T)
0 (Y)
100 ms
500 ms
1s
5s
60 s
100 ms
500 ms
1s
5s
60 s
deadline D (q0.99 latency)
deadline D (q0.99 latency)
Fig. 6. Applicability map (Proposition 3): per framework variant and deadline, the number of decision classes for which timeliness (T), type (Y), coverage (C) and cardinality (K) all hold at α = 0.05, cd = 0.5 on a dedicated GPU; “0 (P)” names the most frequent violated predicate. †: extension with trained L2 heads.
TABLE VI P RE - REGISTERED ENDPOINTS OF THE ANALYSES REPORTED IN THIS PAPER AND THEIR OUTCOMES ( ALL REPORTED AS REGISTERED ). Endpoint
Measured
Outcome
Label efficiency (ratio ≥ 3 on ≥ 3 families) Retention on new templates (FT − base) Rewording (Laya drop ≤3 pts where AnyJev L2 ≥10) Type mismatch: questionreaders safe or within bound; Laya FT violates R1 gate escalates only where π2 > max(m, β ), equals the deadline check at ≤ 2× 1 s class infeasible at c=1 (Lemma 1)
2 of 4 families
failed
+0.138 [+0.114, +0.163]
held; questionblind held; insensitivity
4 of 4 families
Jev 0, AnyJev 0, Laya FT 0 half held violations only SLA breach; identical at ≤ held 2× 41 tokens vs. ≥45
held
VIII. D ISCUSSION Typed dispatch as an architectural pattern. The results argue for moving the notion of type from the output domain to the decision itself. A holistic control stack can declare, for every decision it issues, the question template and option schema together with the deadline class, and route the decision on that declaration: to a local specialist behind a per-type gate when the type is known and calibrated, to a question-reading framework when the type is new and the deadline allows it, and otherwise to the safe default. The declared type then carries the certificate of Theorem 2, and Proposition 2 explains why no confidence signal can replace it. This pattern is compatible with agentic O-RAN designs that already exchange typed policy instances between the non-RT and near-RT tiers [5], and with arbitration layers that govern actions after they are proposed [7], [8]; the presented characterization supplies the per-type timing, coverage and label figures that such a dispatcher needs. Evaluation methodology. Two pre-registered generality endpoints passed while the underlying capability was absent, which carries a lesson beyond the frameworks evaluated here. Invariance tests, i.e., rewording a question and checking that
the answer does not change, cannot distinguish a model that understands the question from one that ignores it. Counterfactual questions, which keep the state and options and change the question so that the correct answer changes, can. Characterizations of decision models for network control should therefore include counterfactual questions, and should report the share of unchanged answers next to accuracy. Energy and sovereignty. Energy per decision separates the design points by more than an order of magnitude at equal option count (Section VI), and in the presented measurements the most energy efficient choice for known types, namely the plain encoder, was also the most accurate one. The hosted framework removes the local energy cost from the operator’s accounting but moves the decision inputs, i.e., network state and KPIs, outside the operator’s domain. Both properties belong in the deployment decision next to latency and accuracy, and neither is visible in an accuracy-only comparison. IX. L IMITATIONS The presented characterization carries the following limitations, which are stated plainly. (1) The pre-registered labelefficiency claim held in only 2 of 4 families, as summarized in Table VI. (2) Fine-tuned Laya is question-blind, and two pre-registered generality endpoints passed only because of it. (3) Plane A (E2/near-RT RIC) is not closed, so the near-RT results are measured in isolation and by replay. (4) There is no real RAN, and RAN timing is scoped by a real-time factor. (5) Labels come from simulator rules. (6) One 8B LLM serves as both System-2 and accuracy ceiling. (7) The shift split barely moves the state distribution, so the shift results are weak evidence, with the one confidence-level exception noted in Section VI. (8) Jev is hosted and versioned (jev1.13.0), and its behaviour may change. (9) The analyses of the group-and-final reduction and of the per-framework core-policy loop, together with their pre-registered endpoints, are not part of this paper; the map uses only the measured accuracy, latency and gate of the reduction. (10) Questioncontrastive fine-tuning and GPU partitioning (NVIDIA Multi-
11
Process Service, MPS, not run on the shared edge GPU) were not measured, and the applicability map is computed for a dedicated GPU. (11) The label-efficiency comparison does not separate the typed head from the backbone size, since Laya’s backbone is ModernBERT-large and the plain encoder’s is ModernBERT-base. (12) Certified robustness radii are bounded by the certification sample budget (n = 100). X. C ONCLUSION This paper presented a theory-driven characterization of type-safe decision frameworks for agentic 5G control, which aims to establish where each design point can be applied rather than to rank them. The manuscript not only presented a theoretical framework whose results compose into checkable applicability predicates, but also measured every predicate for a hosted, an open fine-tunable and a retrofit framework on an Open5GS/UERANSIM platform with a closed core-policy loop and an NVIDIA L4 edge GPU. The characterization showed that type safety makes agentic control executable, but that safety of the output type is not safety of the decision: a fine-tuned typed encoder answered its training question for 98–99.5% of changed questions while its calibrated gate kept acting, whereas the question-reading frameworks acted wrongly on at most 0.143 and 0.137 of any changed question, at 11–29 times the latency for the hosted Jev and at an 8B language model’s energy cost for AnyJev. In addition, freetext escalation to the evaluated 8B-class on-edge LLM proved infeasible at or below 1 s, a co-located LLM made a batch-1 20 Hz loop late at any escalation rate unless System-1 batched its due decisions, and certified action required at least 19 labels per new decision type even for zero-shot frameworks. In contrast to evaluations that rank models by accuracy, the presented applicability map assigns known decision types to local specialists from 100 ms on a dedicated GPU (from about 250 ms when a decoding System-2 shares it) and runtimedeclared types to question-reading frameworks from 0.5–1 s, subject to energy and sovereignty constraints. Future work will focus on question-contrastive fine-tuning of typed encoders, on closing the near-RT loop through the E2 interface with a real RAN, and on GPU partitioning between System-1 and System-2. ACKNOWLEDGMENT The research leading to these results has been supported by the SHARC project (Grant Agreement No. 101290994) and the PQ-NEXT project (Grant Agreement No. 101225759). R EFERENCES [1] M. Ameur, A. Mekrache, B. Brik, and A. Ksentini, “LLM-powered agentic AI for 5G/6G networks: A tutorial and survey on architectures, protocols, and standardization,” arXiv:2607.16066, 2026. [2] Z. He, Y. Luo, X. Liu, M. B. Mashhadi, M. Shojafar, M. Debbah et al., “Agentic AI-RAN: Enabling intent-driven, explainable and self-evolving Open RAN intelligence,” arXiv:2602.24115, 2026. [3] Z. R. Tam, C.-K. Wu, Y.-L. Tsai, C.-Y. Lin, H.-y. Lee, and Y.-N. Chen, “Let me speak freely? A study on the impact of format restrictions on large language model performance,” in Proc. Conf. Empirical Methods in Natural Language Processing: Industry Track (EMNLP), 2024. [4] B. T. Willard and R. Louf, “Efficient guided generation for large language models,” arXiv:2307.09702, 2023.
[5] H. Li, D. Xu, M. Chen, and Y. Liu, “Agentic Open RAN: A deterministic and auditable framework for intent-driven radio control,” arXiv:2604.13384, 2026. [6] F. A. Bimo, C.-K. Lai, Z.-Y. Yang, and R.-G. Cheng, “Contractbased agentic intent framework for network slicing in O-RAN,” arXiv:2603.01663, 2026. [7] L. Xia, R. Q. Hu, P. S. Kudyba, Z. An, and H. Sun, “xTRUCE: A provably safe arbiter for multi-xApp conflict mitigation in agentic ORAN,” arXiv:2608.28532, 2026. [8] S. B. Hashemi Natanzi and B. Tang, “Taming the agentic RAN: Stability-guaranteed arbitration of autonomous AI agents in O-RAN,” arXiv:2609.18857, 2026. [9] D. Almeida, “Introducing System One models & Jev,” TypeSafe AI blog, https://typesafe.ai/blog/introducing-system-one-models-and-jev, Sep. 2026, model version evaluated: jev-1.13.0. [10] Convai Innovations, “Laya: a non-autoregressive typed decision model,” Hugging Face model card, https://huggingface.co/convaiinnovations/ laya, 2026, version 0.3.5, English checkpoint, 421M parameters, Apache-2.0. [11] B. Warner, A. Chaffin, B. Clavié, O. Weller, O. Hallström, S. Taghadouini et al., “Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference,” arXiv:2412.13663, 2024. [12] J. Zhang, T. Yang, Y. Shi, and L. Wu, “AnyJev: Turn any LLM into a Jev-style decision model,” https://github.com/nokia-applied-research/ AnyJev, 2026, commit 3cd8c6f, Apache-2.0. [13] O-RAN Alliance, “O-RAN architecture description,” O-RAN Alliance, Work Group 1 (Use Cases and Overall Architecture), Tech. Rep. ORAN.WG1.OAD-R003-v10.00, 2023. [14] A. Yang et al., “Qwen3 technical report,” arXiv:2505.09388, 2025. [15] Z. Zhang, M. S. Hossen, D. Ron, V. K. Shah, and Y. Liu, “OpenTwin: Closed-loop digital twins for trustworthy policy deployment in Open RAN,” arXiv:2605.24662, 2026. [16] C. K. Chow, “On optimum recognition error and reject tradeoff,” IEEE Trans. Inf. Theory, vol. 16, no. 1, pp. 41–46, 1970. [17] Y. Geifman and R. El-Yaniv, “Selective classification for deep neural networks,” in Advances in Neural Information Processing Systems (NeurIPS), 2017. [18] V. Vovk, A. Gammerman, and G. Shafer, Algorithmic Learning in a Random World. Springer, 2005. [19] A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, and T. Schuster, “Conformal risk control,” in Proc. Int. Conf. Learning Representations (ICLR), 2024. [20] M. A. Rahman, M. A. Rahman, N. H. Samin, K. R. Tasnia, M. H. Amin, S. R. Ahona et al., “Beyond aggregate risk: Role-stratified conformal risk control for LLM tool calls,” arXiv:2607.24343, 2026. [21] S. Yoo, S. Park, P. Popovski, J. Kang, and O. Simeone, “Calibrating wireless AI via meta-learned context-dependent conformal prediction,” arXiv:2501.14566, 2025. [22] X. Su, M. Zhu, O. Simeone, and C. Fischione, “Post-hoc conformal prediction for reliable wireless communications,” arXiv:2609.26625, 2026. [23] A. Qchohi, J. Moysen Cortes, and M. Zecchin, “Confounding-valid conformal inference for counterfactual KPIs in wireless networks,” arXiv:2609.05073, 2026. [24] Z. Liu, Y. Zeng, Y. Chang, and L. Lin, “Forced deferral: Manipulating routing decisions in multimodal LLM cascades,” arXiv:2606.15308, 2026. [25] 3GPP, “Architecture enhancements for 5G system (5GS) to support network data analytics services,” 3rd Generation Partnership Project, Technical Specification TS 23.288, V18.0.0, 2022. [26] ——, “5G system; policy authorization service; stage 3,” 3rd Generation Partnership Project, Technical Specification TS 29.514, V20.0.0, 2026. [27] O. Adamuz-Hinojosa, L. Zanzi, V. Sciancalepore, A. Garcia-Saavedra, and X. Costa-Pérez, “ORANUS: Latency-tailored orchestration via stochastic network calculus in 6G O-RAN,” arXiv:2401.03812, 2024. [28] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng et al., “Efficient memory management for large language model serving with PagedAttention,” in Proc. ACM Symp. Operating Systems Principles (SOSP), 2023. [29] J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen et al., “AWQ: Activationaware weight quantization for on-device LLM compression and acceleration,” in Proc. Machine Learning and Systems (MLSys), 2024. [30] R. Barker, T. Seyfi, A. E. Dorcheh, J. Boone, F. Afghah, and J. Boccuzzi, “AtlasRAN: Timing-aware evaluation of open-source 5G platforms for integrated wireless testbeds,” arXiv:2603.14661, 2026. [31] Open5GS, “Open5GS: open-source 5G core,” https://open5gs.org, version 2.8.0.
12
[32] A. Güngör, “UERANSIM: open-source 5G UE and RAN (gNodeB) simulator,” https://github.com/aligungr/UERANSIM. [33] Software Radio Systems, “srsRAN project,” https://www.srsran.com. [34] cem8kaya, “open5gs-nwdaf,” https://github.com/cem8kaya/ open5gs-nwdaf, commit 5619630 with one local patch. [35] J. Cohen, E. Rosenfeld, and J. Z. Kolter, “Certified adversarial robustness via randomized smoothing,” in Proc. Int. Conf. Machine Learning (ICML), 2019. [36] J. Zeng, J. Xu, X. Zheng, and X. Huang, “Certified robustness to text adversarial attacks by randomized [MASK],” Computational Linguistics, vol. 49, no. 2, pp. 395–427, 2023. [37] E. Altman, Constrained Markov Decision Processes. Chapman & Hall/CRC, 1999. [38] R. M. Loynes, “The stability of a queue with non-independent interarrival and service times,” Math. Proc. Cambridge Philos. Soc., vol. 58, no. 3, pp. 497–520, 1962.
A PPENDIX P ROOFS Proof of Theorem 1: The program of Section IV is a linear program over randomized rules; under Slater’s condition (some rule meets both constraints strictly) strong duality holds [37], and the Lagrangian, an expectation over (m, σ ) of the chosen action’s value, is maximized pointwise. Under independence, VESC = F2 (π2 − β + µ) − µ − κ. If π2 ≤ m, then VACT −VESC ≥ (1 − F2 )(m − β + µ) + κ ≥ 0 on {m ≥ β }, while on {m < β }, VACT < 0 = VABS . Proof of Corollary 1: If F2 (σ ) = 0, then VESC = −µ −κ ≤ 0. Otherwise, under independence, VESC = F2 (π2 − β + µ) − µ − κ ≤ F2 (π2 − β ) ≤ max(π2 − β , 0) ≤ max(VACT ,VABS ). Proof of Lemma 1: After the first token, decoding emits one token per step, and the concurrency-1 step time is the shortest the server attains, so the sum is a lower bound that is attained at concurrency 1. Proof of Proposition 1: The server’s service rate is 1/sb during System-2’s busy time and 1/si during its idle time, so its long-run capacity is u2 /sb + (1 − u2 )/si decisions per unit time when the alternation is slow relative to sb ; the load condition 1/T < u2 /sb + (1 − u2 )/si of a single server with modulated service [38] rearranges to u2 < sb (T − si )/(T (sb − si )). During a busy period one decision is served per sb while one arrives per T , so the backlog grows by sb − T per period. Proof of Theorem 2: The loss Li (τ) = 1{mi ≥ τ, wrongi } lies in [0, 1] and is non-increasing in τ. With Tn the candidate set of τ̂, let τ̂ ′ be the smallest τ ∈ Tn ∪ {mn+1 } with 1 n+1 ∑i≤n+1 Li (τ) ≤ α. Since Ln+1 ≤ 1, τ̂ is feasible for this problem, so τ̂ ′ ≤ τ̂ and, by monotonicity, Ln+1 (τ̂) ≤ Ln+1 (τ̂ ′ ). The threshold τ̂ ′ is a symmetric function of n+1 exchangeable 1 decisions, hence E[Ln+1 (τ̂ ′ )] = E[ n+1 ∑i Li (τ̂ ′ )] ≤ α, following conformal risk control [19]. For the shift, condition on the calibration set, which fixes τ̂; the loss is a [0, 1]-valued function of Z, so EQ L ≤ EP L + dTV (PZ , QZ ), and averaging over the calibration set returns the marginal bound. Proof of Lemma 2: At τ = ∞ the left side of the criterion is 1/(n + 1); any finite τ therefore needs 1/(n + 1) ≤ α. Proof of Proposition 2: The bound is Theorem 2 with ′ Q = Pt . For tightness, fix the calibration set (hence τ̂t ), give m the same law under t and t ′ , let a = P(m ≥ τ̂t ) be the ACT rate, let the gate of type t have Pt (ACT ∧ wrong) = αt exactly, let every acted decision be wrong under t ′ , and couple correctness
on the non-acted decisions identically under both types. Then the two laws of Z differ only in the mass a − αt that moves ′ from acted-and-correct to acted-and-wrong, so dTV (PZt , PZt ) = a − αt , while Pt ′ (ACT ∧ wrong) = a = αt + dTV : the bound is attained. Proof of Proposition 3: Compose Theorem 2 (risk), Lemma 1 and Proposition 1 (time) and Lemma 2 (labels); (K) holds by the definition of KFmax , with the latency of the reduction entering (T), and (C) is the coverage condition itself.